5.8 KiB
Hermes Model Usage Report
2026-08-09 | 30-Day Window (Jul 10 – Aug 9, 2026)
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)
Headline Numbers (Last 7 Days)
| Metric | deepseek-v4-pro | claude-sonnet-5 |
|---|---|---|
| Call volume | 14,024 | 154 |
| Spend | $39.20 | $7.30 |
| Avg cost/call | $0.0028 | $0.0474 |
| Share of calls | 98.9% | 1.1% |
| Share of spend | 84.3% | 15.7% |
| Est. 30-day spend | ~$183 | ~$44 |
DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.
Sonnet 5 Daily Breakdown
| Date | Calls | Spend | Context |
|---|---|---|---|
| Aug 9 (today) | 3 | $0.05 | Early, still running |
| Aug 8 | 73 | $6.56 | Audit remediation — subagent cascading |
| Aug 7 | 11 | $0.23 | Normal dev day |
| Aug 6 | 6 | $0.01 | Model eval / testing |
| Aug 5 | 20 | $0.14 | |
| Aug 4 | 19 | $0.15 | |
| Aug 3 | 5 | $0.03 | Weekend |
| Aug 2 | 20 | $0.13 | |
| Aug 1 | 43 | $4.17 | Elevated — subagent routing |
| Jul 31 | 146 | $17.78 | Hit $20 daily cap — 89% of day's spend was Sonnet 5 |
| Jul 30 | 62 | $6.83 | |
| Jul 29 | 7 | $0.00 | |
| Jul 28 | 0 | $0.00 | |
| Jul 27 | 2 | $0.00 | |
| Jul 26 | 1 | $0.00 | |
| Jul 25 | 52 | $3.34 | |
| Jul 24 | 87 | $6.72 | |
| Jul 23 | 4 | $0.00 | |
| Jul 22 | 0 | $0.00 | |
| Jul 21 | 0 | $0.00 | |
| Jul 20 | 0 | $0.00 | |
| Jul 19 | 0 | $0.00 | |
| Jul 18 | 1 | $0.00 | |
| Jul 17 | 0 | $0.00 | |
| Jul 16 | 0 | $0.00 | |
| Jul 15 | 0 | $0.00 | |
| Jul 14 | 2 | $0.00 | |
| Jul 13 | 0 | $0.00 | |
| Jul 12 | 34 | $12.75 | Evaluation run |
| Jul 11 | 0 | $0.00 | |
| Jul 10 | 12 | $0.02 |
Typical normal day: ~11 Sonnet 5 calls, ~$0.25/day
Anomaly days: Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month
30-Day Daily Spend Trend
Date Total Spend Sonnet 5 Total Calls Sonnet Calls
Aug 09 $1.42 $0.05 362 3
Aug 08 $16.00 $6.56 3,147 73
Aug 07 $6.44 $0.93 1,787 42 ⬅ all Claude models
Aug 06 $3.06 $0.01 1,493 11
Aug 05 $7.96 $0.14 3,147 20
Aug 04 $5.14 $0.15 1,679 19
Aug 03 $2.91 $0.03 996 5
Aug 02 $5.72 $0.13 2,556 20
Aug 01 $10.29 $4.17 2,354 43
Jul 31 $20.01 $17.78 1,130 146 ⬅ cap hit
Jul 30 $10.29 $6.83 2,073 62
Jul 29 $3.60 $0.00 2,660 7
Jul 28 $3.41 $0.00 2,420 0
Jul 27 $1.38 $0.00 1,192 2
Jul 26 $1.02 $0.00 332 1
Jul 25 $6.29 $3.34 985 52
Jul 24 $59.71 $6.72 1,066 87
Jul 23 $37.92 $0.00 1,023 4
Jul 22 $67.03 $0.00 1,042 0
Jul 21 $4.77 $0.00 140 0
Jul 20 $4.20 $0.00 1,301 0
Jul 19 $0.81 $0.00 109 0
Jul 18 $0.17 $0.00 90 1
Jul 17 $2.17 $0.00 168 0
Jul 16 $1.42 $0.00 396 0
Jul 15 $5.83 $0.00 1,545 0
Jul 14 $2.34 $0.00 1,305 2
Jul 13 $57.49 $0.00 3,061 0
Jul 12 $99.86 $12.75 3,920 34 ⬅ biggest spike
Jul 11 $0.18 $0.00 1,656 0
Jul 10 $21.89 $0.02 4,544 12
Normal daily spend range (August): $3–8/day
Prompt Caching Status
cache_hit = 0 across ALL models, ALL calls, ALL 30 days
Prompt caching is not enabled. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.
Caching Economics
Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):
| Scenario | Input $/M tokens |
|---|---|
| No caching (current) | $2.00 |
| Cache write (5 min TTL) | $2.50 |
| Cache write (1 hr TTL) | $4.00 |
| Cache hit | $0.20 (90% off) |
After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.
Projected savings for Hermes workload (large system prompts, repeated across turns):
| Cache hit rate | Input cost reduction | Monthly savings |
|---|---|---|
| 70% | –53% | ~$15–25 |
| 90% | –81% | ~$20–35 |
Key Spend Anomalies
| Date | Spend | Root Cause |
|---|---|---|
| Jul 12 | $99.86 | Massive outlier — evaluation run or unbounded subagent loop |
| Jul 13 | $57.49 | Continuation |
| Jul 22–24 | $37–67/day | Heavy deepseek-v4-pro throughput (2–3× normal) |
| Jul 31 | $20.01 | $20 daily cap breached — 89% from Sonnet 5 (subagent routing override) |
Verdict
The model chain is correct and working as designed. Not under-resourced.
- DeepSeek V4 Pro: 99% of calls, $5.60/day average — the workhorse
- Sonnet 5: 1.1% of calls, genuine rare override — not a silent runaway
- Monthly spend has stabilized at $3–8/day in August, well within the $20–30/day budget
- The model chain doc now matches reality:
deepseek-v4-proprimary,claude-sonnet-5for critical escalation
One concrete action: Enable Anthropic prompt caching. It's a client-side change with no infrastructure cost — just set cache breakpoints on the system prompt. 80–90% off cached input tokens. Worth doing before Sep 1 when base prices rise 50%.