8.2 KiB
Hermes Model Usage Report
2026-08-09 | 30-Day Window (Jul 10 – Aug 9, 2026)
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)
Headline Numbers (Last 7 Days)
| Metric | deepseek-v4-pro | claude-sonnet-5 |
|---|---|---|
| Call volume | 14,024 | 154 |
| Spend | $39.20 | $7.30 |
| Avg cost/call | $0.0028 | $0.0474 |
| Share of calls | 98.9% | 1.1% |
| Share of spend | 84.3% | 15.7% |
| Est. 30-day spend | ~$183 | ~$44 |
DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.
Sonnet 5 Daily Breakdown
| Date | Calls | Spend | Context |
|---|---|---|---|
| Aug 9 (today) | 3 | $0.05 | Early, still running |
| Aug 8 | 73 | $6.56 | Audit remediation — subagent cascading |
| Aug 7 | 11 | $0.23 | Normal dev day |
| Aug 6 | 6 | $0.01 | Model eval / testing |
| Aug 5 | 20 | $0.14 | |
| Aug 4 | 19 | $0.15 | |
| Aug 3 | 5 | $0.03 | Weekend |
| Aug 2 | 20 | $0.13 | |
| Aug 1 | 43 | $4.17 | Elevated — subagent routing |
| Jul 31 | 146 | $17.78 | Hit $20 daily cap — 89% of day's spend was Sonnet 5 |
| Jul 30 | 62 | $6.83 | |
| Jul 29 | 7 | $0.00 | |
| Jul 28 | 0 | $0.00 | |
| Jul 27 | 2 | $0.00 | |
| Jul 26 | 1 | $0.00 | |
| Jul 25 | 52 | $3.34 | |
| Jul 24 | 87 | $6.72 | |
| Jul 23 | 4 | $0.00 | |
| Jul 22 | 0 | $0.00 | |
| Jul 21 | 0 | $0.00 | |
| Jul 20 | 0 | $0.00 | |
| Jul 19 | 0 | $0.00 | |
| Jul 18 | 1 | $0.00 | |
| Jul 17 | 0 | $0.00 | |
| Jul 16 | 0 | $0.00 | |
| Jul 15 | 0 | $0.00 | |
| Jul 14 | 2 | $0.00 | |
| Jul 13 | 0 | $0.00 | |
| Jul 12 | 34 | $12.75 | Model eval pipeline |
| Jul 11 | 0 | $0.00 | |
| Jul 10 | 12 | $0.02 |
Typical normal day: ~11 Sonnet 5 calls, ~$0.25/day
Anomaly days: Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month
30-Day Daily Spend Trend
Date Total Spend Sonnet 5 Total Calls Sonnet Calls
Aug 09 $1.42 $0.05 362 3
Aug 08 $16.00 $6.56 3,147 73
Aug 07 $6.44 $0.93 1,787 42
Aug 06 $3.06 $0.01 1,493 11
Aug 05 $7.96 $0.14 3,147 20
Aug 04 $5.14 $0.15 1,679 19
Aug 03 $2.91 $0.03 996 5
Aug 02 $5.72 $0.13 2,556 20
Aug 01 $10.29 $4.17 2,354 43
Jul 31 $20.01 $17.78 1,130 146 ⬅ cap hit
Jul 30 $10.29 $6.83 2,073 62
Jul 29 $3.60 $0.00 2,660 7
Jul 28 $3.41 $0.00 2,420 0
Jul 27 $1.38 $0.00 1,192 2
Jul 26 $1.02 $0.00 332 1
Jul 25 $6.29 $3.34 985 52
Jul 24 $59.71 $6.72 1,066 87
Jul 23 $37.92 $0.00 1,023 4
Jul 22 $67.03 $0.00 1,042 0
Jul 21 $4.77 $0.00 140 0
Jul 20 $4.20 $0.00 1,301 0
Jul 19 $0.81 $0.00 109 0
Jul 18 $0.17 $0.00 90 1
Jul 17 $2.17 $0.00 168 0
Jul 16 $1.42 $0.00 396 0
Jul 15 $5.83 $0.00 1,545 0
Jul 14 $2.34 $0.00 1,305 2
Jul 13 $57.49 $0.00 3,061 0
Jul 12 $99.86 $12.75 3,920 34 ⬅ biggest spike
Jul 11 $0.18 $0.00 1,656 0
Jul 10 $21.89 $0.02 4,544 12
August normal days: $3–8/day typical, $10–16/day on heavy remediation days
Prompt Caching Status
cache_hit = 0 across ALL models, ALL calls, ALL 30 days
Prompt caching is not enabled. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.
Caching Economics
Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):
| Scenario | Input $/M tokens |
|---|---|
| No caching (current) | $2.00 |
| Cache write (5 min TTL) | $2.50 |
| Cache write (1 hr TTL) | $4.00 |
| Cache hit | $0.20 (90% off) |
After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.
Projected savings for Hermes workload (large system prompts, repeated across turns):
| Cache hit rate | Input cost reduction | Monthly savings |
|---|---|---|
| 70% | –53% | ~$15–25 |
| 90% | –81% | ~$20–35 |
Key Spend Anomalies — Root Cause Analysis
| Date | Spend | Root Cause |
|---|---|---|
| Jul 12 | $99.86 | gpt-5.5 eval pipeline. 196 calls to gpt-5.5 ($66.42 — 67% of day) with 219K avg prompt tokens. 94 calls alone at 23:00 ($43.55 in one hour). deepseek-v4-flash added 1,846 eval calls ($2.87). Model catalog audit against all 128 models. Not Hermes. |
| Jul 13 | $57.49 | gpt-5.5 eval pipeline (continuation). 30 calls for $56.23 (98% of day) with 458K avg prompt tokens. Three overnight bursts: midnight ($15.36), 3 AM ($29.00), 4 AM ($11.87). DeepSeek V4 Pro handled all other traffic ($1.21). |
| Jul 22–24 | $37–67/day | gpt-5.5 → gpt-5.6-terra eval pipeline. Fewer calls (1,023–1,066) but 10–20× normal cost per call. gpt-5.5 at $63.93 (Jul 22), gpt-5.6-terra at $31.64 (Jul 23) and $50.22 (Jul 24). Avg prompt size: 309K–397K tokens. Each eval call cost $0.30–$1.50 vs normal $0.003. |
| Jul 31 | $20.01 | $20 daily cap breached. 146 Sonnet 5 calls ($17.78 — 89% of spend). Subagent delegation.model was pinned to claude-sonnet-5, bypassing the conductor's model routing. Fixed Aug 1 by switching delegation back to deepseek-v4-pro. |
All four anomalies share a common root: the July model evaluation pipeline hitting gpt-5.5 and gpt-5.6-terra through admin-ai with enormous evaluation-sized contexts. These models were never in Hermes' production chain — the eval runner discovered them in the proxy catalog and tested them. The Jul 31 event was a separate bug: subagent delegation config hard-overriding to Sonnet 5.
Data Source Limitation
This report only covers LiteLLM-proxied traffic (admin-ai). It is blind to direct fallback provider spend.
The fallback chain operates outside admin-ai: deepseek direct → google direct → xai direct → anthropic direct. When admin-ai is unreachable or the daily cap is hit, traffic falls through to these keys. Spend there is invisible to LiteLLM SpendLogs.
Known gap: Aug 5 actual spend was ~$45 (per changelog) but LiteLLM shows only $7.96. The ~$37 delta went through direct provider keys.
Fix needed: Real-time cost monitoring requires a second feed polling each provider's usage API directly. Without it, a fallback cascade can silently burn through provider credits with no alert.
Verdict
The model chain is correct and working as designed. Cost is under control for the proxied path. The fallback path is a blind spot that needs monitoring.
- DeepSeek V4 Pro: 99% of calls, ~$5.60/day — the workhorse
- Sonnet 5: 1.1% of calls (~11/day typical), genuine rare override — not a silent runaway
- July's $271 in anomaly spend (Jul 12–24) was the model evaluation pipeline hitting non-production models — not Hermes
- August baseline: $3–8/day typical, $10–16/day on heavy remediation days
- The model chain doc matches reality:
deepseek-v4-proprimary,claude-sonnet-5for critical escalation
Three action items:
- Enable Anthropic prompt caching — 80–90% off cached input tokens. Client-side change only. Must be done before Sep 1 ($2→$3 base price increase).
- Implement fallback provider monitoring — direct API polling of DeepSeek, Google, xAI, and Anthropic usage endpoints. The LiteLLM SpendLogs are blind to ~30–50% of actual spend on failover days.
- Tag eval pipeline traffic — any automated model testing must use a dedicated LiteLLM key with its own budget cap. The July anomalies contaminated 30 days of production cost data.