174 lines
8.2 KiB
Markdown
174 lines
8.2 KiB
Markdown
# Hermes Model Usage Report
|
||
**2026-08-09** | 30-Day Window (Jul 10 – Aug 9, 2026)
|
||
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)
|
||
|
||
---
|
||
|
||
## Headline Numbers (Last 7 Days)
|
||
|
||
| Metric | deepseek-v4-pro | claude-sonnet-5 |
|
||
|--------|----------------|-----------------|
|
||
| Call volume | 14,024 | 154 |
|
||
| Spend | $39.20 | $7.30 |
|
||
| Avg cost/call | $0.0028 | $0.0474 |
|
||
| Share of calls | 98.9% | 1.1% |
|
||
| Share of spend | 84.3% | 15.7% |
|
||
| Est. 30-day spend | ~$183 | ~$44 |
|
||
|
||
DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.
|
||
|
||
---
|
||
|
||
## Sonnet 5 Daily Breakdown
|
||
|
||
| Date | Calls | Spend | Context |
|
||
|------|-------|-------|---------|
|
||
| Aug 9 (today) | 3 | $0.05 | Early, still running |
|
||
| **Aug 8** | **73** | **$6.56** | Audit remediation — subagent cascading |
|
||
| Aug 7 | 11 | $0.23 | Normal dev day |
|
||
| Aug 6 | 6 | $0.01 | Model eval / testing |
|
||
| Aug 5 | 20 | $0.14 | |
|
||
| Aug 4 | 19 | $0.15 | |
|
||
| Aug 3 | 5 | $0.03 | Weekend |
|
||
| Aug 2 | 20 | $0.13 | |
|
||
| Aug 1 | 43 | $4.17 | Elevated — subagent routing |
|
||
| **Jul 31** | **146** | **$17.78** | Hit $20 daily cap — 89% of day's spend was Sonnet 5 |
|
||
| Jul 30 | 62 | $6.83 | |
|
||
| Jul 29 | 7 | $0.00 | |
|
||
| Jul 28 | 0 | $0.00 | |
|
||
| Jul 27 | 2 | $0.00 | |
|
||
| Jul 26 | 1 | $0.00 | |
|
||
| Jul 25 | 52 | $3.34 | |
|
||
| Jul 24 | 87 | $6.72 | |
|
||
| Jul 23 | 4 | $0.00 | |
|
||
| Jul 22 | 0 | $0.00 | |
|
||
| Jul 21 | 0 | $0.00 | |
|
||
| Jul 20 | 0 | $0.00 | |
|
||
| Jul 19 | 0 | $0.00 | |
|
||
| Jul 18 | 1 | $0.00 | |
|
||
| Jul 17 | 0 | $0.00 | |
|
||
| Jul 16 | 0 | $0.00 | |
|
||
| Jul 15 | 0 | $0.00 | |
|
||
| Jul 14 | 2 | $0.00 | |
|
||
| Jul 13 | 0 | $0.00 | |
|
||
| Jul 12 | 34 | $12.75 | Model eval pipeline |
|
||
| Jul 11 | 0 | $0.00 | |
|
||
| Jul 10 | 12 | $0.02 | |
|
||
|
||
**Typical normal day:** ~11 Sonnet 5 calls, ~$0.25/day
|
||
**Anomaly days:** Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month
|
||
|
||
---
|
||
|
||
## 30-Day Daily Spend Trend
|
||
|
||
```
|
||
Date Total Spend Sonnet 5 Total Calls Sonnet Calls
|
||
Aug 09 $1.42 $0.05 362 3
|
||
Aug 08 $16.00 $6.56 3,147 73
|
||
Aug 07 $6.44 $0.93 1,787 42
|
||
Aug 06 $3.06 $0.01 1,493 11
|
||
Aug 05 $7.96 $0.14 3,147 20
|
||
Aug 04 $5.14 $0.15 1,679 19
|
||
Aug 03 $2.91 $0.03 996 5
|
||
Aug 02 $5.72 $0.13 2,556 20
|
||
Aug 01 $10.29 $4.17 2,354 43
|
||
Jul 31 $20.01 $17.78 1,130 146 ⬅ cap hit
|
||
Jul 30 $10.29 $6.83 2,073 62
|
||
Jul 29 $3.60 $0.00 2,660 7
|
||
Jul 28 $3.41 $0.00 2,420 0
|
||
Jul 27 $1.38 $0.00 1,192 2
|
||
Jul 26 $1.02 $0.00 332 1
|
||
Jul 25 $6.29 $3.34 985 52
|
||
Jul 24 $59.71 $6.72 1,066 87
|
||
Jul 23 $37.92 $0.00 1,023 4
|
||
Jul 22 $67.03 $0.00 1,042 0
|
||
Jul 21 $4.77 $0.00 140 0
|
||
Jul 20 $4.20 $0.00 1,301 0
|
||
Jul 19 $0.81 $0.00 109 0
|
||
Jul 18 $0.17 $0.00 90 1
|
||
Jul 17 $2.17 $0.00 168 0
|
||
Jul 16 $1.42 $0.00 396 0
|
||
Jul 15 $5.83 $0.00 1,545 0
|
||
Jul 14 $2.34 $0.00 1,305 2
|
||
Jul 13 $57.49 $0.00 3,061 0
|
||
Jul 12 $99.86 $12.75 3,920 34 ⬅ biggest spike
|
||
Jul 11 $0.18 $0.00 1,656 0
|
||
Jul 10 $21.89 $0.02 4,544 12
|
||
```
|
||
|
||
**August normal days:** $3–8/day typical, $10–16/day on heavy remediation days
|
||
|
||
---
|
||
|
||
## Prompt Caching Status
|
||
|
||
```
|
||
cache_hit = 0 across ALL models, ALL calls, ALL 30 days
|
||
```
|
||
|
||
Prompt caching is **not enabled**. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.
|
||
|
||
### Caching Economics
|
||
|
||
Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):
|
||
|
||
| Scenario | Input $/M tokens |
|
||
|----------|-----------------|
|
||
| No caching (current) | $2.00 |
|
||
| Cache write (5 min TTL) | $2.50 |
|
||
| Cache write (1 hr TTL) | $4.00 |
|
||
| Cache hit | **$0.20** (90% off) |
|
||
|
||
After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.
|
||
|
||
**Projected savings for Hermes workload** (large system prompts, repeated across turns):
|
||
|
||
| Cache hit rate | Input cost reduction | Monthly savings |
|
||
|---------------|---------------------|-----------------|
|
||
| 70% | –53% | ~$15–25 |
|
||
| 90% | –81% | ~$20–35 |
|
||
|
||
---
|
||
|
||
## Key Spend Anomalies — Root Cause Analysis
|
||
|
||
| Date | Spend | Root Cause |
|
||
|------|-------|-----------|
|
||
| **Jul 12** | $99.86 | **gpt-5.5 eval pipeline.** 196 calls to gpt-5.5 ($66.42 — 67% of day) with 219K avg prompt tokens. 94 calls alone at 23:00 ($43.55 in one hour). `deepseek-v4-flash` added 1,846 eval calls ($2.87). Model catalog audit against all 128 models. Not Hermes. |
|
||
| **Jul 13** | $57.49 | **gpt-5.5 eval pipeline (continuation).** 30 calls for $56.23 (98% of day) with 458K avg prompt tokens. Three overnight bursts: midnight ($15.36), 3 AM ($29.00), 4 AM ($11.87). DeepSeek V4 Pro handled all other traffic ($1.21). |
|
||
| **Jul 22–24** | $37–67/day | **gpt-5.5 → gpt-5.6-terra eval pipeline.** Fewer calls (1,023–1,066) but 10–20× normal cost per call. gpt-5.5 at $63.93 (Jul 22), gpt-5.6-terra at $31.64 (Jul 23) and $50.22 (Jul 24). Avg prompt size: 309K–397K tokens. Each eval call cost $0.30–$1.50 vs normal $0.003. |
|
||
| **Jul 31** | $20.01 | **$20 daily cap breached.** 146 Sonnet 5 calls ($17.78 — 89% of spend). Subagent `delegation.model` was pinned to `claude-sonnet-5`, bypassing the conductor's model routing. Fixed Aug 1 by switching delegation back to `deepseek-v4-pro`. |
|
||
|
||
All four anomalies share a common root: **the July model evaluation pipeline** hitting gpt-5.5 and gpt-5.6-terra through admin-ai with enormous evaluation-sized contexts. These models were never in Hermes' production chain — the eval runner discovered them in the proxy catalog and tested them. The Jul 31 event was a separate bug: subagent delegation config hard-overriding to Sonnet 5.
|
||
|
||
---
|
||
|
||
## Data Source Limitation
|
||
|
||
> **This report only covers LiteLLM-proxied traffic (admin-ai). It is blind to direct fallback provider spend.**
|
||
|
||
The fallback chain operates outside admin-ai: `deepseek direct → google direct → xai direct → anthropic direct`. When admin-ai is unreachable or the daily cap is hit, traffic falls through to these keys. Spend there is invisible to LiteLLM SpendLogs.
|
||
|
||
**Known gap:** Aug 5 actual spend was ~$45 (per changelog) but LiteLLM shows only $7.96. The ~$37 delta went through direct provider keys.
|
||
|
||
**Fix needed:** Real-time cost monitoring requires a second feed polling each provider's usage API directly. Without it, a fallback cascade can silently burn through provider credits with no alert.
|
||
|
||
---
|
||
|
||
## Verdict
|
||
|
||
> **The model chain is correct and working as designed. Cost is under control for the proxied path. The fallback path is a blind spot that needs monitoring.**
|
||
|
||
- DeepSeek V4 Pro: 99% of calls, ~$5.60/day — the workhorse
|
||
- Sonnet 5: 1.1% of calls (~11/day typical), genuine rare override — not a silent runaway
|
||
- July's $271 in anomaly spend (Jul 12–24) was the model evaluation pipeline hitting non-production models — not Hermes
|
||
- August baseline: $3–8/day typical, $10–16/day on heavy remediation days
|
||
- The model chain doc matches reality: `deepseek-v4-pro` primary, `claude-sonnet-5` for critical escalation
|
||
|
||
**Three action items:**
|
||
|
||
1. **Enable Anthropic prompt caching** — 80–90% off cached input tokens. Client-side change only. Must be done before Sep 1 ($2→$3 base price increase).
|
||
2. **Implement fallback provider monitoring** — direct API polling of DeepSeek, Google, xAI, and Anthropic usage endpoints. The LiteLLM SpendLogs are blind to ~30–50% of actual spend on failover days.
|
||
3. **Tag eval pipeline traffic** — any automated model testing must use a dedicated LiteLLM key with its own budget cap. The July anomalies contaminated 30 days of production cost data.
|