Files
itpp-infrastructure/docs/reports/model-usage-2026-08-09.md
T

174 lines
8.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Hermes Model Usage Report
**2026-08-09** | 30-Day Window (Jul 10 Aug 9, 2026)
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)
---
## Headline Numbers (Last 7 Days)
| Metric | deepseek-v4-pro | claude-sonnet-5 |
|--------|----------------|-----------------|
| Call volume | 14,024 | 154 |
| Spend | $39.20 | $7.30 |
| Avg cost/call | $0.0028 | $0.0474 |
| Share of calls | 98.9% | 1.1% |
| Share of spend | 84.3% | 15.7% |
| Est. 30-day spend | ~$183 | ~$44 |
DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.
---
## Sonnet 5 Daily Breakdown
| Date | Calls | Spend | Context |
|------|-------|-------|---------|
| Aug 9 (today) | 3 | $0.05 | Early, still running |
| **Aug 8** | **73** | **$6.56** | Audit remediation — subagent cascading |
| Aug 7 | 11 | $0.23 | Normal dev day |
| Aug 6 | 6 | $0.01 | Model eval / testing |
| Aug 5 | 20 | $0.14 | |
| Aug 4 | 19 | $0.15 | |
| Aug 3 | 5 | $0.03 | Weekend |
| Aug 2 | 20 | $0.13 | |
| Aug 1 | 43 | $4.17 | Elevated — subagent routing |
| **Jul 31** | **146** | **$17.78** | Hit $20 daily cap — 89% of day's spend was Sonnet 5 |
| Jul 30 | 62 | $6.83 | |
| Jul 29 | 7 | $0.00 | |
| Jul 28 | 0 | $0.00 | |
| Jul 27 | 2 | $0.00 | |
| Jul 26 | 1 | $0.00 | |
| Jul 25 | 52 | $3.34 | |
| Jul 24 | 87 | $6.72 | |
| Jul 23 | 4 | $0.00 | |
| Jul 22 | 0 | $0.00 | |
| Jul 21 | 0 | $0.00 | |
| Jul 20 | 0 | $0.00 | |
| Jul 19 | 0 | $0.00 | |
| Jul 18 | 1 | $0.00 | |
| Jul 17 | 0 | $0.00 | |
| Jul 16 | 0 | $0.00 | |
| Jul 15 | 0 | $0.00 | |
| Jul 14 | 2 | $0.00 | |
| Jul 13 | 0 | $0.00 | |
| Jul 12 | 34 | $12.75 | Model eval pipeline |
| Jul 11 | 0 | $0.00 | |
| Jul 10 | 12 | $0.02 | |
**Typical normal day:** ~11 Sonnet 5 calls, ~$0.25/day
**Anomaly days:** Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month
---
## 30-Day Daily Spend Trend
```
Date Total Spend Sonnet 5 Total Calls Sonnet Calls
Aug 09 $1.42 $0.05 362 3
Aug 08 $16.00 $6.56 3,147 73
Aug 07 $6.44 $0.93 1,787 42
Aug 06 $3.06 $0.01 1,493 11
Aug 05 $7.96 $0.14 3,147 20
Aug 04 $5.14 $0.15 1,679 19
Aug 03 $2.91 $0.03 996 5
Aug 02 $5.72 $0.13 2,556 20
Aug 01 $10.29 $4.17 2,354 43
Jul 31 $20.01 $17.78 1,130 146 ⬅ cap hit
Jul 30 $10.29 $6.83 2,073 62
Jul 29 $3.60 $0.00 2,660 7
Jul 28 $3.41 $0.00 2,420 0
Jul 27 $1.38 $0.00 1,192 2
Jul 26 $1.02 $0.00 332 1
Jul 25 $6.29 $3.34 985 52
Jul 24 $59.71 $6.72 1,066 87
Jul 23 $37.92 $0.00 1,023 4
Jul 22 $67.03 $0.00 1,042 0
Jul 21 $4.77 $0.00 140 0
Jul 20 $4.20 $0.00 1,301 0
Jul 19 $0.81 $0.00 109 0
Jul 18 $0.17 $0.00 90 1
Jul 17 $2.17 $0.00 168 0
Jul 16 $1.42 $0.00 396 0
Jul 15 $5.83 $0.00 1,545 0
Jul 14 $2.34 $0.00 1,305 2
Jul 13 $57.49 $0.00 3,061 0
Jul 12 $99.86 $12.75 3,920 34 ⬅ biggest spike
Jul 11 $0.18 $0.00 1,656 0
Jul 10 $21.89 $0.02 4,544 12
```
**August normal days:** $38/day typical, $1016/day on heavy remediation days
---
## Prompt Caching Status
```
cache_hit = 0 across ALL models, ALL calls, ALL 30 days
```
Prompt caching is **not enabled**. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.
### Caching Economics
Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):
| Scenario | Input $/M tokens |
|----------|-----------------|
| No caching (current) | $2.00 |
| Cache write (5 min TTL) | $2.50 |
| Cache write (1 hr TTL) | $4.00 |
| Cache hit | **$0.20** (90% off) |
After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.
**Projected savings for Hermes workload** (large system prompts, repeated across turns):
| Cache hit rate | Input cost reduction | Monthly savings |
|---------------|---------------------|-----------------|
| 70% | 53% | ~$1525 |
| 90% | 81% | ~$2035 |
---
## Key Spend Anomalies — Root Cause Analysis
| Date | Spend | Root Cause |
|------|-------|-----------|
| **Jul 12** | $99.86 | **gpt-5.5 eval pipeline.** 196 calls to gpt-5.5 ($66.42 — 67% of day) with 219K avg prompt tokens. 94 calls alone at 23:00 ($43.55 in one hour). `deepseek-v4-flash` added 1,846 eval calls ($2.87). Model catalog audit against all 128 models. Not Hermes. |
| **Jul 13** | $57.49 | **gpt-5.5 eval pipeline (continuation).** 30 calls for $56.23 (98% of day) with 458K avg prompt tokens. Three overnight bursts: midnight ($15.36), 3 AM ($29.00), 4 AM ($11.87). DeepSeek V4 Pro handled all other traffic ($1.21). |
| **Jul 2224** | $3767/day | **gpt-5.5 → gpt-5.6-terra eval pipeline.** Fewer calls (1,0231,066) but 1020× normal cost per call. gpt-5.5 at $63.93 (Jul 22), gpt-5.6-terra at $31.64 (Jul 23) and $50.22 (Jul 24). Avg prompt size: 309K397K tokens. Each eval call cost $0.30$1.50 vs normal $0.003. |
| **Jul 31** | $20.01 | **$20 daily cap breached.** 146 Sonnet 5 calls ($17.78 — 89% of spend). Subagent `delegation.model` was pinned to `claude-sonnet-5`, bypassing the conductor's model routing. Fixed Aug 1 by switching delegation back to `deepseek-v4-pro`. |
All four anomalies share a common root: **the July model evaluation pipeline** hitting gpt-5.5 and gpt-5.6-terra through admin-ai with enormous evaluation-sized contexts. These models were never in Hermes' production chain — the eval runner discovered them in the proxy catalog and tested them. The Jul 31 event was a separate bug: subagent delegation config hard-overriding to Sonnet 5.
---
## Data Source Limitation
> **This report only covers LiteLLM-proxied traffic (admin-ai). It is blind to direct fallback provider spend.**
The fallback chain operates outside admin-ai: `deepseek direct → google direct → xai direct → anthropic direct`. When admin-ai is unreachable or the daily cap is hit, traffic falls through to these keys. Spend there is invisible to LiteLLM SpendLogs.
**Known gap:** Aug 5 actual spend was ~$45 (per changelog) but LiteLLM shows only $7.96. The ~$37 delta went through direct provider keys.
**Fix needed:** Real-time cost monitoring requires a second feed polling each provider's usage API directly. Without it, a fallback cascade can silently burn through provider credits with no alert.
---
## Verdict
> **The model chain is correct and working as designed. Cost is under control for the proxied path. The fallback path is a blind spot that needs monitoring.**
- DeepSeek V4 Pro: 99% of calls, ~$5.60/day — the workhorse
- Sonnet 5: 1.1% of calls (~11/day typical), genuine rare override — not a silent runaway
- July's $271 in anomaly spend (Jul 1224) was the model evaluation pipeline hitting non-production models — not Hermes
- August baseline: $38/day typical, $1016/day on heavy remediation days
- The model chain doc matches reality: `deepseek-v4-pro` primary, `claude-sonnet-5` for critical escalation
**Three action items:**
1. **Enable Anthropic prompt caching** — 8090% off cached input tokens. Client-side change only. Must be done before Sep 1 ($2→$3 base price increase).
2. **Implement fallback provider monitoring** — direct API polling of DeepSeek, Google, xAI, and Anthropic usage endpoints. The LiteLLM SpendLogs are blind to ~3050% of actual spend on failover days.
3. **Tag eval pipeline traffic** — any automated model testing must use a dedicated LiteLLM key with its own budget cap. The July anomalies contaminated 30 days of production cost data.