Files
itpp-infrastructure/docs/reports/model-usage-2026-08-09.md
T

155 lines
5.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Hermes Model Usage Report
**2026-08-09** | 30-Day Window (Jul 10 Aug 9, 2026)
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)
---
## Headline Numbers (Last 7 Days)
| Metric | deepseek-v4-pro | claude-sonnet-5 |
|--------|----------------|-----------------|
| Call volume | 14,024 | 154 |
| Spend | $39.20 | $7.30 |
| Avg cost/call | $0.0028 | $0.0474 |
| Share of calls | 98.9% | 1.1% |
| Share of spend | 84.3% | 15.7% |
| Est. 30-day spend | ~$183 | ~$44 |
DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.
---
## Sonnet 5 Daily Breakdown
| Date | Calls | Spend | Context |
|------|-------|-------|---------|
| Aug 9 (today) | 3 | $0.05 | Early, still running |
| **Aug 8** | **73** | **$6.56** | Audit remediation — subagent cascading |
| Aug 7 | 11 | $0.23 | Normal dev day |
| Aug 6 | 6 | $0.01 | Model eval / testing |
| Aug 5 | 20 | $0.14 | |
| Aug 4 | 19 | $0.15 | |
| Aug 3 | 5 | $0.03 | Weekend |
| Aug 2 | 20 | $0.13 | |
| Aug 1 | 43 | $4.17 | Elevated — subagent routing |
| **Jul 31** | **146** | **$17.78** | Hit $20 daily cap — 89% of day's spend was Sonnet 5 |
| Jul 30 | 62 | $6.83 | |
| Jul 29 | 7 | $0.00 | |
| Jul 28 | 0 | $0.00 | |
| Jul 27 | 2 | $0.00 | |
| Jul 26 | 1 | $0.00 | |
| Jul 25 | 52 | $3.34 | |
| Jul 24 | 87 | $6.72 | |
| Jul 23 | 4 | $0.00 | |
| Jul 22 | 0 | $0.00 | |
| Jul 21 | 0 | $0.00 | |
| Jul 20 | 0 | $0.00 | |
| Jul 19 | 0 | $0.00 | |
| Jul 18 | 1 | $0.00 | |
| Jul 17 | 0 | $0.00 | |
| Jul 16 | 0 | $0.00 | |
| Jul 15 | 0 | $0.00 | |
| Jul 14 | 2 | $0.00 | |
| Jul 13 | 0 | $0.00 | |
| Jul 12 | 34 | $12.75 | Evaluation run |
| Jul 11 | 0 | $0.00 | |
| Jul 10 | 12 | $0.02 | |
**Typical normal day:** ~11 Sonnet 5 calls, ~$0.25/day
**Anomaly days:** Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month
---
## 30-Day Daily Spend Trend
```
Date Total Spend Sonnet 5 Total Calls Sonnet Calls
Aug 09 $1.42 $0.05 362 3
Aug 08 $16.00 $6.56 3,147 73
Aug 07 $6.44 $0.93 1,787 42 ⬅ all Claude models
Aug 06 $3.06 $0.01 1,493 11
Aug 05 $7.96 $0.14 3,147 20
Aug 04 $5.14 $0.15 1,679 19
Aug 03 $2.91 $0.03 996 5
Aug 02 $5.72 $0.13 2,556 20
Aug 01 $10.29 $4.17 2,354 43
Jul 31 $20.01 $17.78 1,130 146 ⬅ cap hit
Jul 30 $10.29 $6.83 2,073 62
Jul 29 $3.60 $0.00 2,660 7
Jul 28 $3.41 $0.00 2,420 0
Jul 27 $1.38 $0.00 1,192 2
Jul 26 $1.02 $0.00 332 1
Jul 25 $6.29 $3.34 985 52
Jul 24 $59.71 $6.72 1,066 87
Jul 23 $37.92 $0.00 1,023 4
Jul 22 $67.03 $0.00 1,042 0
Jul 21 $4.77 $0.00 140 0
Jul 20 $4.20 $0.00 1,301 0
Jul 19 $0.81 $0.00 109 0
Jul 18 $0.17 $0.00 90 1
Jul 17 $2.17 $0.00 168 0
Jul 16 $1.42 $0.00 396 0
Jul 15 $5.83 $0.00 1,545 0
Jul 14 $2.34 $0.00 1,305 2
Jul 13 $57.49 $0.00 3,061 0
Jul 12 $99.86 $12.75 3,920 34 ⬅ biggest spike
Jul 11 $0.18 $0.00 1,656 0
Jul 10 $21.89 $0.02 4,544 12
```
**Normal daily spend range (August):** $38/day
---
## Prompt Caching Status
```
cache_hit = 0 across ALL models, ALL calls, ALL 30 days
```
Prompt caching is **not enabled**. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.
### Caching Economics
Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):
| Scenario | Input $/M tokens |
|----------|-----------------|
| No caching (current) | $2.00 |
| Cache write (5 min TTL) | $2.50 |
| Cache write (1 hr TTL) | $4.00 |
| Cache hit | **$0.20** (90% off) |
After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.
**Projected savings for Hermes workload** (large system prompts, repeated across turns):
| Cache hit rate | Input cost reduction | Monthly savings |
|---------------|---------------------|-----------------|
| 70% | 53% | ~$1525 |
| 90% | 81% | ~$2035 |
---
## Key Spend Anomalies
| Date | Spend | Root Cause |
|------|-------|-----------|
| Jul 12 | $99.86 | Massive outlier — evaluation run or unbounded subagent loop |
| Jul 13 | $57.49 | Continuation |
| Jul 2224 | $3767/day | Heavy deepseek-v4-pro throughput (23× normal) |
| Jul 31 | $20.01 | $20 daily cap breached — 89% from Sonnet 5 (subagent routing override) |
---
## Verdict
> **The model chain is correct and working as designed. Not under-resourced.**
- DeepSeek V4 Pro: 99% of calls, $5.60/day average — the workhorse
- Sonnet 5: 1.1% of calls, genuine rare override — not a silent runaway
- Monthly spend has stabilized at $38/day in August, well within the $2030/day budget
- The model chain doc now matches reality: `deepseek-v4-pro` primary, `claude-sonnet-5` for critical escalation
**One concrete action:** Enable Anthropic prompt caching. It's a client-side change with no infrastructure cost — just set cache breakpoints on the system prompt. 8090% off cached input tokens. Worth doing before Sep 1 when base prices rise 50%.