Files
itpp-infrastructure/docs/reports/model-usage-2026-08-09.md
T

5.8 KiB
Raw Blame History

Hermes Model Usage Report

2026-08-09 | 30-Day Window (Jul 10 Aug 9, 2026)
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)


Headline Numbers (Last 7 Days)

Metric deepseek-v4-pro claude-sonnet-5
Call volume 14,024 154
Spend $39.20 $7.30
Avg cost/call $0.0028 $0.0474
Share of calls 98.9% 1.1%
Share of spend 84.3% 15.7%
Est. 30-day spend ~$183 ~$44

DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.


Sonnet 5 Daily Breakdown

Date Calls Spend Context
Aug 9 (today) 3 $0.05 Early, still running
Aug 8 73 $6.56 Audit remediation — subagent cascading
Aug 7 11 $0.23 Normal dev day
Aug 6 6 $0.01 Model eval / testing
Aug 5 20 $0.14
Aug 4 19 $0.15
Aug 3 5 $0.03 Weekend
Aug 2 20 $0.13
Aug 1 43 $4.17 Elevated — subagent routing
Jul 31 146 $17.78 Hit $20 daily cap — 89% of day's spend was Sonnet 5
Jul 30 62 $6.83
Jul 29 7 $0.00
Jul 28 0 $0.00
Jul 27 2 $0.00
Jul 26 1 $0.00
Jul 25 52 $3.34
Jul 24 87 $6.72
Jul 23 4 $0.00
Jul 22 0 $0.00
Jul 21 0 $0.00
Jul 20 0 $0.00
Jul 19 0 $0.00
Jul 18 1 $0.00
Jul 17 0 $0.00
Jul 16 0 $0.00
Jul 15 0 $0.00
Jul 14 2 $0.00
Jul 13 0 $0.00
Jul 12 34 $12.75 Evaluation run
Jul 11 0 $0.00
Jul 10 12 $0.02

Typical normal day: ~11 Sonnet 5 calls, ~$0.25/day
Anomaly days: Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month


30-Day Daily Spend Trend

Date       Total Spend   Sonnet 5    Total Calls   Sonnet Calls
Aug 09        $1.42         $0.05         362             3
Aug 08       $16.00         $6.56       3,147            73
Aug 07        $6.44         $0.93       1,787            42 ⬅ all Claude models
Aug 06        $3.06         $0.01       1,493            11
Aug 05        $7.96         $0.14       3,147            20
Aug 04        $5.14         $0.15       1,679            19
Aug 03        $2.91         $0.03         996             5
Aug 02        $5.72         $0.13       2,556            20
Aug 01       $10.29         $4.17       2,354            43
Jul 31       $20.01        $17.78       1,130           146 ⬅ cap hit
Jul 30       $10.29         $6.83       2,073            62
Jul 29        $3.60         $0.00       2,660             7
Jul 28        $3.41         $0.00       2,420             0
Jul 27        $1.38         $0.00       1,192             2
Jul 26        $1.02         $0.00         332             1
Jul 25        $6.29         $3.34         985            52
Jul 24       $59.71         $6.72       1,066            87
Jul 23       $37.92         $0.00       1,023             4
Jul 22       $67.03         $0.00       1,042             0
Jul 21        $4.77         $0.00         140             0
Jul 20        $4.20         $0.00       1,301             0
Jul 19        $0.81         $0.00         109             0
Jul 18        $0.17         $0.00          90             1
Jul 17        $2.17         $0.00         168             0
Jul 16        $1.42         $0.00         396             0
Jul 15        $5.83         $0.00       1,545             0
Jul 14        $2.34         $0.00       1,305             2
Jul 13       $57.49         $0.00       3,061             0
Jul 12       $99.86        $12.75       3,920            34 ⬅ biggest spike
Jul 11        $0.18         $0.00       1,656             0
Jul 10       $21.89         $0.02       4,544            12

Normal daily spend range (August): $38/day


Prompt Caching Status

cache_hit = 0 across ALL models, ALL calls, ALL 30 days

Prompt caching is not enabled. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.

Caching Economics

Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):

Scenario Input $/M tokens
No caching (current) $2.00
Cache write (5 min TTL) $2.50
Cache write (1 hr TTL) $4.00
Cache hit $0.20 (90% off)

After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.

Projected savings for Hermes workload (large system prompts, repeated across turns):

Cache hit rate Input cost reduction Monthly savings
70% 53% ~$1525
90% 81% ~$2035

Key Spend Anomalies

Date Spend Root Cause
Jul 12 $99.86 Massive outlier — evaluation run or unbounded subagent loop
Jul 13 $57.49 Continuation
Jul 2224 $3767/day Heavy deepseek-v4-pro throughput (23× normal)
Jul 31 $20.01 $20 daily cap breached — 89% from Sonnet 5 (subagent routing override)

Verdict

The model chain is correct and working as designed. Not under-resourced.

  • DeepSeek V4 Pro: 99% of calls, $5.60/day average — the workhorse
  • Sonnet 5: 1.1% of calls, genuine rare override — not a silent runaway
  • Monthly spend has stabilized at $38/day in August, well within the $2030/day budget
  • The model chain doc now matches reality: deepseek-v4-pro primary, claude-sonnet-5 for critical escalation

One concrete action: Enable Anthropic prompt caching. It's a client-side change with no infrastructure cost — just set cache breakpoints on the system prompt. 8090% off cached input tokens. Worth doing before Sep 1 when base prices rise 50%.