Files
itpp-infrastructure/docs/reports/model-usage-2026-08-09.md
T

8.2 KiB
Raw Blame History

Hermes Model Usage Report

2026-08-09 | 30-Day Window (Jul 10 Aug 9, 2026)
Source: LiteLLM SpendLogs (93,786 requests, PostgreSQL on app1)


Headline Numbers (Last 7 Days)

Metric deepseek-v4-pro claude-sonnet-5
Call volume 14,024 154
Spend $39.20 $7.30
Avg cost/call $0.0028 $0.0474
Share of calls 98.9% 1.1%
Share of spend 84.3% 15.7%
Est. 30-day spend ~$183 ~$44

DeepSeek V4 Pro is 17× cheaper per call and handles 99% of volume.


Sonnet 5 Daily Breakdown

Date Calls Spend Context
Aug 9 (today) 3 $0.05 Early, still running
Aug 8 73 $6.56 Audit remediation — subagent cascading
Aug 7 11 $0.23 Normal dev day
Aug 6 6 $0.01 Model eval / testing
Aug 5 20 $0.14
Aug 4 19 $0.15
Aug 3 5 $0.03 Weekend
Aug 2 20 $0.13
Aug 1 43 $4.17 Elevated — subagent routing
Jul 31 146 $17.78 Hit $20 daily cap — 89% of day's spend was Sonnet 5
Jul 30 62 $6.83
Jul 29 7 $0.00
Jul 28 0 $0.00
Jul 27 2 $0.00
Jul 26 1 $0.00
Jul 25 52 $3.34
Jul 24 87 $6.72
Jul 23 4 $0.00
Jul 22 0 $0.00
Jul 21 0 $0.00
Jul 20 0 $0.00
Jul 19 0 $0.00
Jul 18 1 $0.00
Jul 17 0 $0.00
Jul 16 0 $0.00
Jul 15 0 $0.00
Jul 14 2 $0.00
Jul 13 0 $0.00
Jul 12 34 $12.75 Model eval pipeline
Jul 11 0 $0.00
Jul 10 12 $0.02

Typical normal day: ~11 Sonnet 5 calls, ~$0.25/day
Anomaly days: Jul 31 ($17.78), Aug 1 ($4.17), Aug 8 ($6.56) account for 63% of all Sonnet 5 spend this month


30-Day Daily Spend Trend

Date       Total Spend   Sonnet 5    Total Calls   Sonnet Calls
Aug 09        $1.42         $0.05         362             3
Aug 08       $16.00         $6.56       3,147            73
Aug 07        $6.44         $0.93       1,787            42
Aug 06        $3.06         $0.01       1,493            11
Aug 05        $7.96         $0.14       3,147            20
Aug 04        $5.14         $0.15       1,679            19
Aug 03        $2.91         $0.03         996             5
Aug 02        $5.72         $0.13       2,556            20
Aug 01       $10.29         $4.17       2,354            43
Jul 31       $20.01        $17.78       1,130           146 ⬅ cap hit
Jul 30       $10.29         $6.83       2,073            62
Jul 29        $3.60         $0.00       2,660             7
Jul 28        $3.41         $0.00       2,420             0
Jul 27        $1.38         $0.00       1,192             2
Jul 26        $1.02         $0.00         332             1
Jul 25        $6.29         $3.34         985            52
Jul 24       $59.71         $6.72       1,066            87
Jul 23       $37.92         $0.00       1,023             4
Jul 22       $67.03         $0.00       1,042             0
Jul 21        $4.77         $0.00         140             0
Jul 20        $4.20         $0.00       1,301             0
Jul 19        $0.81         $0.00         109             0
Jul 18        $0.17         $0.00          90             1
Jul 17        $2.17         $0.00         168             0
Jul 16        $1.42         $0.00         396             0
Jul 15        $5.83         $0.00       1,545             0
Jul 14        $2.34         $0.00       1,305             2
Jul 13       $57.49         $0.00       3,061             0
Jul 12       $99.86        $12.75       3,920            34 ⬅ biggest spike
Jul 11        $0.18         $0.00       1,656             0
Jul 10       $21.89         $0.02       4,544            12

August normal days: $38/day typical, $1016/day on heavy remediation days


Prompt Caching Status

cache_hit = 0 across ALL models, ALL calls, ALL 30 days

Prompt caching is not enabled. Hermes does not send Anthropic cache control headers. The LiteLLM proxy passes them through natively — enabling requires a client-side change only.

Caching Economics

Anthropic Claude Sonnet 5 introductory pricing (through Aug 31, 2026):

Scenario Input $/M tokens
No caching (current) $2.00
Cache write (5 min TTL) $2.50
Cache write (1 hr TTL) $4.00
Cache hit $0.20 (90% off)

After Sep 1, 2026: base input rises to $3/M, cache hits to $0.30/M.

Projected savings for Hermes workload (large system prompts, repeated across turns):

Cache hit rate Input cost reduction Monthly savings
70% 53% ~$1525
90% 81% ~$2035

Key Spend Anomalies — Root Cause Analysis

Date Spend Root Cause
Jul 12 $99.86 gpt-5.5 eval pipeline. 196 calls to gpt-5.5 ($66.42 — 67% of day) with 219K avg prompt tokens. 94 calls alone at 23:00 ($43.55 in one hour). deepseek-v4-flash added 1,846 eval calls ($2.87). Model catalog audit against all 128 models. Not Hermes.
Jul 13 $57.49 gpt-5.5 eval pipeline (continuation). 30 calls for $56.23 (98% of day) with 458K avg prompt tokens. Three overnight bursts: midnight ($15.36), 3 AM ($29.00), 4 AM ($11.87). DeepSeek V4 Pro handled all other traffic ($1.21).
Jul 2224 $3767/day gpt-5.5 → gpt-5.6-terra eval pipeline. Fewer calls (1,0231,066) but 1020× normal cost per call. gpt-5.5 at $63.93 (Jul 22), gpt-5.6-terra at $31.64 (Jul 23) and $50.22 (Jul 24). Avg prompt size: 309K397K tokens. Each eval call cost $0.30$1.50 vs normal $0.003.
Jul 31 $20.01 $20 daily cap breached. 146 Sonnet 5 calls ($17.78 — 89% of spend). Subagent delegation.model was pinned to claude-sonnet-5, bypassing the conductor's model routing. Fixed Aug 1 by switching delegation back to deepseek-v4-pro.

All four anomalies share a common root: the July model evaluation pipeline hitting gpt-5.5 and gpt-5.6-terra through admin-ai with enormous evaluation-sized contexts. These models were never in Hermes' production chain — the eval runner discovered them in the proxy catalog and tested them. The Jul 31 event was a separate bug: subagent delegation config hard-overriding to Sonnet 5.


Data Source Limitation

This report only covers LiteLLM-proxied traffic (admin-ai). It is blind to direct fallback provider spend.

The fallback chain operates outside admin-ai: deepseek direct → google direct → xai direct → anthropic direct. When admin-ai is unreachable or the daily cap is hit, traffic falls through to these keys. Spend there is invisible to LiteLLM SpendLogs.

Known gap: Aug 5 actual spend was ~$45 (per changelog) but LiteLLM shows only $7.96. The ~$37 delta went through direct provider keys.

Fix needed: Real-time cost monitoring requires a second feed polling each provider's usage API directly. Without it, a fallback cascade can silently burn through provider credits with no alert.


Verdict

The model chain is correct and working as designed. Cost is under control for the proxied path. The fallback path is a blind spot that needs monitoring.

  • DeepSeek V4 Pro: 99% of calls, ~$5.60/day — the workhorse
  • Sonnet 5: 1.1% of calls (~11/day typical), genuine rare override — not a silent runaway
  • July's $271 in anomaly spend (Jul 1224) was the model evaluation pipeline hitting non-production models — not Hermes
  • August baseline: $38/day typical, $1016/day on heavy remediation days
  • The model chain doc matches reality: deepseek-v4-pro primary, claude-sonnet-5 for critical escalation

Three action items:

  1. Enable Anthropic prompt caching — 8090% off cached input tokens. Client-side change only. Must be done before Sep 1 ($2→$3 base price increase).
  2. Implement fallback provider monitoring — direct API polling of DeepSeek, Google, xAI, and Anthropic usage endpoints. The LiteLLM SpendLogs are blind to ~3050% of actual spend on failover days.
  3. Tag eval pipeline traffic — any automated model testing must use a dedicated LiteLLM key with its own budget cap. The July anomalies contaminated 30 days of production cost data.