v4.1 go-live cut, Moonshot to Mistral swap, v5.x superseded

Mark v5.0/v5.1 SUPERSEDED (error-detection thesis failed at -0.40 delta); v4.1 is canonical. Cut v4.1 proposal with 12 version strings bumped. Moonshot to Mistral across production seats; production worker de-kimi'd 2026-08-18. Data-retention posture corrected 7/9 to 8/9 no-training (DeepSeek sole exception). Reconciled COGS with measured Mistral spend. Committed deployed v4.0 content and research docs to resolve the repo/live fork.
This commit is contained in:
root
2026-08-18 20:28:36 -04:00
parent c1e690a122
commit 3c07727f5c
18 changed files with 3409 additions and 2 deletions
+88
View File
@@ -0,0 +1,88 @@
# VerdictTank v4.0 — Dogfood Verdict Report
**Date:** 2026-08-18
**Subject:** The v4.0 proposal itself, run through the v4.0 single-pass panel (10 seats: 9 scoring + synthesis gate)
**Verdict:** **CONDITIONAL**
---
## 1. Headline Scores
| Metric | Score |
|---|---|
| Proposal Strength | **71** |
| Investor Readiness | **44** |
| Composite | **58** |
| Divergence | 27 (proposal ahead) |
| Blocking findings | **Yes (6)** |
## 2. Dimension Means
| Dimension | Score |
|---|---|
| Structure & Communication | 93.8 |
| Internal Consistency | 86.2 |
| Solution Coherence | 84.9 |
| Problem Clarity | 81.2 |
| Risk & Compliance | 64.3 |
| Financial Integrity | 61.8 |
| Execution Feasibility | 50.0 |
| Team & Capability | 31.7 |
| Evidence Quality | 24.0 |
| Market Reality | 15.7 |
## 3. Panel Read
The panel's one-line take: **a well-designed critique loop that reads as an internal blueprint, not a funded product.** It scores 80+ on the four dimensions that measure writing quality, and collapses on the three that measure proof: market reality (15.7), evidence (24.0), team (31.7). Every margin claim, pricing scenario, and shipped-date assertion is unverified.
## 4. Six Blocking Conditions
1. **Market evidence** — bottom-up SAM, three named competitors with ARR/pricing, current paid account count by tier + trailing 30-day volume.
2. **Production proof** — anonymized logs: 500+ reviews processed, median latency <90s, uptime >99.5% over 30 days.
3. **Verified COGS** — vendor invoices/token logs from 100+ real reviews, reconciled to the $0.36 standard / $3.00 White-Label ceiling.
4. **Zero data-retention agreements** — executed with all nine vendors, or drop those vendors from the roster.
5. **Senior technical hire** — name a platform engineer/cofounder with multi-tenant SaaS track record.
6. **Arithmetic fix** — ceiling volume is 2,690 paid reviews, not 3,940 (a 1,250 gap), and the annual MRR reduction recomputes to $2,619, not $2,946.
---
## 5. Panel Integrity Findings (from running the panel itself)
### 5.1 glm-5.2 was not flaky — it was token-starved
**Root cause:** glm-5.2 is a reasoning model. On the financial seat (the heaviest brief), it consumed its entire token budget on hidden reasoning and returned zero visible tokens.
| Metric | Before (12k floor) | After (24k floor) |
|---|---|---|
| finish_reason | `length` | `stop` |
| reasoning_tokens | 11,998 | 14,635 |
| visible content | 0-3,015 chars | 8,545 chars |
| JSON parsed | no | yes (all 10 keys) |
**Conclusion:** glm-5.2 stays in the roster. The financial seat needs ~17k completion tokens (14.6k reasoning + 2k output). This was a budget bug, not a model defect.
### 5.2 Two production blockers found and fixed
1. **VerdictTank-Key allowlist was stale** (v3.5-era, 11 models). It 403'd on six v4.0 roster models: `sonar`, `gpt-5.2`, `glm-5.2`, `claude-haiku-4-5`, `xai/grok-4.3`, and bare `gemini-pro-latest` (it only knew `gemini/gemini-pro-latest`).
2. **Reasoning-model token floors too low** — deployed worker had glm-5.2 and deepseek-v4-pro at 8,000, which silently truncated reasoning models.
## 6. Deployment Fixes Applied
### Fix A — VerdictTank-Key allowlist (LiteLLM)
Added the 6 missing v4 models, keeping the existing 11 for backward compatibility. Total now 17 models.
### Fix B — MODEL_QUIRKS token floors (`/opt/verdicttank/worker.py`)
- `glm-5.2`: 8,000 → 24,000
- `deepseek-v4-pro`: 8,000 → 12,000
- `gemini-pro-latest`: added at 12,000
- `kimi-k2.6`: unchanged (already 16,000)
### Fix C — Robust JSON extraction (`/opt/verdicttank/worker.py`)
Replaced the fragile non-greedy regex + `rfind` parser (which silently dropped the financial seat) with brace-matching that handles nested objects, strings containing braces, and trailing commas in arrays. Verified against 6 test cases (fenced nested, fenced explicit, truncated, plain nested, trailing comma, empty) — all pass.
## 7. Remaining Work
- [ ] Re-run the full panel against the deployed worker to confirm end-to-end parity with the runner
- [ ] Decide whether glm-5.2 stays as permanent financial seat or whether the 24k floor is acceptable long-term cost
- [ ] Clear the six blocking conditions in §4 before GA
- [ ] Rotate the temp key `verdicttank-v4-test` (or delete it) once the production key is confirmed working