Mark v5.0/v5.1 SUPERSEDED (error-detection thesis failed at -0.40 delta); v4.1 is canonical. Cut v4.1 proposal with 12 version strings bumped. Moonshot to Mistral across production seats; production worker de-kimi'd 2026-08-18. Data-retention posture corrected 7/9 to 8/9 no-training (DeepSeek sole exception). Reconciled COGS with measured Mistral spend. Committed deployed v4.0 content and research docs to resolve the repo/live fork.
83 lines
6.2 KiB
Markdown
83 lines
6.2 KiB
Markdown
# VerdictTank — Changelog
|
|
|
|
## 2026-08-18
|
|
|
|
### v4.0 Dogfood + Panel Integrity Fixes
|
|
- v4.0 proposal run through the v4.0 single-pass panel (9 scoring seats + synthesis gate): **CONDITIONAL** (Proposal Strength 71, Investor Readiness 44, Composite 58)
|
|
- Root-caused glm-5.2 "flakiness": reasoning model token-starved on the financial seat (11,998 reasoning tokens at a 12k floor = zero visible output). Not a model defect — budget bug. glm-5.2 stays.
|
|
- Fixed VerdictTank-Key allowlist (LiteLLM): was v3.5-era with 11 models, 403'd on 6 v4.0 roster models. Now 17 models (11 existing + 6 v4 additions)
|
|
- Fixed MODEL_QUIRKS floors in worker.py: glm-5.2 8k→24k, deepseek-v4-pro 8k→12k, gemini-pro-latest added at 12k
|
|
- Replaced fragile extract_json parser with brace-matching (handles nested JSON, trailing commas, embedded braces) — was silently dropping the financial seat
|
|
- Full dogfood report: /root/projects/verdicttank/dogfood-v4-report.md
|
|
- Fixed three proposal arithmetic errors (Condition 6): ceiling volume 3,940→2,690 reviews/mo; annual-mix MRR reduction $2,946→$2,619; annualized prepaid $187,000→$156,000. Fixed in v4-proposal.md and index-v4.0-new.html.
|
|
- Added bottom-up SAM + competitive research (Condition 1, research half): US SAM $1.2M-$3.7M/yr, base ~$1.8M; 8 named competitors. /root/projects/verdicttank/sam-competitive.md
|
|
- Added nine-vendor data-retention + training research (Condition 4, research half): 7 of 9 no-training by default; Moonshot trains by default (no opt-out), DeepSeek silent + PRC storage. /root/projects/verdicttank/vendor-retention-facts.md
|
|
- Corrected data-handling.html third-party section: "five of nine" → "seven of nine" no-training, precise two-exception wording
|
|
|
|
### v4.1 Cut + Moonshot→Mistral Swap + v5.x Superseded
|
|
- Cut v4.1 proposal as the go-live artifact. Version bumped v4.0 → v4.1 across title, hero badge, document label, all prose references, and footer (12 strings).
|
|
- Moonshot→Mistral vendor swap executed across all production seats (market primary + crosscheck_a/c fallbacks → mistral-large-latest). Production worker de-kimi'd 2026-08-18; py_compile clean; systemd verdicttank-worker restarted active. Defensive Kimi/Moonshot VENDOR_PATTERNS scrub lines retained.
|
|
- Data-retention posture corrected: 7/9 → 8/9 no-training. Moonshot (trains by default) replaced by Mistral (no-training default, EU-resident). DeepSeek is now the sole training exception. data-handling.html updated ("Eight of the nine" / "One provider does not offer") and deployed to app3.
|
|
- Reconciled COGS with measured Mistral token usage (SpendLogs startTime): Team 9,624 prompt / 953 completion / $0.00624; Synthesis Gate 20,675 / 1,841 / $0.01310.
|
|
- Marked v5.0 / v5.1 SUPERSEDED (error-detection-density direction deprecated; its validation thesis failed at -0.40 delta). v4.1 is canonical. Lineage reconciled: v3 (08-10) → v5.x pivot (08-12, superseded) → v4.1 (08-18).
|
|
- Repo/live fork resolved: deployed v4.0 content committed (index-v4.0.html now matches deployed), v4.1 added, superseded banners on v5 pages.
|
|
|
|
## 2026-08-10
|
|
|
|
### v3 Proposal + Technical Architecture
|
|
- v3 proposal drafted with 11 enhancements across 3 tiers and new 4-tier pricing
|
|
- New pricing: Free / Pro ($79/mo) / Enterprise ($499/mo) / White-Label ($1,999+/mo)
|
|
- Technical architecture documented (908 lines, 10 data model tables, 4 API endpoints)
|
|
- Architecture run through independent 3-judge technical review: 3/3 Conditional Go
|
|
- v2 proposal archived with "SUPERSEDED" banner
|
|
- All iamgmb.com references purged — docs, repo, and style guide updated to itpropartner.com
|
|
- verdicttank.com domain live on Core (Caddy)
|
|
- docs.itpropartner.com/verdicttank/ deployed with current state
|
|
|
|
## 2026-08-07
|
|
|
|
### Rebrand (SharkTank → VerdictTank)
|
|
- Project renamed: /root/projects/sharktank → /root/projects/verdicttank
|
|
- Mockup moved: mockups.itpropartner.com/sharktank → mockups.itpropartner.com/verdicttank
|
|
- New proposal: proposals.itpropartner.com/verdicttank (incorporates all pipeline review feedback)
|
|
- Original SharkTank proposal archived with "superseded" banner
|
|
- Mockup landing page: removed all model names, added file upload (doc/docx/pdf), purged all em dashes
|
|
- Proposals index: VerdictTank (primary) + SharkTank (archived, grayed out)
|
|
- Critical review page: rebranded, em dashes purged
|
|
|
|
### Pipeline Review Results
|
|
- First dogfood: ran proposal through full VerdictTank pipeline
|
|
- **Verdict: 3/3 Unanimous Conditional Go**
|
|
- Phase 1 (Research Agent): No direct competitor found — cross-vendor architecture is novel
|
|
- Phase 2 (Critic Agent): 4/10 average — 4 fatal flaws: trademark, pricing, financial model, no validation
|
|
- Phase 3a (Judge A): Ratified 85% of critic — pricing > name in priority, rejected co-founder as condition
|
|
- Phase 3b (Judge B): Caught overage incentive flaw, compound reliability risk (97.5% = 18hr/mo downtime)
|
|
- Phase 3c (Judge C): Meta-layer insight, training-data recursion (Year 3), liability asymmetry, SEO desert
|
|
- Critical review addendum: https://mockups.itpropartner.com/verdicttank/critical-review.html
|
|
- Priority-ranked 11-item action plan with reconciled timeline (Apr 2027 paid launch)
|
|
- Novel insights: product IS content, verdict confidence scoring, degraded-mode fallback needed
|
|
|
|
### Updated Proposal (Post-Review)
|
|
- Full rebrand to VerdictTank throughout
|
|
- Tiered usage-based pricing (no "unlimited"): $19/$49/$99/$299 + $14.99 pay-per-review
|
|
- Financial model with churn (5-7%), CAC ($15-25), LTV ($210-290), LTV:CAC (~10:1)
|
|
- Real break-even estimate: 25-35 users (not 6)
|
|
- Timeline pushed to Mar-Apr 2027 paid launch
|
|
- New risks: training-data recursion, liability asymmetry, alignment drift, SEO desert
|
|
- Pipeline described by roles only (no model names)
|
|
- Pipeline self-review section with verdict banner
|
|
- Added: degraded-mode fallback, benchmark accuracy report, advisory board plan
|
|
- Product origin: dogfooding at IT Pro Partner
|
|
|
|
### Pending
|
|
- Register verdicttank.com domain
|
|
- File VerdictTank trademark (USPTO Class 42)
|
|
- Landing page + waitlist at verdicttank.com
|
|
- Interview 20 target ICP users
|
|
- Publish benchmark accuracy report (20 proposals, known outcomes)
|
|
- Recruit 3+ named advisors
|
|
- FastAPI backend with job queue
|
|
- WeasyPrint PDF generation
|
|
- Email delivery via MXroute
|
|
- LiteLLM virtual key: verdicttank-prod
|