Files
root f5175f1ce0 Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
2026-08-26 02:27:28 -04:00

412 lines
23 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# VerdictTank Judge Pool Spec v2.0 - Primary Reviewer Scorecard
**Reviewer:** Claude Opus 5 (Primary Reviewer, Phase 2)
**Date:** 2026-08-12
**Spec:** https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md (17,641 bytes, v2.0)
**Method:** Document critique + live empirical validation against admin-ai (157-model catalog, real inference runs)
> **Note on version:** `web_extract` returned a cached **v1.0**. Verified against the live origin
> via `curl` - the deployed file is **v2.0**, byte-identical to
> `/root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md`. This review is of v2.0.
---
## Scorecard
| # | Dimension | Score | Verdict |
|---|---|---|---|
| 1 | Role Coverage | 7/10 | FIX |
| 2 | Model Assignments | 6/10 | FIX |
| 3 | Vendor Diversity | 8/10 | DEFEND |
| 4 | admin-ai Paths | 9/10 | DEFEND |
| 5 | Fallback Chains | 4/10 | FIX |
| 6 | Content-Based Rules | 5/10 | FIX |
| 7 | Latency Budgets | 2/10 | FIX |
| 8 | Feedback Loop | 5/10 | FIX |
| 9 | Pricing | 4/10 | FIX |
| 10 | Product Boundary | 7/10 | DEFER |
**Weighted mean: 5.7/10** - Architecturally literate, empirically unvalidated. Two dimensions (7, 9) are
launch-blocking. One unlisted finding (§Critical Finding) threatens the product thesis itself.
---
## CRITICAL FINDING (unlisted dimension - read first)
**The 12-judge panel does not measurably outperform one model run once.**
I tested this directly. Same proposal, same rubric, three dimensions, live models.
**A) Single-model noise floor** - `claude-opus-5` x6 at temp 1.0:
```
market: 5,5,5,4,4,4 stdev 0.50
team: 4,4,4,4,4,4 stdev 0.00
fin: 3,3,3,3,3,3 stdev 0.00
```
**B) Nine distinct judges across 8 vendors, one run each:**
```
Opus5 [5,4,3] Sonnet5 [4,4,3] Sol [6,4,3] GeminiPro [4,3,2] DeepSeek [5,3,2]
Terra [4,4,3] Qwen [4,3,2] MiniMax [6,3,4] Fable5 [6,4,4]
panel stdev: market 0.87, team 0.50, fin 0.74
```
**C) The result that matters:**
| | market | team | financials |
|---|---|---|---|
| Mean, 1 model | 4.50 | 4.00 | 3.00 |
| Mean, 9 judges / 8 vendors | 4.89 | 3.56 | 2.89 |
| **Delta** | **0.39** | **0.44** | **0.11** |
Nine frontier models from eight vendors, ~218s of wall clock and roughly 20x the token cost, move the
final number by **less than half a point on a 10-point scale**. On `market`, panel spread (0.87) is only
**1.7x** the single-model rerun noise (0.50) - statistically indistinguishable from rerunning one model
six times.
The individual judges *do* disagree (market scores span 4-6). But disagreement that averages to the same
answer is **variance, not signal**. The spec sells vendor diversity as the core accuracy mechanism and
never once tests whether diversity changes the output.
**This is the product thesis, and it is unvalidated.** Every downstream claim - premium pricing, the
"most accurate review tool available" positioning, the 12-vendor moat - rests on it.
**Required before build:** run 30-50 real proposals with known outcomes. Report panel-vs-single
correlation against ground truth. If the panel does not beat one good model by a margin that justifies
20x cost, the correct architecture is 3 judges, not 12 - and the pricing story needs rebuilding.
Better to learn this now than after a customer runs the same A/B.
---
## 1. Role Coverage - 7/10 - FIX
v2.0 deserves credit: it closed three of the four gaps flagged in the prior review (Financial Integrity,
Team/Founder, Legal/Regulatory now have dedicated phases). That is real progress.
**Technical Architecture remains uncovered** - the one gap explicitly identified in the v4.1 gap analysis
and silently dropped. Execution-Feasibility is operational (timeline, team, resources), not architectural
(stack, scalability, security posture, technical debt). For a product whose buyers are evaluating
*technical* startups, having no judge that reads the architecture is a conspicuous hole.
Two further gaps neither version names:
- **Traction/Evidence.** Nobody scores whether claims are *substantiated*. Research Agent gathers
citations but never scores; no downstream role is tasked with "the founder asserts 40% MoM growth -
is there evidence?" This is the single most common reason real proposals fail diligence.
- **Narrative/Communication quality.** For a *proposal* review tool, no judge assesses whether the
document actually persuades.
**Fix:** add Technical Architecture (merge into Execution-Feasibility if headcount is capped) and fold
an evidence-substantiation mandate into the Research Agent's brief so grounding produces a scored
claims-verification artifact, not just citations.
---
## 2. Model Assignments - 6/10 - FIX
Most seats are defensible. Grok 4.5 on Research is correct (only pool member with native live web
search). Qwen3.7 Plus on Market-Reality for a non-Western lens is genuinely thoughtful.
**Empirically-grounded objections:**
**Kimi K2.6 on Reasoning-Verification is misassigned.** In my scoring test it returned **empty content
after burning all 900 output tokens** - and on the 12-page critical-path run it took 15.1s. The role
requires emitting a structured contradiction report; a model that silently exhausts its budget in the
seat that validates every other judge's consistency is the worst possible placement. It is also the
*only* Moonshot seat, so there is no same-vendor fallback.
**Empty-response rate is a systemic risk the spec never models.** At `max_tokens=12`, four of six probed
models (MiniMax-M3, Kimi K2.6, Claude Fable 5, Gemini Pro) returned **empty content** - reasoning tokens
consumed the entire budget. §4.1 handles "malformed score" but not "well-formed empty response," which is
the actual failure mode I observed. Token budgets must be set per-model with reasoning-token headroom.
**Claude Sonnet 5 sits in two scoring seats** (Validation + Legal/Regulatory). The spec waves this off as
"read/analytical, not scoring-intense," but §1.1 marks Legal/Regulatory as **"Yes - scores."** The
justification contradicts the table two rows above it. Same weights, same biases, two votes.
**Gemini Pro Latest doubles as Cross-Check B and Audit Agent.** The Audit Agent's stated job is catching
*groupthink* - it cannot audit a panel it already voted in. This directly violates the spirit of §3.1's
own audit-independence rule.
---
## 3. Vendor Diversity - 8/10 - DEFEND
The strongest dimension. Eight vendors, max 25% - a genuine improvement over v1.0's 44% Anthropic
violation, and the rule was tightened from 40% to 33% rather than loosened to fit. That is the right
instinct and the spec should defend it.
**Two caveats worth documenting rather than fixing:**
*Vendor diversity is not architecture diversity.* Nearly every pool member is a transformer trained on
overlapping web corpora with similar RLHF conventions. My independence test showed all three "independent"
cross-checks scoring `market` at **exactly 6 - spread 0.00**. Eight logos, one prior. The correlation
data from the Critical Finding is the real story here.
*Infrastructure concentration.* All 12 models route through a single admin-ai/LiteLLM proxy. Eight-vendor
diversity buys nothing if the proxy is down - that is the actual SPOF, and the spec's §4.1 fallback table
implicitly assumes the proxy always answers.
---
## 4. admin-ai Paths - 9/10 - DEFEND
**Verified live, not taken on faith.** I queried the admin-ai catalog (157 models) and confirmed all 13
spec paths resolve exactly, then ran real inference against every one:
```
OK xai/grok-4.5 OK claude-opus-5 OK claude-sonnet-5
OK gpt-5.6-sol OK gemini/gemini-pro-latest OK deepseek-v4-pro
OK kimi-k2.6 OK gpt-5.6-terra OK qwen3.7-plus
OK MiniMax-M3 OK claude-fable-5 OK gemini/gemini-3.6-flash
OK gemini/gemini-3.5-flash-lite
```
All 11 judge models returned live responses. The spec's claim "all model paths confirmed against admin-ai"
is **true** - rare enough in a v2.0 draft to call out. This dimension should be defended as-is.
**The one point off:** `gemini/gemini-pro-latest` is a **floating alias**, not a pinned version. The
catalog carries pinned alternatives (`gemini/gemini-3.1-pro-preview`, `openrouter/google/gemini-3-pro-preview`).
Google can repoint that alias with no notice and silently change two seats - Cross-Check B *and* Audit.
My determinism probe on that alias returned `[6,3,3] / [6,4,3] / [6,4,2]` across three identical calls.
For a product whose entire value is score reproducibility, and which promises T+90/180/365 longitudinal
outcome tracking, an unpinned model destroys year-over-year comparability. **Pin every judge seat.**
---
## 5. Fallback Chains - 4/10 - FIX
§3.1 asserts "Fallback chains always cross vendor boundaries - Enforced at the assignment table (§2.3)."
Reading §2.3 against that claim, the rule is **violated in the first fallback hop of 4 of 11 rows**:
| Role | Primary | Fallback 1 | Violation |
|---|---|---|---|
| Primary Reviewer | Claude Opus 5 | **Claude Sonnet 5** | Anthropic → Anthropic |
| Cross-Check A | GPT-5.6 Sol | **GPT-5.6 Terra** | OpenAI → OpenAI |
| Cross-Check B | Gemini Pro Latest | **Gemini 3.6 Flash** | Google → Google |
| Team/Founder | Claude Fable 5 | **Claude Sonnet 5** | Anthropic → Anthropic |
This is the *identical* defect flagged in the v1.0 review ("§4.1 says different vendor but §2.3 says
Opus→Sonnet"). It was marked as a top-5 prioritized fix, and it was **not fixed** - the assertion text was
added to §3.1 without correcting the table it points at. A rule that is stated but not enforced is worse
than no rule: it will pass code review as "already handled."
The failure mode is precisely what fallbacks exist to prevent. Anthropic has a regional outage → Primary
Reviewer fails over to Sonnet 5 → also Anthropic → also down. Same for the OpenAI and Google rows.
**Additional defects:**
- **Convergence, not diversity.** DeepSeek V4 Pro appears as a fallback in **7 of 11 chains**. Under
broad degradation the 8-vendor panel collapses toward a single DeepSeek-dominated panel, blowing the
33% vendor cap at exactly the moment it matters most. No runtime re-check of the cap after failover.
- **Kimi K2.6 has no same-tier peer** - the sole Moonshot seat degrades straight to DeepSeek (Tier C).
- **§4.2 is unexecutable as written.** "Tier B for Primary/Validation only, Tier C for others" - Tier C
is `DeepSeek V4 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite`: two vendors, three models, for up to
five simultaneous seats. The cost-optimization path *cannot* satisfy the vendor rule. This is the same
unexecutability flagged in v1.0 ("Tier B has no OpenAI/Google model") reappearing in a new form.
---
## 6. Content-Based Rules - 5/10 - FIX
The vertical/stage triggers are directionally sensible and v2.0 adds a genuine improvement: the §3.3
missing-vertical handler with explicit user disclosure is honest product design.
**But the weighting system is undefined.** Every rule adjusts "weight ±N%" and **the spec never states
what the weights weight.** There is no aggregation formula anywhere in 302 lines - no statement of how 11
judges' 10-dimension scores combine into a final number. "Market-Reality weight +30%" is meaningless
without knowing the base weight, the combination function, and whether weights renormalize. This is the
single largest specification hole in the document: it is the core scoring algorithm, and it is absent.
**Composition rule is underspecified.** §3.9 caps stacking at +50% per role but doesn't say what happens
when the cap binds - are all triggers scaled proportionally, or does first-match win? Different answers
give different scores for the same proposal.
**Detection is hand-waved.** Every row says "Detected from proposal text" with no mechanism, no confidence
threshold, and no misclassification path. Vertical detection *changes the score* - a fintech misread as
SaaS loses its +30% Legal/Regulatory weight. That needs a confidence score and a human-review fallback,
and multi-vertical proposals (fintech + healthcare) have no representation at all.
**Page count is a poor complexity proxy.** A 10-page dense technical proposal triggers the *reduced*
panel; a 50-page deck with 40 pages of appendix triggers the full one. Use token count or content
density.
**Substantively questionable:** Seed-stage sets Financial Integrity **-20%**. Seed is where financial
models are *most* fictional and cap-table mistakes are most permanent. De-weighting financial scrutiny
where founders most need it inverts the diligence priority.
---
## 7. Latency Budgets - 2/10 - FIX (launch-blocking)
The spec revised Enterprise from 30s → 45s and annotated it *"v1.0 was unrealistic for 12-judge pipeline."*
The revision is still off by nearly 5x. I measured it.
**Sequential critical path, 12-page proposal (~5,500 tokens), following the spec's own dependency graph:**
```
P1 Research xai/grok-4.5 21.5s
P2 Primary claude-opus-5 20.0s
P3 Validation claude-sonnet-5 15.8s
P5 Reasoning kimi-k2.6 15.1s
P6 Execution gpt-5.6-terra 16.4s
P7 Market qwen3.7-plus 67.2s <-- single judge exceeds Enterprise MAX alone
P8 Financial MiniMax-M3 26.1s
P9 Team claude-fable-5 22.9s
P11 Audit gemini/gemini-pro-latest 13.0s
------------------------------------------------
TOTAL 218.0s
```
**218s against a 45s target and a 90s max - 4.8x over target, 2.4x over the stated maximum.**
This is not a tuning problem, it is structural:
- **The pipeline is inherently sequential.** Phases 2→3→5→6→7→8→9→11 each consume prior output by design.
Only 4a/4b/4c parallelize. You cannot fan out a dependency chain.
- **A single judge blows the entire budget.** Qwen3.7 Plus took **67.2s alone** - 1.5x the full Enterprise
target, before any other judge runs. It burned 1,450 reasoning tokens on a trivial scoring task.
- **Even perfect parallelism fails.** All 11 models fired simultaneously on a *short* prompt still took
**27.4s wall clock** - 61% of the Enterprise budget consumed by the slowest model, with zero
orchestration, retries, or assembly.
- **Fallbacks make it worse.** A 30s timeout + Fallback 1 retry adds 30s+ to an already-blown budget.
The §4.1 retry path and the §4.3 budget are mutually unsatisfiable.
- **Tier ordering is still inverted.** Enterprise runs the deepest pipeline (12 judges) on the tightest
budget (45s); Free runs 3 judges on 90s. v1.0 had this backwards and v2.0 preserved the inversion while
only adjusting magnitudes.
50+ page proposals (the spec's own "High complexity," which *adds* a secondary pass) will run
substantially past 218s.
**Fix:** these are not real-time interactions - they are deep-analysis jobs. Re-architect as an
**asynchronous job model**: submit → progress streaming → notify on completion. Budget **5-10 minutes**
for Enterprise and sell the depth. Then set per-phase timeouts against measured p95, not aspiration.
A 45s promise that reliably takes 218s is a support-ticket generator and a churn driver.
---
## 8. Feedback Loop - 5/10 - FIX
v2.0 genuinely improved here - §5.4 cold-start gate, min N=30/N=50 thresholds, and the immediate-recusal
rule for >2.0σ bias are all correct additions that address prior findings.
**The remaining problems are foundational:**
**The accuracy metric is not sound.** "Correlation between dimension score and T+90 outcome" -
- **Ninety days is far too short.** Seed rounds take 3-9 months; the T+90 signal is mostly noise about
fundraising *timing*, not proposal quality.
- **Survivorship and selection bias are unaddressed.** Response rates on outcome surveys skew heavily to
founders who succeeded. The spec plans to weight models on a systematically biased sample.
- **Confounding is total.** A proposal that scores 4/10, gets rewritten using the Fix-It plan, and then
raises successfully - did the judge score correctly or incorrectly? The product *intervenes* on the
outcome it measures. This is unfixable by more data; it needs a holdout design.
- **N is unreachable.** N≥50 per model per dimension per vertical, with 10 dimensions, 11 models and 5+
verticals, implies thousands of tracked reviews before a single threshold fires. At 100 reviews/mo
Enterprise capacity, that is **years**. The entire feedback loop is aspirational at realistic volume,
and the "data moat" narrative rests on it.
**Rotation rules:** "Remove from Tier A, demote to Tier B" as a demotion path is odd - a model with
<0.3 outcome correlation is not a *cheaper* model, it is an *inaccurate* one. Demoting it means budget
users get the judge known to be wrong. Also, the model-deprecation row says replace with "Fallback 1
from §2.3" - but §2.3 fallbacks are same-vendor in 4 rows, so vendor-caused deprecation cascades to a
sibling that may be deprecated in the same wave.
---
## 9. Pricing - 4/10 - FIX
I have no objection to premium positioning, and the "don't race to the bottom" instinct is right. The
objection is that **the price is not connected to demonstrated value**, and v2.0 raised it by
**3.3x-5.0x** (from $79/$299 to $249/$799/$1,499) on positioning reasoning alone, with zero customers,
zero LOIs, and - per the Critical Finding - no evidence the 12-judge panel beats one model.
**The benchmark is misapplied.** GC AI at $500/seat/mo is cited as the anchor comp, and I verified it
independently (gc.ai, corroborated by vaquill.ai's 2026 benchmark; note haqq.ai could not source it to
GC AI directly). But GC AI serves **in-house legal teams at 1,900+ companies** - daily-use workflow
software with seat-level lock-in. VerdictTank is **episodic**: a founder reviews a proposal during a
fundraise, then churns. Anchoring episodic tooling to daily-workflow pricing is a category error.
**The usage math undermines the tiers.** Enterprise at $799/mo = 100 reviews. Real founders raising a
round need **3-8 reviews over a 2-3 month window**. That is ~$100/review nominal at a utilization
almost no customer will reach - and Pro at $249 for 20 reviews has the same problem. Customers pay for
capacity they cannot consume, notice, and churn. **Per-review or credit-pack pricing fits actual
consumption far better than monthly seats**, and the spec never considers it.
**Margin honesty.** The v4.1 reference concedes "COGS 72-97% margin - pure positioning play." My cost
sampling supports that: a full panel run is roughly $0.30-0.80 in tokens. A 99.9% gross margin at
$799/mo is not premium positioning, it is an unanchored price waiting for a competitor to undercut it
with the same off-the-shelf models. The moat is claimed to be the outcome-tracking corpus - which
dimension 8 shows is years away at this volume.
**Free tier at 1/mo is too stingy** for a trust-first product. The entire pitch is "we tell you the
brutal truth" - that requires *experiencing* the depth. A 1-review Tier-C sample (3 judges, no Fix-It)
demonstrates the weakest possible version of the product to every prospective buyer.
**Fix:** validate willingness-to-pay with 10-20 design partners before locking. Offer per-review pricing
alongside subscriptions. Anchor to *outcome value* (a better raise) rather than to a competitor in an
adjacent category.
---
## 10. Product Boundary - 7/10 - DEFER
Conceptually clean and easy to communicate: VerdictTank = "is this good?", RFP Tank = "does this match
what they asked for?" The superset framing is right, and inheriting one engine is the correct build
decision.
**Unresolved, but not urgent:**
- **RFP Tank is the better business and it is the side project.** RFP responses are recurring, deadline-
driven, budgeted, and B2B - structurally superior to episodic founder fundraising. The spec treats it
as a discount add-on. Strategically inverted.
- **Cannibalization is unpriced.** RFP Tank is a strict superset at (presumably) a higher price. A
rational buyer needing both buys RFP Tank only. Enterprise VT + 20% off RFP Tank is then a discount
on a product that replaces the one just paid for.
- **10-dimension rubric is asserted, never enumerated.** The document references "10-dimension scoring"
in §7 and §8 but **never lists the ten dimensions.** For a build spec, the thing being scored should
be defined; I am scoring against dimensions the spec assumes I already know.
- No shared-account model, no cross-product SSO, no migration path.
**DEFER** because the boundary is directionally correct and none of this blocks the judge-pool build.
Revisit before RFP Tank pricing is set.
---
## Priority Actions
| # | Action | Dim | Severity |
|---|---|---|---|
| 1 | **Validate panel-vs-single-model accuracy on 30-50 real proposals.** Product thesis is unproven. | - | **BLOCKER** |
| 2 | **Re-architect to async jobs; budget 5-10 min.** Measured 218s vs 45s target. | 7 | **BLOCKER** |
| 3 | **Fix 4 same-vendor Fallback-1 hops.** Flagged in v1.0 review, still unfixed. | 5 | **BLOCKER** |
| 4 | **Define the score aggregation formula.** "Weight +30%" is meaningless; core algorithm absent. | 6 | **BLOCKER** |
| 5 | Pin all model versions - drop floating `gemini-pro-latest` alias. | 4 | High |
| 6 | Reassign Kimi K2.6 off Reasoning-Verification (empty output under budget). | 2 | High |
| 7 | Set per-model token budgets with reasoning-token headroom; handle empty-but-valid responses. | 2 | High |
| 8 | Split Gemini Pro's Cross-Check B / Audit double-seat; auditor cannot audit itself. | 2 | High |
| 9 | Add runtime vendor-cap re-check after failover (DeepSeek is fallback in 7 of 11 chains). | 5 | High |
| 10 | Validate pricing with design partners; add per-review option. | 9 | High |
| 11 | Add Technical Architecture + evidence-substantiation coverage. | 1 | Medium |
| 12 | Replace page-count complexity proxy with token count. | 6 | Medium |
| 13 | Enumerate the 10 scoring dimensions in the spec. | 10 | Medium |
| 14 | Reconsider Financial Integrity -20% at seed stage. | 6 | Medium |
---
## Bottom Line
v2.0 is a real improvement over v1.0 - the vendor rule was tightened rather than loosened to fit, the
cold-start gates are correct, the role gaps were mostly closed, and **every model path is genuinely live**,
which I verified rather than assumed. The document is architecturally literate.
But it is a **design document wearing a build spec's clothes**. Its two most load-bearing quantitative
claims fail on contact with the live system: the latency budget is off by 4.8x, and the fallback
cross-vendor guarantee is contradicted by its own assignment table in four rows - *the same defect flagged
in the v1.0 review and marked as a prioritized fix.* The assertion was added; the table was not corrected.
Most seriously: the panel-diversity premise that justifies the pricing, the vendor spread, and the entire
category claim has **never been tested**, and my measurement suggests it may not survive testing. Nine
judges across eight vendors moved the score by under half a point versus a single model.
**Recommendation: do not proceed to build.** Run the accuracy validation (Action 1) and re-baseline
latency (Action 2) first. If the panel advantage is real, this architecture is worth building and the
premium price defensible. If it is not, the correct product is 3 judges at a third of the price - and it
is far cheaper to discover that now than after the first enterprise customer runs the same A/B I just ran.