v5.0 · Pre-Revenue · Validation-Tested

14 blind errors your solo model missed

A solo frontier model gives you a smooth, confident score. An 11-judge panel gives you the 14 things it was wrong about. VerdictTank does not sell you a higher number. It sells you the errors that number was hiding.

Category: AI Proposal Review & Error Detection Panel: 11 seats · 9 vendors Architecture: Full technical document Judge pool: Spec v2.3

01The Thesis Changed, Because the Data Said So

VerdictTank v4.0 was sold on a claim we could not defend: that a multimodel panel produces a better score than a single strong model. On 2026-08-12 we ran that claim against real proposals and it failed. What we found instead is a stronger product.

14
Material errors caught
-0.40
Aggregate score delta
3
Real proposals scored
11
Judge seats, 9 vendors

What failed

The score-elevation thesis. Across three real proposals the panel mean came in 0.40 points below the solo baseline. The panel did not lift scores. On two of three proposals it pushed them down. We are publishing that result rather than burying it, because the reason it happened is the product.

What worked

Error detection density. The same panel run surfaced 14 material errors that the solo baseline missed or underweighted: revenue arithmetic that was wrong by a factor of seven, a funded direct competitor the solo pass never named, a launch-blocking compliance cost larger than projected first-year revenue. None of those show up as a score. All of them decide whether the proposal wins.

You do not buy VerdictTank to get a higher score.
You buy it to find the $60K compliance hole and the broken contractor budget before you ship.

Why the spread is the signal

A single model scoring alone produces low variance. It reads the document once, forms one coherent opinion, and every dimension it emits is downstream of that opinion. The result feels authoritative precisely because nothing inside it disagrees.

An 11-seat panel of nine different vendors cannot produce that coherence, and the incoherence is diagnostic. When a Financial Integrity judge scores a proposal 8.2 while an Execution Feasibility judge scores the same document 2.8, that 5.4-point spread is not noise. It is a precise statement: the money works, the delivery plan does not. A solo model averages that tension away into a single confident 6.1 and tells you nothing actionable.

Measured, not asserted: on every one of the three validated proposals, panel spread exceeded solo spread. RFP Tank 3.7 vs 1.4. VentureBuilt 5.4 vs 2.4. CartMySupply 2.6 vs 1.8. Widening variance is the intended behavior, not a defect to tune out.