VerdictTank is a proposal review engine, not a proposal writer. It ingests a finished document and returns scored dimensions plus a ranked list of concrete Fix-It items. The pipeline was not designed in the abstract. It was hardened across four architectural generations, and then it was pointed at its own proposals.
v2 established the core insight: a proposal has two independent quality axes. Narrative quality (clarity, structure, persuasion) and compliance quality (does it actually answer the scored requirements). One model scored both from one prompt. It proved the concept and exposed the flaw: the axes bled together. A beautifully written section that missed a mandatory requirement scored too high, because the same reasoning pass that admired the prose also graded the compliance.
v3 split scoring into two independent passes with two purpose-built prompts. The narrative pass never sees the compliance rubric. The compliance pass never rewards eloquence. This is the decision that makes the dual score trustworthy: the two numbers can now disagree, and their disagreement carries information. A 9/10 narrative next to a 4/10 compliance is a proposal about to lose.
A single model scoring in isolation is confidently wrong at a predictable rate. v4 introduced a multimodel pipeline: a fast model produces first-pass scores and Fix-It candidates, then a stronger model reviews that output adversarially, challenging every deduction and confirming each Fix-It maps to real proposal text. Scores stopped drifting between runs. This is the generation that made the output defensible.
v5 replaces the adversarial pair with an 11-seat specialist panel across nine vendors, and adds a synthesis seat whose only job is to compute panel statistics, flag scores more than 1.5 standard deviations from the mean, and reconcile the verdict against the evidence. The output is no longer a number. It is a number, a spread, an outlier list, and a ranked set of material errors with the judge that caught each one.
Before selling a review engine we ran the engine on our own work. We assembled the panel and scored three real proposals, in full, with the same prompts and rubric a paying customer would get. One of the three was VerdictTank's own sibling product. The panel returned a NO GO on it.
That run cost roughly $150 in inference and returned 8 of 11 seats. Two seats were lost to a provider credit wall hit mid-run and one to a model family that could not be dispatched at all. The incomplete panel is why v2.3 of the judge pool spec now requires a pre-flight health gate and a pre-baked failover roster, covered in section 5. The results below are what those 8 seats produced, and we report them at 8 seats rather than extrapolating to 11.
Every figure in this section comes from the 2026-08-12 validation run. Nothing is modeled, projected, or illustrative. The solo baseline is Claude Opus 5 scoring the same documents against the same 10-dimension rubric.
| Proposal | Panel mean | Solo baseline | Delta | Panel spread | Solo spread | Verdict |
|---|---|---|---|---|---|---|
| RFP Tank v1.0 | 4.40 | 4.93 | -0.53 | 3.7 | 1.4 | NO GO |
| VentureBuilt v2 | 6.14 | 6.10 | +0.04 | 5.4 | 2.4 | CONDITIONAL GO |
| CartMySupply | 4.29 | 5.00 | -0.71 | 2.6 | 1.8 | NO GO |
| Aggregate | 4.94 | 5.34 | -0.40 | Panel spread exceeded solo spread on all 3 | THESIS FAIL | |
Thesis under test: the panel must show a greater than 0.5 point advantage over the solo mean to justify premium pricing. Result: FAIL on all three proposals and FAIL in aggregate. Panel composition for this run was 8 reporting judges (4 Band A, 4 Band B) out of 11 specified seats, a 73% coverage rate.
The panel does not elevate scores. It sharpens error detection, and error detection on a flawed document moves the number down. All three proposals contained severe cross-cutting defects that additional specialist scrutiny exposed more precisely: fatal execution gaps, competitive mispositioning, and legal blockers. The solo baseline was directionally correct on all three. The panel added precision, not points.
That is the entire finding, and it inverts the sales pitch. If your proposal is sound, the panel will roughly agree with a good solo model and cost you more. If your proposal has a hole in it, the panel finds the hole and the solo model does not. You are not buying a score. You are buying the probability that a specific, expensive, named mistake gets caught before an evaluator or an investor finds it for you.
| Observation | Evidence from the run | What it means |
|---|---|---|
| Generalist seats run optimistic | Gemini Pro scored RFP Tank 6.7 as a Band A generalist and 3.0 as the Band B Market specialist. Same model, same document, 3.7 points apart. | Band A generalist scoring without specialist cross-check is systematically over-optimistic. The role, not the model, drives the score. |
| Role divergence beats model divergence | DeepSeek V4 Pro scored VentureBuilt 6.4 as Cross-Check C and 2.8 as Execution Feasibility. A 3.6 point split inside one vendor. | Panel diversity is not primarily about buying different vendors. It is about buying different questions. |
| One seat can flip a verdict | Remove the 2.8 Execution score from VentureBuilt and the panel averages 6.6, reading as a clean GO. With it, the verdict is CONDITIONAL GO with a named contractor-budget fix. | The lowest score in the panel is frequently the only one doing work. Averaging is what a solo model already does. |
| Tight clustering is also a signal | CartMySupply produced zero outliers beyond 1.5 sigma and the tightest spread of the three (sigma 0.89). | Unanimity across nine vendors on a low score is a far stronger NO GO than one model's low score. |