02How We Found This: Four Generations and a Self-Review

VerdictTank is a proposal review engine, not a proposal writer. It ingests a finished document and returns scored dimensions plus a ranked list of concrete Fix-It items. The pipeline was not designed in the abstract. It was hardened across four architectural generations, and then it was pointed at its own proposals.

v2

Single-Model Scorer

Breakthrough: the dual-axis rubric

v2 established the core insight: a proposal has two independent quality axes. Narrative quality (clarity, structure, persuasion) and compliance quality (does it actually answer the scored requirements). One model scored both from one prompt. It proved the concept and exposed the flaw: the axes bled together. A beautifully written section that missed a mandatory requirement scored too high, because the same reasoning pass that admired the prose also graded the compliance.

v3

Separated Scoring Passes

Breakthrough: axis isolation

v3 split scoring into two independent passes with two purpose-built prompts. The narrative pass never sees the compliance rubric. The compliance pass never rewards eloquence. This is the decision that makes the dual score trustworthy: the two numbers can now disagree, and their disagreement carries information. A 9/10 narrative next to a 4/10 compliance is a proposal about to lose.

v4

Multimodel Adversarial Review

Breakthrough: cross-model verification

A single model scoring in isolation is confidently wrong at a predictable rate. v4 introduced a multimodel pipeline: a fast model produces first-pass scores and Fix-It candidates, then a stronger model reviews that output adversarially, challenging every deduction and confirming each Fix-It maps to real proposal text. Scores stopped drifting between runs. This is the generation that made the output defensible.

v5 ยท Current

Specialist Panel and the Integrity Gate

Breakthrough: disagreement as output

v5 replaces the adversarial pair with an 11-seat specialist panel across nine vendors, and adds a synthesis seat whose only job is to compute panel statistics, flag scores more than 1.5 standard deviations from the mean, and reconcile the verdict against the evidence. The output is no longer a number. It is a number, a spread, an outlier list, and a ranked set of material errors with the judge that caught each one.

The self-review that broke the old thesis

Before selling a review engine we ran the engine on our own work. We assembled the panel and scored three real proposals, in full, with the same prompts and rubric a paying customer would get. One of the three was VerdictTank's own sibling product. The panel returned a NO GO on it.

That run cost roughly $150 in inference and returned 8 of 11 seats. Two seats were lost to a provider credit wall hit mid-run and one to a model family that could not be dispatched at all. The incomplete panel is why v2.3 of the judge pool spec now requires a pre-flight health gate and a pre-baked failover roster, covered in section 5. The results below are what those 8 seats produced, and we report them at 8 seats rather than extrapolating to 11.

The pipeline is its own reference implementation. The full architecture is documented in the companion technical architecture document to a standard where an engineer can implement it from the spec alone. VerdictTank is pre-revenue. We make zero claims about users, beta cohorts, or external validation. What we claim is narrower and verifiable: the architecture is built, the pipeline runs, it was executed against three real proposals on 2026-08-12, and it failed its own headline thesis in public.

03Validation Run: 3 Real Proposals, 8 Reporting Seats

Every figure in this section comes from the 2026-08-12 validation run. Nothing is modeled, projected, or illustrative. The solo baseline is Claude Opus 5 scoring the same documents against the same 10-dimension rubric.

Panel mean vs solo baseline

ProposalPanel meanSolo baselineDeltaPanel spreadSolo spreadVerdict
RFP Tank v1.04.404.93 -0.533.71.4 NO GO
VentureBuilt v26.146.10 +0.045.42.4 CONDITIONAL GO
CartMySupply4.295.00 -0.712.61.8 NO GO
Aggregate4.945.34 -0.40Panel spread exceeded solo spread on all 3 THESIS FAIL

Thesis under test: the panel must show a greater than 0.5 point advantage over the solo mean to justify premium pricing. Result: FAIL on all three proposals and FAIL in aggregate. Panel composition for this run was 8 reporting judges (4 Band A, 4 Band B) out of 11 specified seats, a 73% coverage rate.

Why the panel scored lower

The panel does not elevate scores. It sharpens error detection, and error detection on a flawed document moves the number down. All three proposals contained severe cross-cutting defects that additional specialist scrutiny exposed more precisely: fatal execution gaps, competitive mispositioning, and legal blockers. The solo baseline was directionally correct on all three. The panel added precision, not points.

That is the entire finding, and it inverts the sales pitch. If your proposal is sound, the panel will roughly agree with a good solo model and cost you more. If your proposal has a hole in it, the panel finds the hole and the solo model does not. You are not buying a score. You are buying the probability that a specific, expensive, named mistake gets caught before an evaluator or an investor finds it for you.

Specialist divergence, measured

ObservationEvidence from the runWhat it means
Generalist seats run optimistic Gemini Pro scored RFP Tank 6.7 as a Band A generalist and 3.0 as the Band B Market specialist. Same model, same document, 3.7 points apart. Band A generalist scoring without specialist cross-check is systematically over-optimistic. The role, not the model, drives the score.
Role divergence beats model divergence DeepSeek V4 Pro scored VentureBuilt 6.4 as Cross-Check C and 2.8 as Execution Feasibility. A 3.6 point split inside one vendor. Panel diversity is not primarily about buying different vendors. It is about buying different questions.
One seat can flip a verdict Remove the 2.8 Execution score from VentureBuilt and the panel averages 6.6, reading as a clean GO. With it, the verdict is CONDITIONAL GO with a named contractor-budget fix. The lowest score in the panel is frequently the only one doing work. Averaging is what a solo model already does.
Tight clustering is also a signal CartMySupply produced zero outliers beyond 1.5 sigma and the tightest spread of the three (sigma 0.89). Unanimity across nine vendors on a low score is a far stronger NO GO than one model's low score.