05Judge Pool v2.3: 11 Seats, 9 Vendors, Zero Double-Ups

The panel that produced the validation data ran at 8 of 11 seats because two seats hit a provider credit wall mid-run and one model family could not be dispatched at all. v2.3 is the spec written in response to that failure. Full detail lives in the judge pool specification v2.3.

The roster

BandSeatModelVendorScores
0Research AgentGrok 4.5xAINo
APrimary ReviewerClaude Opus 5AnthropicYes
ACross-Check ADeepSeek V4 FlashDeepSeekYes
ACross-Check BGemini Pro LatestGoogleYes
ACross-Check CDeepSeek V4 ProDeepSeekYes
ALegal / RegulatoryClaude Sonnet 5AnthropicYes
BFinancial IntegrityMiniMax-M3MiniMaxYes
BTeam / FounderClaude Fable 5AnthropicYes
BMarket RealityQwen3.7 PlusAlibabaYes
BExecution FeasibilityGPT-5.2 ProOpenAIYes
CSynthesis & Integrity GateKimi K2.6MoonshotNo

Nine distinct vendors across eleven seats. Nine distinct scoring models. Zero model double-ups: no single model occupies two scoring seats, which is the constraint that keeps correlated failure out of the panel mean. Maximum vendor concentration is Anthropic at 3 of 11 (27.3%), comfortably inside the 40% ceiling. DeepSeek holds 2 of 11 (18.2%). Every remaining vendor holds exactly one seat.

What changed in v2.3

Pre-flight health gate

Before any scoring begins, the orchestrator pings every rostered model with a 5-second probe and writes the result to a per-model health file. Any model returning HTTP 400, HTTP 429, or a no-healthy-deployments error is swapped for its pre-assigned failover before a single scoring call is spent. The 2026-08-12 run burned roughly 12 dispatches discovering dead models at runtime. That failure mode is now closed.

Pre-baked failover roster

Every seat carries a named failover from a different vendor, resolved at gate time rather than improvised mid-run. Failover selection preserves both the vendor-diversity ceiling and the no-double-up rule, so a degraded panel is still a valid panel rather than an accidentally correlated one.

Credit-wall resilience

The provider credit exhaustion that cost two seats mid-run is now detected at the gate and treated as an availability failure, not an error. Anthropic capacity has been restored and the Primary Reviewer seat runs Claude Opus 5 as specified.

Permanent exclusions

One frontier model family proved structurally incapable of running as a panel seat under our orchestration and is permanently excluded from the roster, not merely deprioritized. Excluded models cannot be selected as a failover target either.

Latency

v2.3 targets a critical path of roughly 113 seconds, against 218 seconds measured on the v2.1 architecture. The improvement comes from band parallelism: Band A and Band B seats execute concurrently rather than sequentially, and the Synthesis seat is the only stage that must wait for all scoring seats to return.

The ten scored dimensions

Every scoring seat rates the proposal 1 to 10 on the same ten dimensions, so panel spread is computed dimension by dimension and not only in aggregate:

Problem Clarity · Market Opportunity · Product Differentiation Revenue Model Viability · Go-to-Market Strategy · Competitive Moat Financial Projections · Team / Execution · Risk Mitigation · Legal / Compliance

The Synthesis and Integrity Gate

The Band C seat never scores. It reads all scoring output and performs a fixed checklist: verify score arithmetic, compute panel means and per-dimension spread, compute the delta against the solo baseline, flag every score more than 1.5 standard deviations from the panel mean with a written rationale, and confirm the verdict follows from panel evidence rather than from the Primary Reviewer alone. That gate is what turns eleven opinions into one auditable report.

06Worked Example: VentureBuilt v2, Where the Score Said Nothing

This is the clearest case in the validation set, because it is the one where score elevation delivered exactly zero and error detection delivered everything. Real scores from the 2026-08-12 run.

Panel mean · 8 judges
6.14/10

Median 6.45 · standard deviation 1.758 · spread 5.4 (min 2.8, max 8.2)

Solo baseline · Claude Opus 5
6.10/10

Spread 2.4 · delta +0.04 · statistically tied with the panel

Individual seat scores

SeatModelScoreSigma from meanFlag
Band B · Financial IntegrityMiniMax8.2+1.17Within 1.5 sigma
Band B · Market RealityGemini7.7+0.89Within 1.5 sigma
Band A · Primary ReviewerOpus 56.9+0.43Within 1.5 sigma
Band A · Legal / RegulatoryQwen6.5+0.21Within 1.5 sigma
Band A · Cross-Check CDeepSeek V4 Pro6.4+0.15Within 1.5 sigma
Band A · Cross-Check BGemini6.2+0.04Within 1.5 sigma
Band B · Team / FounderKimi4.4-0.99Within 1.5 sigma
Band B · Execution FeasibilityDeepSeek V4 Pro2.8-1.91OUTLIER
The average is a lie of composition. A 6.14 reads as a solid, fundable proposal with room to improve. The distribution says something completely different: the money is excellent (8.2) and the delivery plan is close to unworkable (2.8). Those are not two opinions about one thing. They are two accurate findings about two different things, and averaging them produces a number that describes neither.

What the outlier actually found

The 2.8 was not a grumpy model. The Integrity Gate challenged it at 1.91 sigma and it survived the challenge on evidence: a contractor budget broken 4x to 7x, an architecture that regressed from v1 with no schema and no API contract, and a founder workload of 37 engagements plus 950 hours alongside an MSP day job. The Team seat (4.4) and the Band A generalists (6.2 to 6.9) all acknowledged the same workload problem. They weighted it less severely. The Execution specialist is the only seat that forced it into the verdict.

Remove that one seat and the panel averages 6.6, which reads as a clean GO and ships a proposal with an 800-hour founder gap in it. The specialist seat cost a few cents of inference and changed the verdict.

Fix-It items, ranked by materiality

  1. CriticalExecution
    Fix the contractor budget or cut the scope. $1,500/mo buys 105 to 140 hours at market rates, not the volume the plan assumes. Raise to roughly $7,500/mo or reduce scope to fit the hours actually purchased.
  2. CriticalFinancial
    Restate Year 2 on a recognized-revenue basis. On run-rate the year looks healthy. On recognized revenue it loses roughly $11K to $18K. Present both.
  3. HighCompetitive
    Withdraw or qualify the uniqueness claim. LivePlan Plan Review and IdeaProof already ship in this space. Reposition on a defensible axis.
  4. HighTeam
    Name the contractor and the sourcing plan before Phase 2. A budget line with no named person is not a capacity plan.
  5. MediumGo-to-market
    Map the Year 1 to Year 2 GTM bridge. Eight net-new signups per month appear in the model with no acquisition mechanism behind them.
  6. MediumLegal
    Complete data protection and trademark clearance before Phase 0 to 1.

Panel verdict: CONDITIONAL GO. Estimated rework 15 to 20 hours. The solo baseline returned a 6.10 and none of the six conditions above.