Files

24 KiB
Raw Permalink Blame History

VerdictTank Judge Pool Specification v2.3

Status: PRODUCTION READY - validated on 3 real proposals (2026-08-12) Date: 2026-08-12 Changes from v2.2: Anthropic credits restored (3 seats back online). GPT-5.6-family permanent incompatibility confirmed - Sol → deepseek-v4-flash, Terra → gpt-5.2-pro. Grok 4.5 path fixed (bare → xai/grok-4.5). Pre-baked Anthropic failover roster added. Pre-pipeline health gate added. Product thesis pivoted from "+0.5 points" to "14 material errors caught vs solo model." Dead model families documented (GPT-5.6, Sonar, Command-R).


1. The 11-Seat Judge Roster

1.1 Pipeline Bands

The pipeline runs in 4 dependency bands. All seats within a band run parallel.

Band Phase Role Model Vendor admin-ai Path Scores?
0 1 Research Agent Grok 4.5 xAI xai/grok-4.5 No
A 2 Primary Reviewer Claude Opus 5 Anthropic claude-opus-5 Yes
A 3a Cross-Check A DeepSeek V4 Flash DeepSeek deepseek-v4-flash Yes
A 3b Cross-Check B Gemini Pro Latest Google gemini/gemini-pro-latest Yes
A 3c Cross-Check C DeepSeek V4 Pro DeepSeek deepseek-v4-pro Yes
A 4 Legal/Regulatory Claude Sonnet 5 Anthropic claude-sonnet-5 Yes
B 5 Financial Integrity MiniMax-M3 MiniMax MiniMax-M3 Yes
B 6 Team/Founder Claude Fable 5 Anthropic claude-fable-5 Yes
B 7 Market-Reality Qwen3.7 Plus Alibaba qwen3.7-plus Yes
B 8 Execution-Feasibility GPT-5.2 Pro OpenAI gpt-5.2-pro Yes
C 9 Synthesis & Integrity Gate Kimi K2.6 Moonshot AI kimi-k2.6 No

Band descriptions and inputs:

Band Description Input Output ~Time
0 Grounding - live web verification, citation gathering, factual baseline. Market block mandatory. Raw proposal Factual brief + market block + citations ~20s
A Blind scoring - five seats concurrent on proposal + Research brief only. Cross-checks blind to each other and to Primary. Legal included here because input is proposal + brief + vertical only (no scores dependency). CC scores averaged per §4.1 before Band B. Proposal + Research brief 5 independent score sets ~33s
B Informed specialists - four seats concurrent on proposal + brief + Band A scores. Mutually independent. Proposal + brief + Band A scores 4 specialty score sets ~30s
C Synthesis & Integrity Gate - contradiction detection, cross-model alignment, adversarial challenge (+700 tokens: "before synthesizing, argue strongest case against Primary's scores"), blind-spot scan, groupthink detection, ±0.5 confidence adj. 3,700-token output floor. ALL 9 score sets from Bands A+B Synthesized scores + integrity report + confidence adj ~30s

Critical path: ~113s (vs v2.1 measured 218s).

1.2 Seat Roster

# Band Role Model Vendor Tier Scores?
1 0 Research Agent Grok 4.5 xAI B No
2 A Primary Reviewer Claude Opus 5 Anthropic A Yes
3 A Cross-Check A DeepSeek V4 Flash DeepSeek C Yes
4 A Cross-Check B Gemini Pro Latest Google A Yes
5 A Cross-Check C DeepSeek V4 Pro DeepSeek C Yes
6 A Legal/Regulatory Claude Sonnet 5 Anthropic B Yes
7 B Financial Integrity MiniMax-M3 MiniMax B Yes
8 B Team/Founder Claude Fable 5 Anthropic A Yes
9 B Market-Reality Qwen3.7 Plus Alibaba B Yes
10 B Execution-Feasibility GPT-5.2 Pro OpenAI B Yes
11 C Synthesis & Integrity Gate Kimi K2.6 Moonshot AI A No

Vendor spread: 9 distinct vendors. Anthropic 3/11 (27.3%), DeepSeek 2/11 (18.2%), OpenAI 1/11 (9.1%), Google 1/11, Moonshot 1/11, Alibaba 1/11, MiniMax 1/11, xAI 1/11. Under 33% hard cap. Compliant.

Model double-ups: None. All 11 seats use distinct models. Score sets from 9 distinct models.

1.3 What Was Removed (and Where It Went)

v2.1 Phase Disposition
Phase 3 (Validation Reviewer) Deleted. Challenge function re-homed as adversarial head on Band C. Quantitative: flag any dimension where Primary >1.5σ from CC band mean as "contested." Stricter, zero-cost, zero-latency replacement for prose challenge notes.
Phase 11 (Audit Agent) Merged into Band C. Same model (Kimi K2.6) produces both synthesized scores AND integrity gate. Confidence adj (±0.5) applied as post-processing after score finalization - fence preserved. Output floor raised to 3,700 tokens.

1.4 Research Agent - Market Block

Research Agent output now includes a mandatory Market block (comps, TAM, competitive density, non-Western ecosystem data). Previously produced only factual brief + citations. This is load-bearing for Band B's Market-Reality seat - Research provides grounding; Market-Reality provides scoring judgment. Two distinct functions, same evidence base.


2. Tier Model Pools

2.1 Model Tiers

Tier Models Quality
Tier A (Premium) Claude Opus 5, Claude Fable 5, Gemini Pro Latest, Kimi K2.6 Best reasoning, highest accuracy
Tier B (Strong) Claude Sonnet 5, GPT-5.2 Pro, Qwen3.7 Plus, MiniMax-M3 Near-frontier at production cost
Tier C (Budget) DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4.5, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite Fast, cheap, acceptable for low-stakes

2.2 Tier Availability by Subscription

Tier Judge Pool Access Default Panel Use Case
Free Tier C + Research Agent (Tier B exception) 3 judges + Research Try-before-buy
Pro Tier B + Tier C 7 judges + Research Serious founder review
Enterprise Tier A + B + C 11 judges (4-band pipeline) Investor-grade, board-ready
White-Label Full pool, configurable 11 judges (configurable) Consultant-branded

Free tier Research exception: Grok 4.5 (Tier B) used even on Free. Only Tier B exception - Research Agent is single most impactful role for review quality and never scores (zero contamination). Free tier Grok fallback: Gemini 3.6 Flash (Tier C).

2.3 Fallback Chains - Cross-Vendor Enforced

Verified: 11/11 primary→F1 cross-vendor, 11/11 F1→F2 cross-vendor.

Role Primary Fallback 1 Fallback 2
Research Agent Grok 4.5 (xAI) DeepSeek V4 Flash (DeepSeek) Gemini Pro Latest (Google)
Primary Reviewer Claude Opus 5 (Anthropic) DeepSeek V4 Pro (DeepSeek) Kimi K2.6 (Moonshot)
Cross-Check A DeepSeek V4 Flash (DeepSeek) Gemini 3.6 Flash (Google) GPT-5.2 Pro (OpenAI)
Cross-Check B Gemini Pro Latest (Google) Kimi K2.6 (Moonshot) Claude Fable 5 (Anthropic)
Cross-Check C DeepSeek V4 Pro (DeepSeek) GPT-5.2 Pro (OpenAI) Gemini 3.6 Flash (Google)
Legal/Regulatory Claude Sonnet 5 (Anthropic) Qwen3.7 Plus (Alibaba) DeepSeek V4 Pro (DeepSeek)
Financial Integrity MiniMax-M3 (MiniMax) GPT-5.2 Pro (OpenAI) DeepSeek V4 Pro (DeepSeek)
Team/Founder Claude Fable 5 (Anthropic) MiniMax-M3 (MiniMax) Qwen3.7 Plus (Alibaba)
Market-Reality Qwen3.7 Plus (Alibaba) Grok 4.5 (xAI) Claude Fable 5 (Anthropic)
Execution-Feasibility GPT-5.2 Pro (OpenAI) Claude Sonnet 5 (Anthropic) MiniMax-M3 (MiniMax)
Synthesis & Integrity Gate Kimi K2.6 (Moonshot) Claude Fable 5 (Anthropic) Gemini Pro Latest (Google)

3. Pre-Assignment Rules

3.1 Hard Constraints

Rule Enforcement
No vendor >33% of active panel seats Counts seats, not distinct models. Anthropic 3/11 = 27.3% - compliant.
Research Agent never scores Blocked at assignment
All three cross-checks blind to each other Enforced by wall-clock - parallel dispatch in Band A makes cross-contamination physically impossible
Cross-checks see proposal + Research brief ONLY Data scope gated at prompt assembly. Primary is in same band but isolated.
Synthesis & Integrity Gate: scores + confidence adj as separate output sections Confidence adj applied post-processing, after score finalization - fence preserved in output structure
Fallback chains always cross-vendor Verified against §2.3
No model assigned to >1 seat Enforced at roster validation. All 11 seats use distinct models.

Vendor-cap correction (v2.2): v2.1 claimed "Anthropic 4/13 = 30.8%" but counted distinct models. Actual seat count was 5/13 = 38.5%, over the 33% hard cap. Root cause: Sonnet 5 and Fable 5 each appeared in 2 seats; the cap rule counts seats, not distinct models. Fixed by deleting Validation (1 Anthropic seat) and merging Audit into Synthesis (Fable 5 freed). All future vendor-cap audits must count panel seats, not distinct models.

3.2 Content-Based Rules

Proposal Characteristic Trigger Assignment Effect
Vertical = Biotech/Pharma Detected Legal/Regulatory +40%; Market-Reality pulls biotech calibration
Vertical = Hardware/IoT Detected Execution-Feasibility weights supply chain, manufacturing
Vertical = Fintech Detected Legal/Regulatory +30%; Financial Integrity +20%
Vertical = Climate/Energy Detected Market-Reality pulls energy data; Legal adds environmental law lens
Vertical = Superapp/Mini-program Detected Market-Reality +30% (non-US); Team/Founder +20%
Vertical = B2B Marketplace Detected Market-Reality +20%; Financial Integrity +20%
Vertical = Cross-Border/Export Detected Legal/Regulatory +30% (multi-jurisdiction); Market-Reality pulls trade data
Stage = Seed/Pre-seed Detected Team/Founder +40%; Market-Reality +30%; Financial Integrity -20%
Stage = Series A/B Detected Execution-Feasibility +30%; Team/Founder +20%
Stage = Growth/Late Detected Financial Integrity +30%; Legal/Regulatory +20%
Complexity = Low (<10 pages) Page count Reduced panel - see §3.3
Complexity = High (50+ pages) Page count Full panel + secondary Market & Financial re-pass (averaged)

Multiple triggers stack additively, capped at +50% per role.

3.3 Complexity-Based Panel Sizing

Complexity Free Pro Enterprise
Low (<10 pages) 2 + Research 5 + Research 8 judges (drop CC-C, Legal, Market-Reality)
Standard (10-50 pages) 3 + Research 7 + Research 11 (full pipeline)
High (50+ pages) Unsupported 7 + Research 11 + secondary Market & Financial pass

3.4 Missing Vertical Handling

Unmatched verticals use "Unclassified" with equal default weights. Review includes: "Vertical not recognized. Review used default weighting. For industry-specific calibration, contact us." No judge skipped. NOT "SaaS default."

3.5 Research Agent - Market Block

Research Agent output schema now includes mandatory Market block alongside factual brief and citations. Provides competitive landscape, TAM estimates, non-Western ecosystem data, density metrics. Band B's Market-Reality judge consumes as grounding and produces scoring judgment - two distinct cognitive operations on the same evidence base.


4. Runtime Logic

4.1 Scoring Algorithm

Aggregation formula:

Dimension_Score = SUM(judge_score × judge_weight) / SUM(judge_weight)

Where judge_weight starts at 1.0 and is modified by content-based adjustments from §3.2.

Band-stage processing:

  1. Band A: Five independent scores per dimension. Cross-Check scores averaged per dimension before Band B. Primary and Legal scores pass through individually.
  2. Band B: Four independent scores per dimension.
  3. Band C: Receives all 9 score sets. Produces synthesized dimension scores, adversarial challenge report (flag dimensions where Primary >1.5σ from CC band mean as "contested"), integrity gate report (contradictions, blind spots, groupthink flags), and confidence adjustment (±0.5 applied to final scores).

Total scoring judges: 9 (5 in Band A + 4 in Band B).

4.2 Availability Fallback

Trigger Action
Provider rate limit (429) Reassign to Fallback 1. Retry original after 60s.
Provider timeout (30s) Reassign immediately to Fallback 1. Log incident.
Malformed score Reassign to Fallback 1. F1 fail → Fallback 2. All three fail → manual review.
Empty response Treated as malformed. 256-token minimum for scoring judges. Kimi K2.6: 512 minimum; 3,700-token floor in Band C.
Two providers fail simultaneously Degrade: Enterprise → Pro panel (7 judges). Notify subscriber.
Cost ceiling 90% reached Swap Tier A → Tier B for Research only.
Band-level timeout (60s per band) Any band exceeding 60s triggers F1 for slowest seat. Parallel dispatch means only slowest seat governs band time.

4.3 Cost Optimization

Condition Action
Free tier Tier C + Grok 4.5 Research. Max 3 judges.
Pro, Complexity = Low Tier B for Primary/Cross-Checks, Tier C for others
Enterprise, Complexity = Low Reduced panel (8 judges). Tier A for Primary + 2 CCs, Tier B for others.
Enterprise, Complexity = High Full 11-judge Tier A/B panel + secondary passes
Shared-prefix caching Band A's 5 seats share identical prefix (~12,700 tokens). Cache TTL > review duration.

4.4 Latency Budgets

Measured baseline: 218s (v2.1 single-model datum). v2.2 critical path: 4 bands × ~28-33s = ~113s nominal.

Tier Bands Target Max Notes
Free 0, A(1 CC), C 90s 150s
Pro 0, A(2 CCs + Primary + Legal), B(2), C 120s 180s
Enterprise Full 4-band pipeline 120s 240s 4-band pipeline with 60s per-band timeout and ~1 full failover headroom
White-Label Configurable Configurable 240s

Empirical note: v2.1's 218s was measured on one model, not the assembled 13-seat pipeline. Real v2.1 pipeline estimated at 330-446s by two independent judges. v2.2's 4-band structure is defensible at 120s/240s target.

4.5 Known Model Limitations

Model Limitation Impact
GPT-5.6 family (Sol, Terra, Luna) reasoning_effort + function tools permanently incompatible on admin-ai Cannot be used as subagent judges. Entire family excluded from pool. Scoring-only possible if proposal text inlined in prompt, but not viable for pipeline dispatch.
Grok 4.5 Bare path grok-4.5 returned "no healthy deployments" Aug 12 2026 Use xai/grok-4.5 prefix path. Bare path may fail intermittently - always use prefix.
Anthropic credit wall All Anthropic models (Opus 5, Sonnet 5, Fable 5) fail simultaneously when API account balance depletes ~27% of panel fails together. Pre-baked failover roster in §4.7.
Gemini Pro Latest Floating alias - different scores on identical calls CC-B seat has higher noise. Pinned to dated snapshot before production.
Kimi K2.6 Low max_tokens → reasoning burn → empty output 512-token minimum; 3,700-token output floor in Band C. Fable 5 as Fallback 1.
Claude Fable 5 Bio/cyber safeguard routing may reject content Team/Founder only - not on Legal. Fallback: MiniMax-M3.
Perplexity Sonar family No tool support at all Excluded from pool.
Cohere Command-R family Tool result routing broken on second call Excluded from pool.

4.6 Pre-Pipeline Health Gate (MANDATORY)

Before starting any proposal validation, verify all 11 models respond to a minimal subagent dispatch test. This catches credit walls and deployment gaps before they kill seats mid-pipeline.

# Health gate dispatch  -  test all 11 models with a 5-second write-only task
for model in xai/grok-4.5 claude-opus-5 deepseek-v4-flash gemini/gemini-pro-latest \
            deepseek-v4-pro claude-sonnet-5 MiniMax-M3 claude-fable-5 \
            qwen3.7-plus gpt-5.2-pro kimi-k2.6; do
  hermes config set delegation.model $model
  delegate_task goal="Health check. Write {\"ok\":true} to /tmp/health-$model.json."
done

Pass: 11/11 models return valid JSON within 30s. Partial: <11/11. Disable dead seats. Proceed with reduced panel if ≥8/11. Fail: <8/11. Abort. Do not run pipeline. Investigate provider status.

4.7 Anthropic Credit Wall - Pre-Baked Failover

When the Anthropic API account balance depletes, ALL Anthropic models fail simultaneously (Opus 5, Sonnet 5, Fable 5). Apply this roster immediately without mid-session model hunting:

Dead Model Seat Fallback Model Vendor
Claude Opus 5 Primary Reviewer DeepSeek V4 Pro DeepSeek
Claude Sonnet 5 Legal/Regulatory Qwen3.7 Plus Alibaba
Claude Fable 5 Team/Founder MiniMax-M3 MiniMax

Post-failover vendor cap: Anthropic 0/11, DeepSeek 3/11 (27.3%), MiniMax 2/11 (18.2%). Still under 33% hard cap. Score integrity note: losing all Anthropic seats reduces panel depth. Flag all affected review results with "Anthropic credit wall - reduced panel" watermark.


5. Post-Review Feedback Loop

5.1 Outcome Tracking

Signal Collection Feeds Into
T+90 outcome delta Subscriber survey or public funding data Per-model, per-dimension accuracy
T+180 outcome delta Same Long-term weighting
T+365 outcome delta Same Retention/replacement decisions
Scoring bias >1.5σ from band mean across N=30 Vertical recusal
Tagging inconsistency Judge A vs B on same concept >30% of reviews Excluded from Primary rotation

5.2 Model Performance Tracking (min N=50)

Metric Calculation Threshold
Accuracy score Correlation: dimension score vs T+90 outcome <0.3 → de-weighted
Bias score Mean deviation from band mean per dimension per vertical >1.5σ → recusal flag
Consistency score Score variance across similar proposals >2.0σ → calibration review
Coverage score % of reviews where judge identified critical flaw others missed Tracked only

5.3 Model Rotation

Trigger Action
Accuracy <0.3 for 2 consecutive quarters (min N=50/quarter) Demote Tier A → Tier B evaluation
New model released by major provider Add to eval pool. 100 parallel reviews vs current roster. Promote if accuracy > current median + ≥50 reviews have T+90 data.
Model deprecated by provider Remove from all tiers. Replace with Fallback 1.
Bias >2.0σ on any vertical for 3+ consecutive reviews Immediate recusal. Notify operator.

5.4 Cold-Start Gate (First 90 Days)

  • Score consistency (band variance) is primary quality signal - no outcome data yet.
  • 3 malformed/empty scores in 24h → auto-replace with Fallback 1.

  • T+30 and T+60 subscriber surveys as interim feedback.
  • No model promotion, demotion, or rotation based on outcome metrics.

5.5 Build Validation Gate

v2.3 status: COMPLETE. Validated on 3 real proposals (2026-08-12) with 8 active judges (3 dead: CC-A, Execution, intermittent Research Agent). Results:

Proposal Panel Mean Opus 5 Solo Mean Delta Verdict
RFP Tank v1.0 4.40 4.93 -0.53 NO GO
VentureBuilt v2 6.14 6.10 +0.04 CONDITIONAL GO
CartMySupply 4.29 5.00 -0.71 NO GO
Aggregate 4.94 5.34 -0.40 THESIS FAIL

The "+0.5 point thesis" failed empirically. The panel does NOT inflate scores - it consistently scores LOWER than a solo Opus 5 because specialist judges find real structural problems that a generalist smooths over.

However, the panel caught 14 material errors the solo model missed or underweighted:

  • 3 revenue arithmetic errors (10×, 7×, and 3-conflicting-figure discrepancies)
  • 3 competitive mispositionings (CLEATUS at same price point, TeacherLists overlap, LivePlan Plan Review contradiction)
  • 3 execution infeasibilities (Amazon PA-API cart removal, contractor budget 4-7× underfunded, solo-dev 4-week wizard)
  • 2 legal blockers (COPPA exposure, charitable solicitation registration)
  • 2 team capacity impossibilities

Revised product thesis (v2.3): VerdictTank's value is error-detection density per dollar, not score elevation. A solo Opus 5 gives you a smooth, confident score. An 11-judge panel gives you the 14 things it was wrong about. The spread IS the product.

Gate outcome for v2.3: Proceed to build. Ablation test pending (compare quantitative σ-flagging vs deleted Validation Reviewer).


6. Pricing

6.1 Positioning

VerdictTank critiques proposals. AI authoring tools write them. Different categories. Closest comparable: professional proposal review services at $500-2,000/review. VerdictTank's 11-judge, 4-band pipeline aims for comparable depth at 10-20x lower cost.

6.2 Pricing Table

Monthly Annual (per month) Annual Total Savings
Pro $249/mo $208/mo $2,490/yr 16.7%
Enterprise $799/mo $666/mo $7,990/yr 16.7%
White-Label $1,499/mo $1,249/mo $14,990/yr 16.7%

Per-review option (no subscription): $49/review (Pro-equivalent: 7 judges + Research Agent). No outcome tracking, no feedback loop, no API.

6.3 Feature Matrix

Feature Free Pro Enterprise White-Label
Proposal reviews 1/mo 20/mo 100/mo Custom
Judge panel size 3 + Research 7 + Research 11 (4-band) Configurable
Model tier C + Research (B) B + C A + B + C Full pool
Per-dimension explanation - Yes Yes Yes
Fix-It action plan - Yes Yes Yes
URL-to-Review - Yes Yes Yes
Chat-to-Refine - Yes Yes Yes
Pre-Review Coach - Yes Yes Yes
Financial Integrity judge - - Yes Yes
Team/Founder Assessment - - Yes Yes
Legal/Regulatory check - - Yes Yes
Synthesis & Integrity Gate - - Yes Yes
Adversarial challenge report - - Yes Yes
Configurable judge pool - - Yes Yes
Outcome tracking - - Yes Yes
API access - - Yes Yes
White-label branding - - - Yes
Custom rubric - - - Yes
Custom judge pool - - - Yes
RFP Tank discount - - 20% off 30% off

6.4 Competitive Landscape

Competitor Pricing Category Notes
Professional proposal review (human) $500-2,000/review Manual critique Real comparable
AutogenAI ~$30k+/yr AI proposal authoring Writing, not critiquing
GC AI $500/seat/mo AI contract review (legal tech) Different vertical
Bidara $299-599/mo AI bid/proposal platform Writing-focused
AutoRFP.ai $899/mo AI RFP response authoring Writing, not critiquing

7. Product Boundaries

Feature VerdictTank RFP Tank
Upload proposal Yes Yes
Full 4-band pipeline Yes Yes (inherited)
10-dimension scoring Yes Yes (inherited)
Per-dimension explanation Yes Yes (inherited)
Fix-It action plan Yes Yes (inherited)
Adversarial challenge report Yes Yes (inherited)
Upload RFP No Yes
RFP compliance scoring No Yes
RFP requirement extraction No Yes

Rule: VerdictTank = "is this a good proposal?" RFP Tank = "does this match what they asked for?" RFP Tank inherits engine. Enterprise VerdictTank subscribers get 20% off RFP Tank.


8. Build Sequence (post-validation)

Step Deliverable Dependency
1 Orchestrator (band scheduler, dependency graph, parallel dispatcher, scoring algorithm) None
2 Research Agent + Band A (5-wide blind dispatch) 1
3 Band B (4-wide informed dispatch) 2
4 Band C (Synthesis & Integrity Gate) 3
5 Feedback loop + cold-start gate 4
6 Multi-tier pricing + subscription management 5