# VerdictTank v2.1 — ONE Consolidation Recommendation ## Research Agent Grounding Lens (Grok 4.5) **Date:** 2026-08-12 **Author role:** Research Agent (only judge with native live web + social search) **Target:** ~15% reduction without losing effectiveness **Constraints satisfied:** no model >2 roles · Audit + Research survive · ≥2 cross-checks --- ## Research-lens thesis I produce the factual brief every downstream judge consumes. From that position, two seats add **latency and tokens without new signal**: 1. **Market-Reality (Phase 7)** mostly **re-summarizes my brief** — competitive density, TAM defensibility, and citation-backed market claims are already Research outputs. Its unique residual (non-Western lens) is a *prompt/schema* requirement, not a reason for a full serial phase. 2. **Cross-Check C as a third parallel scorer** is **highly correlated** with A/B. Internal Opus 5 empirical validation on this stack found all three cross-checks scored `market` at **exactly 6** (spread **0.00**). External ensemble guidance: if judges agree on everything, you bought one verdict three times (orq.ai LLM juries; Verga et al. "Judges → Juries"). Two diverse cross-checks capture the ensemble; a third correlated scorer is mostly cost. Specialist phases that do **not** rehash Research (Financial math, Team/Founder fit, Execution ops, Legal/regulatory code) stay. Audit and Reasoning stay as non-scoring synthesis/gates. --- ## (1) Changes — single consolidation: **Grounding-Redundancy Cut** ### A. Eliminate standalone Market-Reality (Phase 7) | Action | Detail | |---|---| | Remove | Phase 7 seat: Market-Reality / Qwen3.7 Plus (default) | | Absorb into Research | Expand Phase 1 brief schema with a mandatory **Market block**: competitive density table, TAM/SAM defensibility notes, ≥1 non-Western comps when vertical warrants, citation URLs, "unknown/unverifiable" flags | | Absorb residual scoring | Market *dimension* continues via Primary + Validation + 2 cross-checks (already score all 10 dimensions). No separate market narrative phase | | Vertical weights (§3.2) | Market-Reality weight boosts → apply to **Research brief depth** (more market citations) + **Primary/Validation market-dimension weight**, not a missing judge | ### B. Collapse cross-checks 3 → 2 (keep minimum) | Action | Detail | |---|---| | Remove | Phase 4c Cross-Check C as default Enterprise seat | | Keep | Phase 4a Cross-Check A (GPT-5.6 Sol) + Phase 4b Cross-Check B | | Diversity preserve | **Reassign DeepSeek V4 Pro → Cross-Check B** (replace Gemini Pro Latest in default lineup). Reasons: (i) keeps Chinese-lab / non-Western error surface that Market+CC-C previously carried; (ii) avoids Gemini floating-alias longitudinal noise documented in v2.1 §4.5 and review results; (iii) Gemini remains Fallback 1 for Research and available in pool | | Blindness | Both remaining CCs still see **proposal + Research brief ONLY** — unchanged isolation | ### C. What does *not* change - Research Agent never scores - Audit Agent remains fenced, post-scoring, read-only - Reasoning-Verification still synthesizes all score sets - Financial, Team/Founder, Execution, Legal stay as serial specialists - Fable double: Team + Audit (still 2) - No other model exceeds 2 roles ### D. Lineup delta (Enterprise default) | # | Phase | Role | Model (after) | Vendor | Change | |---|---|---|---|---|---| | 1 | 1 | Research Agent | Grok 4.5 | xAI | **Survives** — brief schema expanded | | 2 | 2 | Primary Reviewer | Claude Opus 5 | Anthropic | Unchanged | | 3 | 3 | Validation Reviewer | Claude Sonnet 5 | Anthropic | Unchanged | | 4 | 4a | Cross-Check A | GPT-5.6 Sol | OpenAI | Unchanged | | 5 | 4b | Cross-Check B | **DeepSeek V4 Pro** | DeepSeek | **Was Gemini; DeepSeek moved here** | | — | ~~4c~~ | ~~Cross-Check C~~ | — | — | **REMOVED** | | 6 | 5 | Financial Integrity | MiniMax-M3 | MiniMax | Unchanged | | 7 | 6 | Team/Founder | Claude Fable 5 | Anthropic | Unchanged | | — | ~~7~~ | ~~Market-Reality~~ | — | — | **REMOVED** (folded into Research) | | 8 | 7' | Execution-Feasibility | GPT-5.6 Terra | OpenAI | Unchanged | | 9 | 8' | Legal/Regulatory | **Qwen3.7 Plus** | Alibaba | **Vendor-cap patch** (was Sonnet) | | 10 | 9' | Reasoning-Verification | Kimi K2.6 | Moonshot | Unchanged | | 11 | 10' | Audit Agent | Claude Fable 5 | Anthropic | **Survives** | **Seats:** 13 → **11** **Phases:** 11 → **10** (4a–4b parallel; Market gone) **Scoring judges:** 10 → **8** (Primary, Validation, CC-A, CC-B, Financial, Team, Execution, Legal) **Non-scoring:** Research, Reasoning, Audit (3) **Total:** 8 + 3 = **11** Qwen3.7 Plus remains on the default path via Legal (not dropped from Enterprise). Gemini stays in pool + Research fallback. ### E. Fallback table patches (must ship with this cut) - Delete Market-Reality row **or** keep as optional on-demand role for White-Label only - Cross-Check B primary becomes DeepSeek V4 Pro; F1 Gemini Pro Latest; F2 Claude Fable 5 (still cross-vendor) - Cross-Check C row retired from Enterprise default (pool may retain for WL custom panels) - Legal primary: Qwen3.7 Plus; F1 Gemini Pro Latest; F2 DeepSeek V4 Pro (cross-vendor) - Research F2 can stay Terra; optional: add Qwen as Research F3 for market-heavy verticals ### F. Research brief schema addition (contract) ``` ## Market block (mandatory) - Competitors: name, category (authoring vs critique vs other), price anchor, source URL - Non-Western / non-US comps: min 1 when vertical in {superapp, cross-border, unclassified-global}; else "N/A — US-centric vertical" - TAM/SAM claims: verified | overstated | unverifiable — with citation or explicit gap - Density judgment: sparse | contested | saturated — one paragraph, citations only ``` Downstream judges **cite the Market block ID**, not re-crawl the open web (except Legal/Financial specialists on their narrow facts). --- ## (2) Reduction % | Metric | Before (v2.1) | After | Delta | |---|---|---|---| | Seats (Enterprise default) | 13 | 11 | **−15.4%** | | Phases | 11 | 10 | −9.1% | | Parallel CC calls | 3 | 2 | −33% of CC fan-out | | Scoring judges | 10 | 8 | −20% | | Serial specialist after scores | Market + Exec + Legal… | Exec + Legal… | −1 serial hop | | Est. E2E latency (vs 218s baseline) | 218s | ~185–195s | **~11–15%** wall-clock | | Est. token/COGS share | full panel | −Market narrative −1 full score JSON | **~14–18%** cost | **Headline reduction: ~15%** (seat count exact at −15.4%; cost/latency band 14–18% / 11–15%). Why not larger: cutting Audit or a second CC would violate constraints or erase the disagreement signal that makes panels worth running (selection-bottleneck literature: judge-based selection > synthesis; arxiv 2603.20324). --- ## (3) What survives (effectiveness preserved) | Kept | Why it still works | |---|---| | **Research Agent** | Stronger, not weaker — owns market grounding explicitly | | **Audit Agent** | Groupthink / blind-spot gate unchanged | | **≥2 cross-checks** | Sol + DeepSeek: US frontier + non-Western lab; blind; proposal+brief only | | **Primary + Validation** | Dual Anthropic pass with sober second read (different tiers A/B) | | **Financial, Team, Execution, Legal** | True specialists; low overlap with Research crawl | | **Reasoning-Verification** | Contradiction synthesis over fewer, higher-signal score sets | | **Vendor cap** | After Legal→Qwen: Anthropic 3/11 = **27.3%** (under 33%) | | **Pool diversity** | Gemini remains in pool/fallbacks; White-Label can still pin full custom panels | | **Model role cap** | Fable Team+Audit = 2; all others ≤1 | --- ## (4) Risks | Risk | Severity | Mitigation | |---|---|---| | Market nuance loss on superapp / cross-border verticals | Med | Research Market block mandatory; §3.2 superapp trigger increases Research market citation quota + Primary market-dimension weight +30% | | DeepSeek as sole non-Western CC may under-challenge US-centric Primary | Med | Legal→Qwen adds second non-Western seat; Audit watches for US-default groupthink | | Anthropic share 4/11 if Legal stays Sonnet | High (rule break) | **Must** reassign Legal → Qwen (included in this same cut) | | Sol still cannot tool-use on admin-ai | Low | CC-A is scoring-only JSON — already OK per §4.5 | | Dropping third CC reduces disagreement surface | Med-Low | Literature + internal 0.00 market spread say third was fake diversity; monitor panel sigma for 30 days; if sigma collapses, restore CC-C on High complexity only (§3.3 already expands high-complexity) | | Research brief becomes single point of market failure | Med | Research fallback chain unchanged (Gemini grounding → Terra); malformed Research → subscriber notify + degrade already specified | | Pro tier "7 judges" matrix needs rewrite | Low | Define Pro as: Research + Primary + Validation + 2 CC + Execution + Reasoning (no Financial/Team/Legal/Audit) — document explicitly | | Build-gate still unproven (panel delta <0.5 vs solo) | Existential (pre-existing) | This cut helps the thesis: fewer correlated seats make a ≥0.5 panel advantage more plausible if one exists; still run 3–5 real proposals before build | --- ## (5) New diagram ### Before (v2.1) — 11 phases / 13 seats ``` ┌─────────────────────┐ │ 1 Research (Grok) │ grounding brief └──────────┬──────────┘ ▼ ┌─────────────────────┐ │ 2 Primary (Opus 5) │ scores └──────────┬──────────┘ ▼ ┌─────────────────────┐ │ 3 Validation (Sonnet)│ challenge └──────────┬──────────┘ ▼ ┌────────────────┼────────────────┐ ▼ ▼ ▼ ┌──────────┐ ┌──────────┐ ┌──────────┐ │4a CC-A │ │4b CC-B │ │4c CC-C │ blind re-scores │ Sol │ │ Gemini │ │ DeepSeek │ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────────────────┼────────────────┘ ▼ ┌─────────────────────┐ │ 5 Financial (M3) │ └──────────┬──────────┘ ▼ ┌─────────────────────┐ │ 6 Team (Fable) │ └──────────┬──────────┘ ▼ ┌─────────────────────┐ │ 7 Market (Qwen) │ ← re-summarizes Research └──────────┬──────────┘ ▼ ┌─────────────────────┐ │ 8 Execution (Terra) │ └──────────┬──────────┘ ▼ ┌─────────────────────┐ │ 9 Legal (Sonnet) │ └──────────┬──────────┘ ▼ ┌─────────────────────┐ │10 Reasoning (Kimi) │ synthesize └──────────┬──────────┘ ▼ ┌─────────────────────┐ │11 Audit (Fable) │ fenced ±0.5 └─────────────────────┘ ``` ### After — Grounding-Redundancy Cut — 10 phases / 11 seats (−15.4%) ``` ┌──────────────────────────────────┐ │ 1 Research (Grok 4.5) │ │ + mandatory Market block │ LIVE WEB (only) └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │ 2 Primary (Opus 5) │ └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │ 3 Validation (Sonnet 5) │ └────────────────┬─────────────────┘ ▼ ┌───────────┴───────────┐ ▼ ▼ ┌─────────────┐ ┌─────────────┐ │4a CC-A Sol │ │4b CC-B │ 2 blind CCs │ │ │ DeepSeek │ (min met) └──────┬──────┘ └──────┬──────┘ └───────────┬───────────┘ ▼ ┌──────────────────────────────────┐ │ 5 Financial (MiniMax-M3) │ └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │ 6 Team/Founder (Fable 5) │ └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │ 7 Execution (Terra) │ Market phase GONE └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │ 8 Legal (Qwen3.7 Plus) │ vendor-cap patch └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │ 9 Reasoning-Verification (Kimi) │ └────────────────┬─────────────────┘ ▼ ┌──────────────────────────────────┐ │10 Audit (Fable 5) │ SURVIVES └──────────────────────────────────┘ ``` ### Signal flow (Research lens) ``` Research brief ├─ facts/citations ──────────────► all scorers (unchanged) ├─ Market block (NEW) ───────────► Primary / Validation / CCs │ (no Phase-7 rewrite) └─ gaps/unverifiable ────────────► Audit watches for overclaim Dropped edges (were low-signal): Research ══re-summary══► Market narrative ══► Reasoning CC-C ══correlate≈0══► CC average ``` --- ## Constraint checklist | Constraint | Status | |---|---| | ~15% reduction | **15.4% seats**; ~14–18% COGS; ~11–15% latency | | No model >2 roles | Fable Team+Audit=2; all others ≤1 | | Audit survives | Phase 10' Fable, fenced | | Research survives | Phase 1 Grok, expanded | | Minimum 2 cross-checks | Sol + DeepSeek | | Vendor ≤33% | After Legal→Qwen: Anthropic 3/11 = 27.3% | --- ## Implementation note (one PR) 1. Spec §1.1 / §1.2 / §2.3 / §3.2 / §3.3 / §4.1 / §6.3 — role counts must all read **11 seats / 10 phases** 2. Cross-section reconciliation (lesson from v2.0→v2.1 role expansion) before any re-review 3. Cache-bust deploy `judge-pool-spec.md?v=2.2` 4. Do **not** build until 3–5 real proposals still clear the ≥0.5 panel-vs-solo gate on the *reduced* panel --- ## Bottom line **ONE consolidation:** *Grounding-Redundancy Cut* — delete Market-Reality as a serial phase (it re-summarizes Research) and delete the third cross-check (empirically zero incremental spread); fold market into the Research brief; keep DeepSeek as CC-B; move Legal to Qwen for vendor-cap compliance. **−15.4% seats, Research + Audit intact, 2 blind cross-checks retained, specialists preserved.**