# VerdictTank Pipeline Consolidation Recommendation ## Lens: Reasoning / Logic Overlap Detection (Contradiction Detector) **Author:** Kimi K2.6 (Reasoning-Verification specialist) **Spec:** judge-pool-spec.md v2.1 **Target:** ~15% reduction **Date:** 2026-08-12 --- ## Executive finding **ONE recommendation:** Collapse **Phase 10 (Reasoning-Verification)** and **Phase 11 (Audit Agent)** into a single post-scoring synthesis seat: **Synthesis and Integrity Gate**. These two roles ask the same meta-question with different labels: | Role | Stated question | Actual job | |---|---|---| | Phase 10 Reasoning-Verification | Where do judges contradict? | Read all score sets, find disagreements, synthesize | | Phase 11 Audit Agent | What did the panel miss / groupthink? | Read finished review, find blind spots / groupthink, adjust confidence | Both are **non-scoring meta-reviewers of the panel's own output**. Neither sees the raw proposal as primary input in a unique way that the other does not eventually consume. Both produce a **post-hoc integrity adjustment**, not an independent domain score. Under a contradiction-detection lens, they are statistically and logically collinear seats. --- ## 1. Changes ### Collapse | Remove | Absorb into | |---|---| | Phase 10: Reasoning-Verification (standalone) | **New Phase 10: Synthesis and Integrity Gate** | | Phase 11: Audit Agent (standalone) | (same merged seat) | ### New merged role definition | Field | Value | |---|---| | **Phase** | 10 (final) | | **Role** | **Synthesis and Integrity Gate** | | **Model** | Kimi K2.6 (Moonshot) — keeps contradiction specialty as primary | | **Fallback 1** | Claude Fable 5 (Anthropic) — former Audit primary | | **Fallback 2** | GPT-5.6 Sol (OpenAI) — former Reasoning F2 | | **Scores?** | No (preserves non-scoring fence) | | **Input** | ALL score sets from phases 2-9 + assembled draft review | | **Output** | (a) Contradiction / alignment report, (b) synthesized dimension scores via section 4.1 formula, (c) blind-spot / groupthink notes, (d) single confidence adjustment +/-0.5 | ### Prompt contract (must stay explicit) Merged prompt has **two ordered sections**, not a vague do-both: 1. **Contradiction pass** (former Phase 10): pairwise/cluster disagreement map across Primary, Validation, Cross-Check avg, Financial, Team, Market, Execution, Legal. Emit synthesized scores. 2. **Integrity pass** (former Phase 11): given the synthesis above, scan for groupthink (panel sigma too low), missed domain blind spots, and apply **one** confidence adjustment (+/-0.5). Single model, single call, single output schema. No second serial hop. ### Spec section edits required | Section | Edit | |---|---| | section 1.1 Pipeline Phases | Drop row 11; rewrite Phase 10 as Synthesis and Integrity Gate; phase count **11 to 10** | | section 1.2 Enterprise Lineup | **13 seats to 11 seats** (12 unique models still available in pool; one fewer active seat) | | section 1.2 Model double-ups | Remove Fable 5: Team + Audit double-up. Fable becomes Reasoning/Audit **fallback only** (Team primary only) | | section 2.3 Fallback table | Delete Audit Agent row; replace Reasoning-Verification row with merged role chain above | | section 3.1 Hard constraints | Replace "Audit reads FINISHED review only" with "Synthesis and Integrity Gate is post-scoring, read-only, cannot open new domain scores" | | section 3.3 Low-complexity Enterprise | Already drops Audit — update text to drop CC-C, Team, Legal (Audit no longer separate) | | section 4.1 Scoring algorithm | Merge the two non-scoring bullets into one: Synthesis and Integrity Gate produces contradiction report + synthesized scores + confidence +/-0.5 | | section 4.2 Cost ceiling rule | Swap Tier A to B for non-scoring roles first (Research, Synthesis) — drop separate Audit | | section 4.4 Latency | Enterprise sequential phases lose one hop — recalibrate target downward slightly (e.g. 180s to ~165s aspirational; keep 300s max) | | section 6.3 Feature matrix | Replace Audit Agent (Phase 11) row with Synthesis and Integrity Gate (includes audit) — still Enterprise/WL only if desired | ### What this is NOT - Not dropping Cross-Check C (that is diversity, not logic-overlap). - Not merging Primary with Validation (Validation is deliberately contaminated by Primary; different question). - Not merging Market with Execution (different input scopes and vertical triggers). - Not touching Research (unique tool surface; never scores). --- ## 2. Reduction % ### Seat / phase accounting (Enterprise standard panel) | Metric | Before (v2.1) | After | Delta | |---|---|---|---| | Pipeline phases | 11 | **10** | -1 phase (-9.1%) | | Active seats | 13 | **11** | -2 seats (-15.4%) | | Serial post-score hops | 2 (P10 then P11) | **1** | -50% of meta-tail latency | | Non-scoring roles | 3 (Research, Reasoning, Audit) | **2** (Research, Synthesis) | -33% meta overhead | | Unique models in default lineup | 11 distinct (Sonnet x2, Fable x2) | **11 distinct** (Sonnet still x2 Validation+Legal; Fable Team-only) | pool size unchanged | | Anthropic share of active seats | 4/13 = 30.8% | **3/11 = 27.3%** | more headroom under 33% cap | ### Cost reduction (directional) - Eliminate one full Tier-A (or Tier-A-fallback) completion on every Enterprise review. - v2.1 cited about $1.30-2.10 per standard full review — expect **~12-18% COGS cut** on the judge tail (one fewer expensive synthesis call), landing near **~15%** all-in when weighted by token volume of P10+P11 vs earlier short scoring calls. - Latency: remove one sequential dependency (~15-40s depending on model) from the critical path after phase 9. **Headline reduction: ~15%** (seats -15.4%; cost band -12-18%; phases -9%). --- ## 3. What survives ### Intact domain + scoring spine ``` 1 Research Agent (Grok 4.5) - grounding, never scores 2 Primary Reviewer (Claude Opus 5) - 10-dimension critique 3 Validation Reviewer (Claude Sonnet 5) - challenges Primary 4a Cross-Check A (GPT-5.6 Sol) - blind re-score 4b Cross-Check B (Gemini Pro Latest) - blind re-score 4c Cross-Check C (DeepSeek V4 Pro) - blind re-score 5 Financial Integrity (MiniMax-M3) - cap table / unit econ 6 Team/Founder Assessment (Claude Fable 5) - founder-market fit 7 Market-Reality (Qwen3.7 Plus) - TAM / competitive / non-Western 8 Execution-Feasibility (GPT-5.6 Terra) - ops / timeline realism 9 Legal/Regulatory (Claude Sonnet 5) - multi-jurisdiction risk 10 Synthesis and Integrity Gate (Kimi K2.6) - contradictions + audit + +/-0.5 ``` ### Preserved invariants - Research never scores - Three blind cross-checks, isolated from Primary/Validation - Domain specialists (Financial, Team, Market, Execution, Legal) remain **separate scoring questions** - Section 4.1 weighted aggregation formula unchanged - Cross-vendor fallback rule unchanged (merged chain still cross-vendor at every hop) - Vendor cap still satisfied (Anthropic drops to 27.3%) - Build validation gate (section 5.5) unchanged — still require >=0.5 panel advantage - Fenced post-score mode retained (merged seat cannot reopen domain scores) ### Why these roles were NOT collapsed | Pair | Why keep separate | |---|---| | Primary vs Validation | Validation is **score-conditioned** (sees Primary). Different epistemic job than independent re-score. | | Validation vs Cross-Checks | CCs are blind; Validation is not. Collapsing destroys the blind control. | | Cross-Check A/B/C | Same question by design — but **vendor diversity is the product**. Overlap is intentional ensemble, not accidental duplication. Do not collapse under this lens. | | Market vs Execution | Different section 3.2 triggers, different evidence (competitive density vs ops timeline). | | Financial vs Execution | Cap-table math is not shipping realism. | | Legal vs Audit | Legal is domain-scoring on statutes; Audit was meta. Legal stays; Audit folds into Synthesis. | | Team vs Audit (same model Fable) | Team scores founders; Audit was meta. After merge, Fable stops double-duty — **removes same-model Team-to-Audit soft contamination path**. | --- ## 4. Risks | Risk | Severity | Mitigation | |---|---|---| | **Single-call attention split** — model skimps on contradiction map OR groupthink scan | Medium | Enforce two-section JSON schema; reject output missing either contradictions[] or integrity.confidence_delta. | | **Loss of independent second meta-reader** — two weak meta-passes can catch what one strong pass misses | Medium | Keep dual-section prompt; optionally run Fable fallback as **shadow audit** on 10% of Enterprise reviews for 90 days (not on critical path). | | **K2.6 empty-output / reasoning-token burn** (section 4.5) now single-points the entire meta-tail | High | Raise min tokens to 768-1024 for merged seat; on empty go immediate F1 (Fable). Circuit breaker already in section 4.2. | | **Confidence +/-0.5 applied by same model that synthesized scores** — self-grading bias | Medium | Constrain: confidence delta may only move on documented groupthink (panel sigma under threshold) or explicit missed-claim list; ban free-form "I feel +/-0.5". | | **Pro tier marketing** currently sells no Audit; matrix row rename may confuse | Low | Feature matrix: Synthesis and Integrity (Enterprise) — capability retained, seat count honest. | | **Low-complexity Enterprise drop list** referenced Audit explicitly | Low | Update section 3.3 text only. | | **Outcome-tracking granularity** — cannot attribute T+90 errors to Reasoning vs Audit separately | Low | Accept; meta-tail was always under-identified. Track merged seat as one accuracy series. | | **Over-collapse temptation** — if 15% feels good, next PR kills CC-C or Validation | Process | This lens only authorizes **logic-collinear non-scoring** merges. Ensemble scorers are out of scope. | ### Explicit non-risks - Vendor diversity of **scoring** panel unchanged (still 8 vendors available; scoring seats still multi-lab). - No new same-vendor Primary/meta self-audit (Kimi is not Anthropic Primary/Validation). - Fable no longer audits after sitting Team — **contamination risk decreases**. --- ## 5. New diagram ### Before (v2.1) — 11 phases / 13 seats ``` [Proposal] | v +-----------------+ | 1 Research | Grok 4.5 (no score) +--------+--------+ v +-----------------+ | 2 Primary | Opus 5 (scores) +--------+--------+ v +-----------------+ | 3 Validation | Sonnet 5 (scores, sees Primary) +--------+--------+ v +----+----+ v v v +----+ +----+ +----+ |4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores) +--+-+ +--+-+ +--+-+ +-----+-----+ v +-----------------+ | 5 Financial | MiniMax-M3 +--------+--------+ v +-----------------+ | 6 Team/Founder | Fable 5 +--------+--------+ v +-----------------+ | 7 Market | Qwen3.7 Plus +--------+--------+ v +-----------------+ | 8 Execution | Terra +--------+--------+ v +-----------------+ | 9 Legal | Sonnet 5 +--------+--------+ v +-----------------+ |10 Reasoning | Kimi K2.6 (contradictions + synth) +--------+--------+ v +-----------------+ |11 Audit | Fable 5 (groupthink + +/-0.5) +--------+--------+ v [Final Report] ``` ### After (recommended) — 10 phases / 11 seats (~15% seat cut) ``` [Proposal] | v +-----------------+ | 1 Research | Grok 4.5 (no score) +--------+--------+ v +-----------------+ | 2 Primary | Opus 5 (scores) +--------+--------+ v +-----------------+ | 3 Validation | Sonnet 5 (scores, sees Primary) +--------+--------+ v +----+----+ v v v +----+ +----+ +----+ |4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores) +--+-+ +--+-+ +--+-+ +-----+-----+ v +-----------------+ | 5 Financial | MiniMax-M3 +--------+--------+ v +-----------------+ | 6 Team/Founder | Fable 5 (Team only — no Audit double) +--------+--------+ v +-----------------+ | 7 Market | Qwen3.7 Plus +--------+--------+ v +-----------------+ | 8 Execution | Terra +--------+--------+ v +-----------------+ | 9 Legal | Sonnet 5 +--------+--------+ v +-------------------------------------+ |10 Synthesis and Integrity Gate | Kimi K2.6 | - contradiction / alignment map | | - section 4.1 synthesized scores | | - groupthink / blind-spot notes | | - confidence delta +/-0.5 (gated) | | F1: Fable 5 · F2: Sol | +------------------+------------------+ v [Final Report] ``` ### Overlap evidence (why this pair, not another) ``` INPUT SCOPE proposal scores finished-review Primary ### . . Validation ## ### . Cross-Checks ### . . intentional ensemble (KEEP x3) Financial ## ## . Team ## ## . Market ## ## . Execution ## ## . Legal ## # . Reasoning (P10) . ### ## ] Audit (P11) . ## ### ] SAME META QUESTION Research tools . . (unique — KEEP) Overlap score P10 intersect P11: question similarity ~0.85 (contradiction equiv missed-agreement/groupthink) input Jaccard ~0.80 (both post-domain, panel-output-primary) output collinearity ~0.75 (both emit integrity adjustment, not domain scores) -> COLLAPSE CANDIDATE under reasoning lens ``` ### Opus empirical echo (from v2.1 review results) Opus 5 panel-vs-single test already showed cross-checks can be **statistically tight** (market dimension spread 0.00 across three CCs). That finding pressures **ensemble scorer** count — out of scope for *this* lens. The safer 15% cut is the **duplicated meta-tail**, not the diversity scorers the product sells. Killing Audit+Reasoning redundancy reduces COGS without touching the scoring majority (still 9 scoring judges + Research + 1 meta). **Scoring judges retained: 9** (Primary, Validation, 3xCC, Financial, Team, Market, Execution, Legal) **Meta retained: 1** (merged) + Research --- ## Decision line > **Collapse Reasoning-Verification and Audit Agent into one Synthesis and Integrity Gate (Kimi K2.6).** > **-2 seats (-15.4%), -1 serial phase, ~15% meta-tail cost, cleaner vendor math, removes Fable Team-to-Audit double-duty.** > Domain ensemble untouched. Ship as v2.2 delta on top of v2.1 blocker fixes. --- ## Implementation checklist (for conductor) 1. Patch judge-pool-spec.md sections listed in section 1 2. Bump version header to v2.2; status remains FINAL DRAFT pending 3-5 proposal gate 3. Update feature matrix + latency table 4. Rewrite section 2.3 fallback row; delete Audit row 5. Add output JSON schema for two-section merged seat 6. Re-run vendor % assertion (expect 3/11 Anthropic) 7. Do **not** change Pro 7-judge composition in this PR (Enterprise matrix row rename only) 8. Optional: 90-day 10% shadow Fable integrity pass off critical path