- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs - disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md - clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy) - projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture - proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review - docs/super-search/firecrawl-provider-strategy.md - updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io - .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
15 KiB
VerdictTank Pipeline Consolidation Recommendation
Lens: Reasoning / Logic Overlap Detection (Contradiction Detector)
Author: Kimi K2.6 (Reasoning-Verification specialist) Spec: judge-pool-spec.md v2.1 Target: ~15% reduction Date: 2026-08-12
Executive finding
ONE recommendation: Collapse Phase 10 (Reasoning-Verification) and Phase 11 (Audit Agent) into a single post-scoring synthesis seat: Synthesis and Integrity Gate.
These two roles ask the same meta-question with different labels:
| Role | Stated question | Actual job |
|---|---|---|
| Phase 10 Reasoning-Verification | Where do judges contradict? | Read all score sets, find disagreements, synthesize |
| Phase 11 Audit Agent | What did the panel miss / groupthink? | Read finished review, find blind spots / groupthink, adjust confidence |
Both are non-scoring meta-reviewers of the panel's own output. Neither sees the raw proposal as primary input in a unique way that the other does not eventually consume. Both produce a post-hoc integrity adjustment, not an independent domain score. Under a contradiction-detection lens, they are statistically and logically collinear seats.
1. Changes
Collapse
| Remove | Absorb into |
|---|---|
| Phase 10: Reasoning-Verification (standalone) | New Phase 10: Synthesis and Integrity Gate |
| Phase 11: Audit Agent (standalone) | (same merged seat) |
New merged role definition
| Field | Value |
|---|---|
| Phase | 10 (final) |
| Role | Synthesis and Integrity Gate |
| Model | Kimi K2.6 (Moonshot) — keeps contradiction specialty as primary |
| Fallback 1 | Claude Fable 5 (Anthropic) — former Audit primary |
| Fallback 2 | GPT-5.6 Sol (OpenAI) — former Reasoning F2 |
| Scores? | No (preserves non-scoring fence) |
| Input | ALL score sets from phases 2-9 + assembled draft review |
| Output | (a) Contradiction / alignment report, (b) synthesized dimension scores via section 4.1 formula, (c) blind-spot / groupthink notes, (d) single confidence adjustment +/-0.5 |
Prompt contract (must stay explicit)
Merged prompt has two ordered sections, not a vague do-both:
- Contradiction pass (former Phase 10): pairwise/cluster disagreement map across Primary, Validation, Cross-Check avg, Financial, Team, Market, Execution, Legal. Emit synthesized scores.
- Integrity pass (former Phase 11): given the synthesis above, scan for groupthink (panel sigma too low), missed domain blind spots, and apply one confidence adjustment (+/-0.5).
Single model, single call, single output schema. No second serial hop.
Spec section edits required
| Section | Edit |
|---|---|
| section 1.1 Pipeline Phases | Drop row 11; rewrite Phase 10 as Synthesis and Integrity Gate; phase count 11 to 10 |
| section 1.2 Enterprise Lineup | 13 seats to 11 seats (12 unique models still available in pool; one fewer active seat) |
| section 1.2 Model double-ups | Remove Fable 5: Team + Audit double-up. Fable becomes Reasoning/Audit fallback only (Team primary only) |
| section 2.3 Fallback table | Delete Audit Agent row; replace Reasoning-Verification row with merged role chain above |
| section 3.1 Hard constraints | Replace "Audit reads FINISHED review only" with "Synthesis and Integrity Gate is post-scoring, read-only, cannot open new domain scores" |
| section 3.3 Low-complexity Enterprise | Already drops Audit — update text to drop CC-C, Team, Legal (Audit no longer separate) |
| section 4.1 Scoring algorithm | Merge the two non-scoring bullets into one: Synthesis and Integrity Gate produces contradiction report + synthesized scores + confidence +/-0.5 |
| section 4.2 Cost ceiling rule | Swap Tier A to B for non-scoring roles first (Research, Synthesis) — drop separate Audit |
| section 4.4 Latency | Enterprise sequential phases lose one hop — recalibrate target downward slightly (e.g. 180s to ~165s aspirational; keep 300s max) |
| section 6.3 Feature matrix | Replace Audit Agent (Phase 11) row with Synthesis and Integrity Gate (includes audit) — still Enterprise/WL only if desired |
What this is NOT
- Not dropping Cross-Check C (that is diversity, not logic-overlap).
- Not merging Primary with Validation (Validation is deliberately contaminated by Primary; different question).
- Not merging Market with Execution (different input scopes and vertical triggers).
- Not touching Research (unique tool surface; never scores).
2. Reduction %
Seat / phase accounting (Enterprise standard panel)
| Metric | Before (v2.1) | After | Delta |
|---|---|---|---|
| Pipeline phases | 11 | 10 | -1 phase (-9.1%) |
| Active seats | 13 | 11 | -2 seats (-15.4%) |
| Serial post-score hops | 2 (P10 then P11) | 1 | -50% of meta-tail latency |
| Non-scoring roles | 3 (Research, Reasoning, Audit) | 2 (Research, Synthesis) | -33% meta overhead |
| Unique models in default lineup | 11 distinct (Sonnet x2, Fable x2) | 11 distinct (Sonnet still x2 Validation+Legal; Fable Team-only) | pool size unchanged |
| Anthropic share of active seats | 4/13 = 30.8% | 3/11 = 27.3% | more headroom under 33% cap |
Cost reduction (directional)
- Eliminate one full Tier-A (or Tier-A-fallback) completion on every Enterprise review.
- v2.1 cited about $1.30-2.10 per standard full review — expect ~12-18% COGS cut on the judge tail (one fewer expensive synthesis call), landing near ~15% all-in when weighted by token volume of P10+P11 vs earlier short scoring calls.
- Latency: remove one sequential dependency (~15-40s depending on model) from the critical path after phase 9.
Headline reduction: ~15% (seats -15.4%; cost band -12-18%; phases -9%).
3. What survives
Intact domain + scoring spine
1 Research Agent (Grok 4.5) - grounding, never scores
2 Primary Reviewer (Claude Opus 5) - 10-dimension critique
3 Validation Reviewer (Claude Sonnet 5) - challenges Primary
4a Cross-Check A (GPT-5.6 Sol) - blind re-score
4b Cross-Check B (Gemini Pro Latest) - blind re-score
4c Cross-Check C (DeepSeek V4 Pro) - blind re-score
5 Financial Integrity (MiniMax-M3) - cap table / unit econ
6 Team/Founder Assessment (Claude Fable 5) - founder-market fit
7 Market-Reality (Qwen3.7 Plus) - TAM / competitive / non-Western
8 Execution-Feasibility (GPT-5.6 Terra) - ops / timeline realism
9 Legal/Regulatory (Claude Sonnet 5) - multi-jurisdiction risk
10 Synthesis and Integrity Gate (Kimi K2.6) - contradictions + audit + +/-0.5
Preserved invariants
- Research never scores
- Three blind cross-checks, isolated from Primary/Validation
- Domain specialists (Financial, Team, Market, Execution, Legal) remain separate scoring questions
- Section 4.1 weighted aggregation formula unchanged
- Cross-vendor fallback rule unchanged (merged chain still cross-vendor at every hop)
- Vendor cap still satisfied (Anthropic drops to 27.3%)
- Build validation gate (section 5.5) unchanged — still require >=0.5 panel advantage
- Fenced post-score mode retained (merged seat cannot reopen domain scores)
Why these roles were NOT collapsed
| Pair | Why keep separate |
|---|---|
| Primary vs Validation | Validation is score-conditioned (sees Primary). Different epistemic job than independent re-score. |
| Validation vs Cross-Checks | CCs are blind; Validation is not. Collapsing destroys the blind control. |
| Cross-Check A/B/C | Same question by design — but vendor diversity is the product. Overlap is intentional ensemble, not accidental duplication. Do not collapse under this lens. |
| Market vs Execution | Different section 3.2 triggers, different evidence (competitive density vs ops timeline). |
| Financial vs Execution | Cap-table math is not shipping realism. |
| Legal vs Audit | Legal is domain-scoring on statutes; Audit was meta. Legal stays; Audit folds into Synthesis. |
| Team vs Audit (same model Fable) | Team scores founders; Audit was meta. After merge, Fable stops double-duty — removes same-model Team-to-Audit soft contamination path. |
4. Risks
| Risk | Severity | Mitigation |
|---|---|---|
| Single-call attention split — model skimps on contradiction map OR groupthink scan | Medium | Enforce two-section JSON schema; reject output missing either contradictions[] or integrity.confidence_delta. |
| Loss of independent second meta-reader — two weak meta-passes can catch what one strong pass misses | Medium | Keep dual-section prompt; optionally run Fable fallback as shadow audit on 10% of Enterprise reviews for 90 days (not on critical path). |
| K2.6 empty-output / reasoning-token burn (section 4.5) now single-points the entire meta-tail | High | Raise min tokens to 768-1024 for merged seat; on empty go immediate F1 (Fable). Circuit breaker already in section 4.2. |
| Confidence +/-0.5 applied by same model that synthesized scores — self-grading bias | Medium | Constrain: confidence delta may only move on documented groupthink (panel sigma under threshold) or explicit missed-claim list; ban free-form "I feel +/-0.5". |
| Pro tier marketing currently sells no Audit; matrix row rename may confuse | Low | Feature matrix: Synthesis and Integrity (Enterprise) — capability retained, seat count honest. |
| Low-complexity Enterprise drop list referenced Audit explicitly | Low | Update section 3.3 text only. |
| Outcome-tracking granularity — cannot attribute T+90 errors to Reasoning vs Audit separately | Low | Accept; meta-tail was always under-identified. Track merged seat as one accuracy series. |
| Over-collapse temptation — if 15% feels good, next PR kills CC-C or Validation | Process | This lens only authorizes logic-collinear non-scoring merges. Ensemble scorers are out of scope. |
Explicit non-risks
- Vendor diversity of scoring panel unchanged (still 8 vendors available; scoring seats still multi-lab).
- No new same-vendor Primary/meta self-audit (Kimi is not Anthropic Primary/Validation).
- Fable no longer audits after sitting Team — contamination risk decreases.
5. New diagram
Before (v2.1) — 11 phases / 13 seats
[Proposal]
|
v
+-----------------+
| 1 Research | Grok 4.5 (no score)
+--------+--------+
v
+-----------------+
| 2 Primary | Opus 5 (scores)
+--------+--------+
v
+-----------------+
| 3 Validation | Sonnet 5 (scores, sees Primary)
+--------+--------+
v
+----+----+
v v v
+----+ +----+ +----+
|4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores)
+--+-+ +--+-+ +--+-+
+-----+-----+
v
+-----------------+
| 5 Financial | MiniMax-M3
+--------+--------+
v
+-----------------+
| 6 Team/Founder | Fable 5
+--------+--------+
v
+-----------------+
| 7 Market | Qwen3.7 Plus
+--------+--------+
v
+-----------------+
| 8 Execution | Terra
+--------+--------+
v
+-----------------+
| 9 Legal | Sonnet 5
+--------+--------+
v
+-----------------+
|10 Reasoning | Kimi K2.6 (contradictions + synth)
+--------+--------+
v
+-----------------+
|11 Audit | Fable 5 (groupthink + +/-0.5)
+--------+--------+
v
[Final Report]
After (recommended) — 10 phases / 11 seats (~15% seat cut)
[Proposal]
|
v
+-----------------+
| 1 Research | Grok 4.5 (no score)
+--------+--------+
v
+-----------------+
| 2 Primary | Opus 5 (scores)
+--------+--------+
v
+-----------------+
| 3 Validation | Sonnet 5 (scores, sees Primary)
+--------+--------+
v
+----+----+
v v v
+----+ +----+ +----+
|4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores)
+--+-+ +--+-+ +--+-+
+-----+-----+
v
+-----------------+
| 5 Financial | MiniMax-M3
+--------+--------+
v
+-----------------+
| 6 Team/Founder | Fable 5 (Team only — no Audit double)
+--------+--------+
v
+-----------------+
| 7 Market | Qwen3.7 Plus
+--------+--------+
v
+-----------------+
| 8 Execution | Terra
+--------+--------+
v
+-----------------+
| 9 Legal | Sonnet 5
+--------+--------+
v
+-------------------------------------+
|10 Synthesis and Integrity Gate | Kimi K2.6
| - contradiction / alignment map |
| - section 4.1 synthesized scores |
| - groupthink / blind-spot notes |
| - confidence delta +/-0.5 (gated) |
| F1: Fable 5 · F2: Sol |
+------------------+------------------+
v
[Final Report]
Overlap evidence (why this pair, not another)
INPUT SCOPE
proposal scores finished-review
Primary ### . .
Validation ## ### .
Cross-Checks ### . . intentional ensemble (KEEP x3)
Financial ## ## .
Team ## ## .
Market ## ## .
Execution ## ## .
Legal ## # .
Reasoning (P10) . ### ## ]
Audit (P11) . ## ### ] SAME META QUESTION
Research tools . . (unique — KEEP)
Overlap score P10 intersect P11:
question similarity ~0.85 (contradiction equiv missed-agreement/groupthink)
input Jaccard ~0.80 (both post-domain, panel-output-primary)
output collinearity ~0.75 (both emit integrity adjustment, not domain scores)
-> COLLAPSE CANDIDATE under reasoning lens
Opus empirical echo (from v2.1 review results)
Opus 5 panel-vs-single test already showed cross-checks can be statistically tight (market dimension spread 0.00 across three CCs). That finding pressures ensemble scorer count — out of scope for this lens. The safer 15% cut is the duplicated meta-tail, not the diversity scorers the product sells. Killing Audit+Reasoning redundancy reduces COGS without touching the scoring majority (still 9 scoring judges + Research + 1 meta).
Scoring judges retained: 9 (Primary, Validation, 3xCC, Financial, Team, Market, Execution, Legal) Meta retained: 1 (merged) + Research
Decision line
Collapse Reasoning-Verification and Audit Agent into one Synthesis and Integrity Gate (Kimi K2.6). -2 seats (-15.4%), -1 serial phase, ~15% meta-tail cost, cleaner vendor math, removes Fable Team-to-Audit double-duty. Domain ensemble untouched. Ship as v2.2 delta on top of v2.1 blocker fixes.
Implementation checklist (for conductor)
- Patch judge-pool-spec.md sections listed in section 1
- Bump version header to v2.2; status remains FINAL DRAFT pending 3-5 proposal gate
- Update feature matrix + latency table
- Rewrite section 2.3 fallback row; delete Audit row
- Add output JSON schema for two-section merged seat
- Re-run vendor % assertion (expect 3/11 Anthropic)
- Do not change Pro 7-judge composition in this PR (Enterprise matrix row rename only)
- Optional: 90-day 10% shadow Fable integrity pass off critical path