- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs - disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md - clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy) - projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture - proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review - docs/super-search/firecrawl-provider-strategy.md - updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io - .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
18 KiB
VerdictTank v2.1 — ONE Consolidation Recommendation
Research Agent Grounding Lens (Grok 4.5)
Date: 2026-08-12
Author role: Research Agent (only judge with native live web + social search)
Target: ~15% reduction without losing effectiveness
Constraints satisfied: no model >2 roles · Audit + Research survive · ≥2 cross-checks
Research-lens thesis
I produce the factual brief every downstream judge consumes. From that position, two seats add latency and tokens without new signal:
- Market-Reality (Phase 7) mostly re-summarizes my brief — competitive density, TAM defensibility, and citation-backed market claims are already Research outputs. Its unique residual (non-Western lens) is a prompt/schema requirement, not a reason for a full serial phase.
- Cross-Check C as a third parallel scorer is highly correlated with A/B. Internal Opus 5 empirical validation on this stack found all three cross-checks scored
marketat exactly 6 (spread 0.00). External ensemble guidance: if judges agree on everything, you bought one verdict three times (orq.ai LLM juries; Verga et al. "Judges → Juries"). Two diverse cross-checks capture the ensemble; a third correlated scorer is mostly cost.
Specialist phases that do not rehash Research (Financial math, Team/Founder fit, Execution ops, Legal/regulatory code) stay. Audit and Reasoning stay as non-scoring synthesis/gates.
(1) Changes — single consolidation: Grounding-Redundancy Cut
A. Eliminate standalone Market-Reality (Phase 7)
| Action | Detail |
|---|---|
| Remove | Phase 7 seat: Market-Reality / Qwen3.7 Plus (default) |
| Absorb into Research | Expand Phase 1 brief schema with a mandatory Market block: competitive density table, TAM/SAM defensibility notes, ≥1 non-Western comps when vertical warrants, citation URLs, "unknown/unverifiable" flags |
| Absorb residual scoring | Market dimension continues via Primary + Validation + 2 cross-checks (already score all 10 dimensions). No separate market narrative phase |
| Vertical weights (§3.2) | Market-Reality weight boosts → apply to Research brief depth (more market citations) + Primary/Validation market-dimension weight, not a missing judge |
B. Collapse cross-checks 3 → 2 (keep minimum)
| Action | Detail |
|---|---|
| Remove | Phase 4c Cross-Check C as default Enterprise seat |
| Keep | Phase 4a Cross-Check A (GPT-5.6 Sol) + Phase 4b Cross-Check B |
| Diversity preserve | Reassign DeepSeek V4 Pro → Cross-Check B (replace Gemini Pro Latest in default lineup). Reasons: (i) keeps Chinese-lab / non-Western error surface that Market+CC-C previously carried; (ii) avoids Gemini floating-alias longitudinal noise documented in v2.1 §4.5 and review results; (iii) Gemini remains Fallback 1 for Research and available in pool |
| Blindness | Both remaining CCs still see proposal + Research brief ONLY — unchanged isolation |
C. What does not change
- Research Agent never scores
- Audit Agent remains fenced, post-scoring, read-only
- Reasoning-Verification still synthesizes all score sets
- Financial, Team/Founder, Execution, Legal stay as serial specialists
- Fable double: Team + Audit (still 2)
- No other model exceeds 2 roles
D. Lineup delta (Enterprise default)
| # | Phase | Role | Model (after) | Vendor | Change |
|---|---|---|---|---|---|
| 1 | 1 | Research Agent | Grok 4.5 | xAI | Survives — brief schema expanded |
| 2 | 2 | Primary Reviewer | Claude Opus 5 | Anthropic | Unchanged |
| 3 | 3 | Validation Reviewer | Claude Sonnet 5 | Anthropic | Unchanged |
| 4 | 4a | Cross-Check A | GPT-5.6 Sol | OpenAI | Unchanged |
| 5 | 4b | Cross-Check B | DeepSeek V4 Pro | DeepSeek | Was Gemini; DeepSeek moved here |
| — | — | — | REMOVED | ||
| 6 | 5 | Financial Integrity | MiniMax-M3 | MiniMax | Unchanged |
| 7 | 6 | Team/Founder | Claude Fable 5 | Anthropic | Unchanged |
| — | — | — | REMOVED (folded into Research) | ||
| 8 | 7' | Execution-Feasibility | GPT-5.6 Terra | OpenAI | Unchanged |
| 9 | 8' | Legal/Regulatory | Qwen3.7 Plus | Alibaba | Vendor-cap patch (was Sonnet) |
| 10 | 9' | Reasoning-Verification | Kimi K2.6 | Moonshot | Unchanged |
| 11 | 10' | Audit Agent | Claude Fable 5 | Anthropic | Survives |
Seats: 13 → 11
Phases: 11 → 10 (4a–4b parallel; Market gone)
Scoring judges: 10 → 8 (Primary, Validation, CC-A, CC-B, Financial, Team, Execution, Legal)
Non-scoring: Research, Reasoning, Audit (3)
Total: 8 + 3 = 11
Qwen3.7 Plus remains on the default path via Legal (not dropped from Enterprise). Gemini stays in pool + Research fallback.
E. Fallback table patches (must ship with this cut)
- Delete Market-Reality row or keep as optional on-demand role for White-Label only
- Cross-Check B primary becomes DeepSeek V4 Pro; F1 Gemini Pro Latest; F2 Claude Fable 5 (still cross-vendor)
- Cross-Check C row retired from Enterprise default (pool may retain for WL custom panels)
- Legal primary: Qwen3.7 Plus; F1 Gemini Pro Latest; F2 DeepSeek V4 Pro (cross-vendor)
- Research F2 can stay Terra; optional: add Qwen as Research F3 for market-heavy verticals
F. Research brief schema addition (contract)
## Market block (mandatory)
- Competitors: name, category (authoring vs critique vs other), price anchor, source URL
- Non-Western / non-US comps: min 1 when vertical in {superapp, cross-border, unclassified-global}; else "N/A — US-centric vertical"
- TAM/SAM claims: verified | overstated | unverifiable — with citation or explicit gap
- Density judgment: sparse | contested | saturated — one paragraph, citations only
Downstream judges cite the Market block ID, not re-crawl the open web (except Legal/Financial specialists on their narrow facts).
(2) Reduction %
| Metric | Before (v2.1) | After | Delta |
|---|---|---|---|
| Seats (Enterprise default) | 13 | 11 | −15.4% |
| Phases | 11 | 10 | −9.1% |
| Parallel CC calls | 3 | 2 | −33% of CC fan-out |
| Scoring judges | 10 | 8 | −20% |
| Serial specialist after scores | Market + Exec + Legal… | Exec + Legal… | −1 serial hop |
| Est. E2E latency (vs 218s baseline) | 218s | ~185–195s | ~11–15% wall-clock |
| Est. token/COGS share | full panel | −Market narrative −1 full score JSON | ~14–18% cost |
Headline reduction: ~15% (seat count exact at −15.4%; cost/latency band 14–18% / 11–15%).
Why not larger: cutting Audit or a second CC would violate constraints or erase the disagreement signal that makes panels worth running (selection-bottleneck literature: judge-based selection > synthesis; arxiv 2603.20324).
(3) What survives (effectiveness preserved)
| Kept | Why it still works |
|---|---|
| Research Agent | Stronger, not weaker — owns market grounding explicitly |
| Audit Agent | Groupthink / blind-spot gate unchanged |
| ≥2 cross-checks | Sol + DeepSeek: US frontier + non-Western lab; blind; proposal+brief only |
| Primary + Validation | Dual Anthropic pass with sober second read (different tiers A/B) |
| Financial, Team, Execution, Legal | True specialists; low overlap with Research crawl |
| Reasoning-Verification | Contradiction synthesis over fewer, higher-signal score sets |
| Vendor cap | After Legal→Qwen: Anthropic 3/11 = 27.3% (under 33%) |
| Pool diversity | Gemini remains in pool/fallbacks; White-Label can still pin full custom panels |
| Model role cap | Fable Team+Audit = 2; all others ≤1 |
(4) Risks
| Risk | Severity | Mitigation |
|---|---|---|
| Market nuance loss on superapp / cross-border verticals | Med | Research Market block mandatory; §3.2 superapp trigger increases Research market citation quota + Primary market-dimension weight +30% |
| DeepSeek as sole non-Western CC may under-challenge US-centric Primary | Med | Legal→Qwen adds second non-Western seat; Audit watches for US-default groupthink |
| Anthropic share 4/11 if Legal stays Sonnet | High (rule break) | Must reassign Legal → Qwen (included in this same cut) |
| Sol still cannot tool-use on admin-ai | Low | CC-A is scoring-only JSON — already OK per §4.5 |
| Dropping third CC reduces disagreement surface | Med-Low | Literature + internal 0.00 market spread say third was fake diversity; monitor panel sigma for 30 days; if sigma collapses, restore CC-C on High complexity only (§3.3 already expands high-complexity) |
| Research brief becomes single point of market failure | Med | Research fallback chain unchanged (Gemini grounding → Terra); malformed Research → subscriber notify + degrade already specified |
| Pro tier "7 judges" matrix needs rewrite | Low | Define Pro as: Research + Primary + Validation + 2 CC + Execution + Reasoning (no Financial/Team/Legal/Audit) — document explicitly |
| Build-gate still unproven (panel delta <0.5 vs solo) | Existential (pre-existing) | This cut helps the thesis: fewer correlated seats make a ≥0.5 panel advantage more plausible if one exists; still run 3–5 real proposals before build |
(5) New diagram
Before (v2.1) — 11 phases / 13 seats
┌─────────────────────┐
│ 1 Research (Grok) │ grounding brief
└──────────┬──────────┘
▼
┌─────────────────────┐
│ 2 Primary (Opus 5) │ scores
└──────────┬──────────┘
▼
┌─────────────────────┐
│ 3 Validation (Sonnet)│ challenge
└──────────┬──────────┘
▼
┌────────────────┼────────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│4a CC-A │ │4b CC-B │ │4c CC-C │ blind re-scores
│ Sol │ │ Gemini │ │ DeepSeek │
└────┬─────┘ └────┬─────┘ └────┬─────┘
└────────────────┼────────────────┘
▼
┌─────────────────────┐
│ 5 Financial (M3) │
└──────────┬──────────┘
▼
┌─────────────────────┐
│ 6 Team (Fable) │
└──────────┬──────────┘
▼
┌─────────────────────┐
│ 7 Market (Qwen) │ ← re-summarizes Research
└──────────┬──────────┘
▼
┌─────────────────────┐
│ 8 Execution (Terra) │
└──────────┬──────────┘
▼
┌─────────────────────┐
│ 9 Legal (Sonnet) │
└──────────┬──────────┘
▼
┌─────────────────────┐
│10 Reasoning (Kimi) │ synthesize
└──────────┬──────────┘
▼
┌─────────────────────┐
│11 Audit (Fable) │ fenced ±0.5
└─────────────────────┘
After — Grounding-Redundancy Cut — 10 phases / 11 seats (−15.4%)
┌──────────────────────────────────┐
│ 1 Research (Grok 4.5) │
│ + mandatory Market block │ LIVE WEB (only)
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│ 2 Primary (Opus 5) │
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│ 3 Validation (Sonnet 5) │
└────────────────┬─────────────────┘
▼
┌───────────┴───────────┐
▼ ▼
┌─────────────┐ ┌─────────────┐
│4a CC-A Sol │ │4b CC-B │ 2 blind CCs
│ │ │ DeepSeek │ (min met)
└──────┬──────┘ └──────┬──────┘
└───────────┬───────────┘
▼
┌──────────────────────────────────┐
│ 5 Financial (MiniMax-M3) │
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│ 6 Team/Founder (Fable 5) │
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│ 7 Execution (Terra) │ Market phase GONE
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│ 8 Legal (Qwen3.7 Plus) │ vendor-cap patch
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│ 9 Reasoning-Verification (Kimi) │
└────────────────┬─────────────────┘
▼
┌──────────────────────────────────┐
│10 Audit (Fable 5) │ SURVIVES
└──────────────────────────────────┘
Signal flow (Research lens)
Research brief
├─ facts/citations ──────────────► all scorers (unchanged)
├─ Market block (NEW) ───────────► Primary / Validation / CCs
│ (no Phase-7 rewrite)
└─ gaps/unverifiable ────────────► Audit watches for overclaim
Dropped edges (were low-signal):
Research ══re-summary══► Market narrative ══► Reasoning
CC-C ══correlate≈0══► CC average
Constraint checklist
| Constraint | Status |
|---|---|
| ~15% reduction | 15.4% seats; ~14–18% COGS; ~11–15% latency |
| No model >2 roles | Fable Team+Audit=2; all others ≤1 |
| Audit survives | Phase 10' Fable, fenced |
| Research survives | Phase 1 Grok, expanded |
| Minimum 2 cross-checks | Sol + DeepSeek |
| Vendor ≤33% | After Legal→Qwen: Anthropic 3/11 = 27.3% |
Implementation note (one PR)
- Spec §1.1 / §1.2 / §2.3 / §3.2 / §3.3 / §4.1 / §6.3 — role counts must all read 11 seats / 10 phases
- Cross-section reconciliation (lesson from v2.0→v2.1 role expansion) before any re-review
- Cache-bust deploy
judge-pool-spec.md?v=2.2 - Do not build until 3–5 real proposals still clear the ≥0.5 panel-vs-solo gate on the reduced panel
Bottom line
ONE consolidation: Grounding-Redundancy Cut — delete Market-Reality as a serial phase (it re-summarizes Research) and delete the third cross-check (empirically zero incremental spread); fold market into the Research brief; keep DeepSeek as CC-B; move Legal to Qwen for vendor-cap compliance.
−15.4% seats, Research + Audit intact, 2 blind cross-checks retained, specialists preserved.