Files
itpp-infrastructure/proposals/verdicttank/consolidation-research-agent-15pct.md
T
root f5175f1ce0 Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
2026-08-26 02:27:28 -04:00

18 KiB
Raw Blame History

VerdictTank v2.1 — ONE Consolidation Recommendation

Research Agent Grounding Lens (Grok 4.5)

Date: 2026-08-12
Author role: Research Agent (only judge with native live web + social search)
Target: ~15% reduction without losing effectiveness
Constraints satisfied: no model >2 roles · Audit + Research survive · ≥2 cross-checks


Research-lens thesis

I produce the factual brief every downstream judge consumes. From that position, two seats add latency and tokens without new signal:

  1. Market-Reality (Phase 7) mostly re-summarizes my brief — competitive density, TAM defensibility, and citation-backed market claims are already Research outputs. Its unique residual (non-Western lens) is a prompt/schema requirement, not a reason for a full serial phase.
  2. Cross-Check C as a third parallel scorer is highly correlated with A/B. Internal Opus 5 empirical validation on this stack found all three cross-checks scored market at exactly 6 (spread 0.00). External ensemble guidance: if judges agree on everything, you bought one verdict three times (orq.ai LLM juries; Verga et al. "Judges → Juries"). Two diverse cross-checks capture the ensemble; a third correlated scorer is mostly cost.

Specialist phases that do not rehash Research (Financial math, Team/Founder fit, Execution ops, Legal/regulatory code) stay. Audit and Reasoning stay as non-scoring synthesis/gates.


(1) Changes — single consolidation: Grounding-Redundancy Cut

A. Eliminate standalone Market-Reality (Phase 7)

Action Detail
Remove Phase 7 seat: Market-Reality / Qwen3.7 Plus (default)
Absorb into Research Expand Phase 1 brief schema with a mandatory Market block: competitive density table, TAM/SAM defensibility notes, ≥1 non-Western comps when vertical warrants, citation URLs, "unknown/unverifiable" flags
Absorb residual scoring Market dimension continues via Primary + Validation + 2 cross-checks (already score all 10 dimensions). No separate market narrative phase
Vertical weights (§3.2) Market-Reality weight boosts → apply to Research brief depth (more market citations) + Primary/Validation market-dimension weight, not a missing judge

B. Collapse cross-checks 3 → 2 (keep minimum)

Action Detail
Remove Phase 4c Cross-Check C as default Enterprise seat
Keep Phase 4a Cross-Check A (GPT-5.6 Sol) + Phase 4b Cross-Check B
Diversity preserve Reassign DeepSeek V4 Pro → Cross-Check B (replace Gemini Pro Latest in default lineup). Reasons: (i) keeps Chinese-lab / non-Western error surface that Market+CC-C previously carried; (ii) avoids Gemini floating-alias longitudinal noise documented in v2.1 §4.5 and review results; (iii) Gemini remains Fallback 1 for Research and available in pool
Blindness Both remaining CCs still see proposal + Research brief ONLY — unchanged isolation

C. What does not change

  • Research Agent never scores
  • Audit Agent remains fenced, post-scoring, read-only
  • Reasoning-Verification still synthesizes all score sets
  • Financial, Team/Founder, Execution, Legal stay as serial specialists
  • Fable double: Team + Audit (still 2)
  • No other model exceeds 2 roles

D. Lineup delta (Enterprise default)

# Phase Role Model (after) Vendor Change
1 1 Research Agent Grok 4.5 xAI Survives — brief schema expanded
2 2 Primary Reviewer Claude Opus 5 Anthropic Unchanged
3 3 Validation Reviewer Claude Sonnet 5 Anthropic Unchanged
4 4a Cross-Check A GPT-5.6 Sol OpenAI Unchanged
5 4b Cross-Check B DeepSeek V4 Pro DeepSeek Was Gemini; DeepSeek moved here
4c Cross-Check C REMOVED
6 5 Financial Integrity MiniMax-M3 MiniMax Unchanged
7 6 Team/Founder Claude Fable 5 Anthropic Unchanged
7 Market-Reality REMOVED (folded into Research)
8 7' Execution-Feasibility GPT-5.6 Terra OpenAI Unchanged
9 8' Legal/Regulatory Qwen3.7 Plus Alibaba Vendor-cap patch (was Sonnet)
10 9' Reasoning-Verification Kimi K2.6 Moonshot Unchanged
11 10' Audit Agent Claude Fable 5 Anthropic Survives

Seats: 13 → 11
Phases: 11 → 10 (4a4b parallel; Market gone)
Scoring judges: 10 → 8 (Primary, Validation, CC-A, CC-B, Financial, Team, Execution, Legal)
Non-scoring: Research, Reasoning, Audit (3)
Total: 8 + 3 = 11

Qwen3.7 Plus remains on the default path via Legal (not dropped from Enterprise). Gemini stays in pool + Research fallback.

E. Fallback table patches (must ship with this cut)

  • Delete Market-Reality row or keep as optional on-demand role for White-Label only
  • Cross-Check B primary becomes DeepSeek V4 Pro; F1 Gemini Pro Latest; F2 Claude Fable 5 (still cross-vendor)
  • Cross-Check C row retired from Enterprise default (pool may retain for WL custom panels)
  • Legal primary: Qwen3.7 Plus; F1 Gemini Pro Latest; F2 DeepSeek V4 Pro (cross-vendor)
  • Research F2 can stay Terra; optional: add Qwen as Research F3 for market-heavy verticals

F. Research brief schema addition (contract)

## Market block (mandatory)
- Competitors: name, category (authoring vs critique vs other), price anchor, source URL
- Non-Western / non-US comps: min 1 when vertical in {superapp, cross-border, unclassified-global}; else "N/A — US-centric vertical"
- TAM/SAM claims: verified | overstated | unverifiable — with citation or explicit gap
- Density judgment: sparse | contested | saturated — one paragraph, citations only

Downstream judges cite the Market block ID, not re-crawl the open web (except Legal/Financial specialists on their narrow facts).


(2) Reduction %

Metric Before (v2.1) After Delta
Seats (Enterprise default) 13 11 15.4%
Phases 11 10 9.1%
Parallel CC calls 3 2 33% of CC fan-out
Scoring judges 10 8 20%
Serial specialist after scores Market + Exec + Legal… Exec + Legal… 1 serial hop
Est. E2E latency (vs 218s baseline) 218s ~185195s ~1115% wall-clock
Est. token/COGS share full panel Market narrative 1 full score JSON ~1418% cost

Headline reduction: ~15% (seat count exact at 15.4%; cost/latency band 1418% / 1115%).

Why not larger: cutting Audit or a second CC would violate constraints or erase the disagreement signal that makes panels worth running (selection-bottleneck literature: judge-based selection > synthesis; arxiv 2603.20324).


(3) What survives (effectiveness preserved)

Kept Why it still works
Research Agent Stronger, not weaker — owns market grounding explicitly
Audit Agent Groupthink / blind-spot gate unchanged
≥2 cross-checks Sol + DeepSeek: US frontier + non-Western lab; blind; proposal+brief only
Primary + Validation Dual Anthropic pass with sober second read (different tiers A/B)
Financial, Team, Execution, Legal True specialists; low overlap with Research crawl
Reasoning-Verification Contradiction synthesis over fewer, higher-signal score sets
Vendor cap After Legal→Qwen: Anthropic 3/11 = 27.3% (under 33%)
Pool diversity Gemini remains in pool/fallbacks; White-Label can still pin full custom panels
Model role cap Fable Team+Audit = 2; all others ≤1

(4) Risks

Risk Severity Mitigation
Market nuance loss on superapp / cross-border verticals Med Research Market block mandatory; §3.2 superapp trigger increases Research market citation quota + Primary market-dimension weight +30%
DeepSeek as sole non-Western CC may under-challenge US-centric Primary Med Legal→Qwen adds second non-Western seat; Audit watches for US-default groupthink
Anthropic share 4/11 if Legal stays Sonnet High (rule break) Must reassign Legal → Qwen (included in this same cut)
Sol still cannot tool-use on admin-ai Low CC-A is scoring-only JSON — already OK per §4.5
Dropping third CC reduces disagreement surface Med-Low Literature + internal 0.00 market spread say third was fake diversity; monitor panel sigma for 30 days; if sigma collapses, restore CC-C on High complexity only (§3.3 already expands high-complexity)
Research brief becomes single point of market failure Med Research fallback chain unchanged (Gemini grounding → Terra); malformed Research → subscriber notify + degrade already specified
Pro tier "7 judges" matrix needs rewrite Low Define Pro as: Research + Primary + Validation + 2 CC + Execution + Reasoning (no Financial/Team/Legal/Audit) — document explicitly
Build-gate still unproven (panel delta <0.5 vs solo) Existential (pre-existing) This cut helps the thesis: fewer correlated seats make a ≥0.5 panel advantage more plausible if one exists; still run 35 real proposals before build

(5) New diagram

Before (v2.1) — 11 phases / 13 seats

                    ┌─────────────────────┐
                    │ 1 Research (Grok)   │  grounding brief
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 2 Primary (Opus 5)  │  scores
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 3 Validation (Sonnet)│  challenge
                    └──────────┬──────────┘
                               ▼
              ┌────────────────┼────────────────┐
              ▼                ▼                ▼
        ┌──────────┐    ┌──────────┐    ┌──────────┐
        │4a CC-A   │    │4b CC-B   │    │4c CC-C   │  blind re-scores
        │ Sol      │    │ Gemini   │    │ DeepSeek │
        └────┬─────┘    └────┬─────┘    └────┬─────┘
              └────────────────┼────────────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 5 Financial (M3)    │
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 6 Team (Fable)      │
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 7 Market (Qwen)     │  ← re-summarizes Research
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 8 Execution (Terra) │
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │ 9 Legal (Sonnet)    │
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │10 Reasoning (Kimi)  │  synthesize
                    └──────────┬──────────┘
                               ▼
                    ┌─────────────────────┐
                    │11 Audit (Fable)     │  fenced ±0.5
                    └─────────────────────┘

After — Grounding-Redundancy Cut — 10 phases / 11 seats (15.4%)

                    ┌──────────────────────────────────┐
                    │ 1 Research (Grok 4.5)            │
                    │    + mandatory Market block      │  LIVE WEB (only)
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 2 Primary (Opus 5)               │
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 3 Validation (Sonnet 5)          │
                    └────────────────┬─────────────────┘
                                     ▼
                         ┌───────────┴───────────┐
                         ▼                       ▼
                  ┌─────────────┐         ┌─────────────┐
                  │4a CC-A Sol  │         │4b CC-B      │  2 blind CCs
                  │             │         │ DeepSeek    │  (min met)
                  └──────┬──────┘         └──────┬──────┘
                         └───────────┬───────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 5 Financial (MiniMax-M3)         │
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 6 Team/Founder (Fable 5)         │
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 7 Execution (Terra)              │  Market phase GONE
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 8 Legal (Qwen3.7 Plus)           │  vendor-cap patch
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │ 9 Reasoning-Verification (Kimi)  │
                    └────────────────┬─────────────────┘
                                     ▼
                    ┌──────────────────────────────────┐
                    │10 Audit (Fable 5)                │  SURVIVES
                    └──────────────────────────────────┘

Signal flow (Research lens)

 Research brief
   ├─ facts/citations ──────────────► all scorers (unchanged)
   ├─ Market block (NEW) ───────────► Primary / Validation / CCs
   │                                   (no Phase-7 rewrite)
   └─ gaps/unverifiable ────────────► Audit watches for overclaim

 Dropped edges (were low-signal):
   Research ══re-summary══► Market narrative ══► Reasoning
   CC-C ══correlate≈0══► CC average

Constraint checklist

Constraint Status
~15% reduction 15.4% seats; ~1418% COGS; ~1115% latency
No model >2 roles Fable Team+Audit=2; all others ≤1
Audit survives Phase 10' Fable, fenced
Research survives Phase 1 Grok, expanded
Minimum 2 cross-checks Sol + DeepSeek
Vendor ≤33% After Legal→Qwen: Anthropic 3/11 = 27.3%

Implementation note (one PR)

  1. Spec §1.1 / §1.2 / §2.3 / §3.2 / §3.3 / §4.1 / §6.3 — role counts must all read 11 seats / 10 phases
  2. Cross-section reconciliation (lesson from v2.0→v2.1 role expansion) before any re-review
  3. Cache-bust deploy judge-pool-spec.md?v=2.2
  4. Do not build until 35 real proposals still clear the ≥0.5 panel-vs-solo gate on the reduced panel

Bottom line

ONE consolidation: Grounding-Redundancy Cut — delete Market-Reality as a serial phase (it re-summarizes Research) and delete the third cross-check (empirically zero incremental spread); fold market into the Research brief; keep DeepSeek as CC-B; move Legal to Qwen for vendor-cap compliance.

15.4% seats, Research + Audit intact, 2 blind cross-checks retained, specialists preserved.