Files
itpp-infrastructure/proposals/verdicttank/consolidation-research-agent-15pct.md
T
root f5175f1ce0 Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
2026-08-26 02:27:28 -04:00

294 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# VerdictTank v2.1 — ONE Consolidation Recommendation
## Research Agent Grounding Lens (Grok 4.5)
**Date:** 2026-08-12
**Author role:** Research Agent (only judge with native live web + social search)
**Target:** ~15% reduction without losing effectiveness
**Constraints satisfied:** no model >2 roles · Audit + Research survive · ≥2 cross-checks
---
## Research-lens thesis
I produce the factual brief every downstream judge consumes. From that position, two seats add **latency and tokens without new signal**:
1. **Market-Reality (Phase 7)** mostly **re-summarizes my brief** — competitive density, TAM defensibility, and citation-backed market claims are already Research outputs. Its unique residual (non-Western lens) is a *prompt/schema* requirement, not a reason for a full serial phase.
2. **Cross-Check C as a third parallel scorer** is **highly correlated** with A/B. Internal Opus 5 empirical validation on this stack found all three cross-checks scored `market` at **exactly 6** (spread **0.00**). External ensemble guidance: if judges agree on everything, you bought one verdict three times (orq.ai LLM juries; Verga et al. "Judges → Juries"). Two diverse cross-checks capture the ensemble; a third correlated scorer is mostly cost.
Specialist phases that do **not** rehash Research (Financial math, Team/Founder fit, Execution ops, Legal/regulatory code) stay. Audit and Reasoning stay as non-scoring synthesis/gates.
---
## (1) Changes — single consolidation: **Grounding-Redundancy Cut**
### A. Eliminate standalone Market-Reality (Phase 7)
| Action | Detail |
|---|---|
| Remove | Phase 7 seat: Market-Reality / Qwen3.7 Plus (default) |
| Absorb into Research | Expand Phase 1 brief schema with a mandatory **Market block**: competitive density table, TAM/SAM defensibility notes, ≥1 non-Western comps when vertical warrants, citation URLs, "unknown/unverifiable" flags |
| Absorb residual scoring | Market *dimension* continues via Primary + Validation + 2 cross-checks (already score all 10 dimensions). No separate market narrative phase |
| Vertical weights (§3.2) | Market-Reality weight boosts → apply to **Research brief depth** (more market citations) + **Primary/Validation market-dimension weight**, not a missing judge |
### B. Collapse cross-checks 3 → 2 (keep minimum)
| Action | Detail |
|---|---|
| Remove | Phase 4c Cross-Check C as default Enterprise seat |
| Keep | Phase 4a Cross-Check A (GPT-5.6 Sol) + Phase 4b Cross-Check B |
| Diversity preserve | **Reassign DeepSeek V4 Pro → Cross-Check B** (replace Gemini Pro Latest in default lineup). Reasons: (i) keeps Chinese-lab / non-Western error surface that Market+CC-C previously carried; (ii) avoids Gemini floating-alias longitudinal noise documented in v2.1 §4.5 and review results; (iii) Gemini remains Fallback 1 for Research and available in pool |
| Blindness | Both remaining CCs still see **proposal + Research brief ONLY** — unchanged isolation |
### C. What does *not* change
- Research Agent never scores
- Audit Agent remains fenced, post-scoring, read-only
- Reasoning-Verification still synthesizes all score sets
- Financial, Team/Founder, Execution, Legal stay as serial specialists
- Fable double: Team + Audit (still 2)
- No other model exceeds 2 roles
### D. Lineup delta (Enterprise default)
| # | Phase | Role | Model (after) | Vendor | Change |
|---|---|---|---|---|---|
| 1 | 1 | Research Agent | Grok 4.5 | xAI | **Survives** — brief schema expanded |
| 2 | 2 | Primary Reviewer | Claude Opus 5 | Anthropic | Unchanged |
| 3 | 3 | Validation Reviewer | Claude Sonnet 5 | Anthropic | Unchanged |
| 4 | 4a | Cross-Check A | GPT-5.6 Sol | OpenAI | Unchanged |
| 5 | 4b | Cross-Check B | **DeepSeek V4 Pro** | DeepSeek | **Was Gemini; DeepSeek moved here** |
| — | ~~4c~~ | ~~Cross-Check C~~ | — | — | **REMOVED** |
| 6 | 5 | Financial Integrity | MiniMax-M3 | MiniMax | Unchanged |
| 7 | 6 | Team/Founder | Claude Fable 5 | Anthropic | Unchanged |
| — | ~~7~~ | ~~Market-Reality~~ | — | — | **REMOVED** (folded into Research) |
| 8 | 7' | Execution-Feasibility | GPT-5.6 Terra | OpenAI | Unchanged |
| 9 | 8' | Legal/Regulatory | **Qwen3.7 Plus** | Alibaba | **Vendor-cap patch** (was Sonnet) |
| 10 | 9' | Reasoning-Verification | Kimi K2.6 | Moonshot | Unchanged |
| 11 | 10' | Audit Agent | Claude Fable 5 | Anthropic | **Survives** |
**Seats:** 13 → **11**
**Phases:** 11 → **10** (4a4b parallel; Market gone)
**Scoring judges:** 10 → **8** (Primary, Validation, CC-A, CC-B, Financial, Team, Execution, Legal)
**Non-scoring:** Research, Reasoning, Audit (3)
**Total:** 8 + 3 = **11**
Qwen3.7 Plus remains on the default path via Legal (not dropped from Enterprise). Gemini stays in pool + Research fallback.
### E. Fallback table patches (must ship with this cut)
- Delete Market-Reality row **or** keep as optional on-demand role for White-Label only
- Cross-Check B primary becomes DeepSeek V4 Pro; F1 Gemini Pro Latest; F2 Claude Fable 5 (still cross-vendor)
- Cross-Check C row retired from Enterprise default (pool may retain for WL custom panels)
- Legal primary: Qwen3.7 Plus; F1 Gemini Pro Latest; F2 DeepSeek V4 Pro (cross-vendor)
- Research F2 can stay Terra; optional: add Qwen as Research F3 for market-heavy verticals
### F. Research brief schema addition (contract)
```
## Market block (mandatory)
- Competitors: name, category (authoring vs critique vs other), price anchor, source URL
- Non-Western / non-US comps: min 1 when vertical in {superapp, cross-border, unclassified-global}; else "N/A — US-centric vertical"
- TAM/SAM claims: verified | overstated | unverifiable — with citation or explicit gap
- Density judgment: sparse | contested | saturated — one paragraph, citations only
```
Downstream judges **cite the Market block ID**, not re-crawl the open web (except Legal/Financial specialists on their narrow facts).
---
## (2) Reduction %
| Metric | Before (v2.1) | After | Delta |
|---|---|---|---|
| Seats (Enterprise default) | 13 | 11 | **15.4%** |
| Phases | 11 | 10 | 9.1% |
| Parallel CC calls | 3 | 2 | 33% of CC fan-out |
| Scoring judges | 10 | 8 | 20% |
| Serial specialist after scores | Market + Exec + Legal… | Exec + Legal… | 1 serial hop |
| Est. E2E latency (vs 218s baseline) | 218s | ~185195s | **~1115%** wall-clock |
| Est. token/COGS share | full panel | Market narrative 1 full score JSON | **~1418%** cost |
**Headline reduction: ~15%** (seat count exact at 15.4%; cost/latency band 1418% / 1115%).
Why not larger: cutting Audit or a second CC would violate constraints or erase the disagreement signal that makes panels worth running (selection-bottleneck literature: judge-based selection > synthesis; arxiv 2603.20324).
---
## (3) What survives (effectiveness preserved)
| Kept | Why it still works |
|---|---|
| **Research Agent** | Stronger, not weaker — owns market grounding explicitly |
| **Audit Agent** | Groupthink / blind-spot gate unchanged |
| **≥2 cross-checks** | Sol + DeepSeek: US frontier + non-Western lab; blind; proposal+brief only |
| **Primary + Validation** | Dual Anthropic pass with sober second read (different tiers A/B) |
| **Financial, Team, Execution, Legal** | True specialists; low overlap with Research crawl |
| **Reasoning-Verification** | Contradiction synthesis over fewer, higher-signal score sets |
| **Vendor cap** | After Legal→Qwen: Anthropic 3/11 = **27.3%** (under 33%) |
| **Pool diversity** | Gemini remains in pool/fallbacks; White-Label can still pin full custom panels |
| **Model role cap** | Fable Team+Audit = 2; all others ≤1 |
---
## (4) Risks
| Risk | Severity | Mitigation |
|---|---|---|
| Market nuance loss on superapp / cross-border verticals | Med | Research Market block mandatory; §3.2 superapp trigger increases Research market citation quota + Primary market-dimension weight +30% |
| DeepSeek as sole non-Western CC may under-challenge US-centric Primary | Med | Legal→Qwen adds second non-Western seat; Audit watches for US-default groupthink |
| Anthropic share 4/11 if Legal stays Sonnet | High (rule break) | **Must** reassign Legal → Qwen (included in this same cut) |
| Sol still cannot tool-use on admin-ai | Low | CC-A is scoring-only JSON — already OK per §4.5 |
| Dropping third CC reduces disagreement surface | Med-Low | Literature + internal 0.00 market spread say third was fake diversity; monitor panel sigma for 30 days; if sigma collapses, restore CC-C on High complexity only (§3.3 already expands high-complexity) |
| Research brief becomes single point of market failure | Med | Research fallback chain unchanged (Gemini grounding → Terra); malformed Research → subscriber notify + degrade already specified |
| Pro tier "7 judges" matrix needs rewrite | Low | Define Pro as: Research + Primary + Validation + 2 CC + Execution + Reasoning (no Financial/Team/Legal/Audit) — document explicitly |
| Build-gate still unproven (panel delta <0.5 vs solo) | Existential (pre-existing) | This cut helps the thesis: fewer correlated seats make a ≥0.5 panel advantage more plausible if one exists; still run 35 real proposals before build |
---
## (5) New diagram
### Before (v2.1) — 11 phases / 13 seats
```
┌─────────────────────┐
│ 1 Research (Grok) │ grounding brief
└──────────┬──────────┘
┌─────────────────────┐
│ 2 Primary (Opus 5) │ scores
└──────────┬──────────┘
┌─────────────────────┐
│ 3 Validation (Sonnet)│ challenge
└──────────┬──────────┘
┌────────────────┼────────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│4a CC-A │ │4b CC-B │ │4c CC-C │ blind re-scores
│ Sol │ │ Gemini │ │ DeepSeek │
└────┬─────┘ └────┬─────┘ └────┬─────┘
└────────────────┼────────────────┘
┌─────────────────────┐
│ 5 Financial (M3) │
└──────────┬──────────┘
┌─────────────────────┐
│ 6 Team (Fable) │
└──────────┬──────────┘
┌─────────────────────┐
│ 7 Market (Qwen) │ ← re-summarizes Research
└──────────┬──────────┘
┌─────────────────────┐
│ 8 Execution (Terra) │
└──────────┬──────────┘
┌─────────────────────┐
│ 9 Legal (Sonnet) │
└──────────┬──────────┘
┌─────────────────────┐
│10 Reasoning (Kimi) │ synthesize
└──────────┬──────────┘
┌─────────────────────┐
│11 Audit (Fable) │ fenced ±0.5
└─────────────────────┘
```
### After — Grounding-Redundancy Cut — 10 phases / 11 seats (15.4%)
```
┌──────────────────────────────────┐
│ 1 Research (Grok 4.5) │
│ + mandatory Market block │ LIVE WEB (only)
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 2 Primary (Opus 5) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 3 Validation (Sonnet 5) │
└────────────────┬─────────────────┘
┌───────────┴───────────┐
▼ ▼
┌─────────────┐ ┌─────────────┐
│4a CC-A Sol │ │4b CC-B │ 2 blind CCs
│ │ │ DeepSeek │ (min met)
└──────┬──────┘ └──────┬──────┘
└───────────┬───────────┘
┌──────────────────────────────────┐
│ 5 Financial (MiniMax-M3) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 6 Team/Founder (Fable 5) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 7 Execution (Terra) │ Market phase GONE
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 8 Legal (Qwen3.7 Plus) │ vendor-cap patch
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 9 Reasoning-Verification (Kimi) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│10 Audit (Fable 5) │ SURVIVES
└──────────────────────────────────┘
```
### Signal flow (Research lens)
```
Research brief
├─ facts/citations ──────────────► all scorers (unchanged)
├─ Market block (NEW) ───────────► Primary / Validation / CCs
│ (no Phase-7 rewrite)
└─ gaps/unverifiable ────────────► Audit watches for overclaim
Dropped edges (were low-signal):
Research ══re-summary══► Market narrative ══► Reasoning
CC-C ══correlate≈0══► CC average
```
---
## Constraint checklist
| Constraint | Status |
|---|---|
| ~15% reduction | **15.4% seats**; ~1418% COGS; ~1115% latency |
| No model >2 roles | Fable Team+Audit=2; all others ≤1 |
| Audit survives | Phase 10' Fable, fenced |
| Research survives | Phase 1 Grok, expanded |
| Minimum 2 cross-checks | Sol + DeepSeek |
| Vendor ≤33% | After Legal→Qwen: Anthropic 3/11 = 27.3% |
---
## Implementation note (one PR)
1. Spec §1.1 / §1.2 / §2.3 / §3.2 / §3.3 / §4.1 / §6.3 — role counts must all read **11 seats / 10 phases**
2. Cross-section reconciliation (lesson from v2.0→v2.1 role expansion) before any re-review
3. Cache-bust deploy `judge-pool-spec.md?v=2.2`
4. Do **not** build until 35 real proposals still clear the ≥0.5 panel-vs-solo gate on the *reduced* panel
---
## Bottom line
**ONE consolidation:** *Grounding-Redundancy Cut* — delete Market-Reality as a serial phase (it re-summarizes Research) and delete the third cross-check (empirically zero incremental spread); fold market into the Research brief; keep DeepSeek as CC-B; move Legal to Qwen for vendor-cap compliance.
**15.4% seats, Research + Audit intact, 2 blind cross-checks retained, specialists preserved.**