Sync docs, audit artifacts, project notes, and VerdictTank proposal docs

- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
This commit is contained in:
root
2026-08-26 02:27:28 -04:00
parent 23e9751d38
commit f5175f1ce0
55 changed files with 14669 additions and 3 deletions
@@ -0,0 +1,293 @@
# VerdictTank v2.1 — ONE Consolidation Recommendation
## Research Agent Grounding Lens (Grok 4.5)
**Date:** 2026-08-12
**Author role:** Research Agent (only judge with native live web + social search)
**Target:** ~15% reduction without losing effectiveness
**Constraints satisfied:** no model >2 roles · Audit + Research survive · ≥2 cross-checks
---
## Research-lens thesis
I produce the factual brief every downstream judge consumes. From that position, two seats add **latency and tokens without new signal**:
1. **Market-Reality (Phase 7)** mostly **re-summarizes my brief** — competitive density, TAM defensibility, and citation-backed market claims are already Research outputs. Its unique residual (non-Western lens) is a *prompt/schema* requirement, not a reason for a full serial phase.
2. **Cross-Check C as a third parallel scorer** is **highly correlated** with A/B. Internal Opus 5 empirical validation on this stack found all three cross-checks scored `market` at **exactly 6** (spread **0.00**). External ensemble guidance: if judges agree on everything, you bought one verdict three times (orq.ai LLM juries; Verga et al. "Judges → Juries"). Two diverse cross-checks capture the ensemble; a third correlated scorer is mostly cost.
Specialist phases that do **not** rehash Research (Financial math, Team/Founder fit, Execution ops, Legal/regulatory code) stay. Audit and Reasoning stay as non-scoring synthesis/gates.
---
## (1) Changes — single consolidation: **Grounding-Redundancy Cut**
### A. Eliminate standalone Market-Reality (Phase 7)
| Action | Detail |
|---|---|
| Remove | Phase 7 seat: Market-Reality / Qwen3.7 Plus (default) |
| Absorb into Research | Expand Phase 1 brief schema with a mandatory **Market block**: competitive density table, TAM/SAM defensibility notes, ≥1 non-Western comps when vertical warrants, citation URLs, "unknown/unverifiable" flags |
| Absorb residual scoring | Market *dimension* continues via Primary + Validation + 2 cross-checks (already score all 10 dimensions). No separate market narrative phase |
| Vertical weights (§3.2) | Market-Reality weight boosts → apply to **Research brief depth** (more market citations) + **Primary/Validation market-dimension weight**, not a missing judge |
### B. Collapse cross-checks 3 → 2 (keep minimum)
| Action | Detail |
|---|---|
| Remove | Phase 4c Cross-Check C as default Enterprise seat |
| Keep | Phase 4a Cross-Check A (GPT-5.6 Sol) + Phase 4b Cross-Check B |
| Diversity preserve | **Reassign DeepSeek V4 Pro → Cross-Check B** (replace Gemini Pro Latest in default lineup). Reasons: (i) keeps Chinese-lab / non-Western error surface that Market+CC-C previously carried; (ii) avoids Gemini floating-alias longitudinal noise documented in v2.1 §4.5 and review results; (iii) Gemini remains Fallback 1 for Research and available in pool |
| Blindness | Both remaining CCs still see **proposal + Research brief ONLY** — unchanged isolation |
### C. What does *not* change
- Research Agent never scores
- Audit Agent remains fenced, post-scoring, read-only
- Reasoning-Verification still synthesizes all score sets
- Financial, Team/Founder, Execution, Legal stay as serial specialists
- Fable double: Team + Audit (still 2)
- No other model exceeds 2 roles
### D. Lineup delta (Enterprise default)
| # | Phase | Role | Model (after) | Vendor | Change |
|---|---|---|---|---|---|
| 1 | 1 | Research Agent | Grok 4.5 | xAI | **Survives** — brief schema expanded |
| 2 | 2 | Primary Reviewer | Claude Opus 5 | Anthropic | Unchanged |
| 3 | 3 | Validation Reviewer | Claude Sonnet 5 | Anthropic | Unchanged |
| 4 | 4a | Cross-Check A | GPT-5.6 Sol | OpenAI | Unchanged |
| 5 | 4b | Cross-Check B | **DeepSeek V4 Pro** | DeepSeek | **Was Gemini; DeepSeek moved here** |
| — | ~~4c~~ | ~~Cross-Check C~~ | — | — | **REMOVED** |
| 6 | 5 | Financial Integrity | MiniMax-M3 | MiniMax | Unchanged |
| 7 | 6 | Team/Founder | Claude Fable 5 | Anthropic | Unchanged |
| — | ~~7~~ | ~~Market-Reality~~ | — | — | **REMOVED** (folded into Research) |
| 8 | 7' | Execution-Feasibility | GPT-5.6 Terra | OpenAI | Unchanged |
| 9 | 8' | Legal/Regulatory | **Qwen3.7 Plus** | Alibaba | **Vendor-cap patch** (was Sonnet) |
| 10 | 9' | Reasoning-Verification | Kimi K2.6 | Moonshot | Unchanged |
| 11 | 10' | Audit Agent | Claude Fable 5 | Anthropic | **Survives** |
**Seats:** 13 → **11**
**Phases:** 11 → **10** (4a4b parallel; Market gone)
**Scoring judges:** 10 → **8** (Primary, Validation, CC-A, CC-B, Financial, Team, Execution, Legal)
**Non-scoring:** Research, Reasoning, Audit (3)
**Total:** 8 + 3 = **11**
Qwen3.7 Plus remains on the default path via Legal (not dropped from Enterprise). Gemini stays in pool + Research fallback.
### E. Fallback table patches (must ship with this cut)
- Delete Market-Reality row **or** keep as optional on-demand role for White-Label only
- Cross-Check B primary becomes DeepSeek V4 Pro; F1 Gemini Pro Latest; F2 Claude Fable 5 (still cross-vendor)
- Cross-Check C row retired from Enterprise default (pool may retain for WL custom panels)
- Legal primary: Qwen3.7 Plus; F1 Gemini Pro Latest; F2 DeepSeek V4 Pro (cross-vendor)
- Research F2 can stay Terra; optional: add Qwen as Research F3 for market-heavy verticals
### F. Research brief schema addition (contract)
```
## Market block (mandatory)
- Competitors: name, category (authoring vs critique vs other), price anchor, source URL
- Non-Western / non-US comps: min 1 when vertical in {superapp, cross-border, unclassified-global}; else "N/A — US-centric vertical"
- TAM/SAM claims: verified | overstated | unverifiable — with citation or explicit gap
- Density judgment: sparse | contested | saturated — one paragraph, citations only
```
Downstream judges **cite the Market block ID**, not re-crawl the open web (except Legal/Financial specialists on their narrow facts).
---
## (2) Reduction %
| Metric | Before (v2.1) | After | Delta |
|---|---|---|---|
| Seats (Enterprise default) | 13 | 11 | **15.4%** |
| Phases | 11 | 10 | 9.1% |
| Parallel CC calls | 3 | 2 | 33% of CC fan-out |
| Scoring judges | 10 | 8 | 20% |
| Serial specialist after scores | Market + Exec + Legal… | Exec + Legal… | 1 serial hop |
| Est. E2E latency (vs 218s baseline) | 218s | ~185195s | **~1115%** wall-clock |
| Est. token/COGS share | full panel | Market narrative 1 full score JSON | **~1418%** cost |
**Headline reduction: ~15%** (seat count exact at 15.4%; cost/latency band 1418% / 1115%).
Why not larger: cutting Audit or a second CC would violate constraints or erase the disagreement signal that makes panels worth running (selection-bottleneck literature: judge-based selection > synthesis; arxiv 2603.20324).
---
## (3) What survives (effectiveness preserved)
| Kept | Why it still works |
|---|---|
| **Research Agent** | Stronger, not weaker — owns market grounding explicitly |
| **Audit Agent** | Groupthink / blind-spot gate unchanged |
| **≥2 cross-checks** | Sol + DeepSeek: US frontier + non-Western lab; blind; proposal+brief only |
| **Primary + Validation** | Dual Anthropic pass with sober second read (different tiers A/B) |
| **Financial, Team, Execution, Legal** | True specialists; low overlap with Research crawl |
| **Reasoning-Verification** | Contradiction synthesis over fewer, higher-signal score sets |
| **Vendor cap** | After Legal→Qwen: Anthropic 3/11 = **27.3%** (under 33%) |
| **Pool diversity** | Gemini remains in pool/fallbacks; White-Label can still pin full custom panels |
| **Model role cap** | Fable Team+Audit = 2; all others ≤1 |
---
## (4) Risks
| Risk | Severity | Mitigation |
|---|---|---|
| Market nuance loss on superapp / cross-border verticals | Med | Research Market block mandatory; §3.2 superapp trigger increases Research market citation quota + Primary market-dimension weight +30% |
| DeepSeek as sole non-Western CC may under-challenge US-centric Primary | Med | Legal→Qwen adds second non-Western seat; Audit watches for US-default groupthink |
| Anthropic share 4/11 if Legal stays Sonnet | High (rule break) | **Must** reassign Legal → Qwen (included in this same cut) |
| Sol still cannot tool-use on admin-ai | Low | CC-A is scoring-only JSON — already OK per §4.5 |
| Dropping third CC reduces disagreement surface | Med-Low | Literature + internal 0.00 market spread say third was fake diversity; monitor panel sigma for 30 days; if sigma collapses, restore CC-C on High complexity only (§3.3 already expands high-complexity) |
| Research brief becomes single point of market failure | Med | Research fallback chain unchanged (Gemini grounding → Terra); malformed Research → subscriber notify + degrade already specified |
| Pro tier "7 judges" matrix needs rewrite | Low | Define Pro as: Research + Primary + Validation + 2 CC + Execution + Reasoning (no Financial/Team/Legal/Audit) — document explicitly |
| Build-gate still unproven (panel delta <0.5 vs solo) | Existential (pre-existing) | This cut helps the thesis: fewer correlated seats make a ≥0.5 panel advantage more plausible if one exists; still run 35 real proposals before build |
---
## (5) New diagram
### Before (v2.1) — 11 phases / 13 seats
```
┌─────────────────────┐
│ 1 Research (Grok) │ grounding brief
└──────────┬──────────┘
┌─────────────────────┐
│ 2 Primary (Opus 5) │ scores
└──────────┬──────────┘
┌─────────────────────┐
│ 3 Validation (Sonnet)│ challenge
└──────────┬──────────┘
┌────────────────┼────────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│4a CC-A │ │4b CC-B │ │4c CC-C │ blind re-scores
│ Sol │ │ Gemini │ │ DeepSeek │
└────┬─────┘ └────┬─────┘ └────┬─────┘
└────────────────┼────────────────┘
┌─────────────────────┐
│ 5 Financial (M3) │
└──────────┬──────────┘
┌─────────────────────┐
│ 6 Team (Fable) │
└──────────┬──────────┘
┌─────────────────────┐
│ 7 Market (Qwen) │ ← re-summarizes Research
└──────────┬──────────┘
┌─────────────────────┐
│ 8 Execution (Terra) │
└──────────┬──────────┘
┌─────────────────────┐
│ 9 Legal (Sonnet) │
└──────────┬──────────┘
┌─────────────────────┐
│10 Reasoning (Kimi) │ synthesize
└──────────┬──────────┘
┌─────────────────────┐
│11 Audit (Fable) │ fenced ±0.5
└─────────────────────┘
```
### After — Grounding-Redundancy Cut — 10 phases / 11 seats (15.4%)
```
┌──────────────────────────────────┐
│ 1 Research (Grok 4.5) │
│ + mandatory Market block │ LIVE WEB (only)
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 2 Primary (Opus 5) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 3 Validation (Sonnet 5) │
└────────────────┬─────────────────┘
┌───────────┴───────────┐
▼ ▼
┌─────────────┐ ┌─────────────┐
│4a CC-A Sol │ │4b CC-B │ 2 blind CCs
│ │ │ DeepSeek │ (min met)
└──────┬──────┘ └──────┬──────┘
└───────────┬───────────┘
┌──────────────────────────────────┐
│ 5 Financial (MiniMax-M3) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 6 Team/Founder (Fable 5) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 7 Execution (Terra) │ Market phase GONE
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 8 Legal (Qwen3.7 Plus) │ vendor-cap patch
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│ 9 Reasoning-Verification (Kimi) │
└────────────────┬─────────────────┘
┌──────────────────────────────────┐
│10 Audit (Fable 5) │ SURVIVES
└──────────────────────────────────┘
```
### Signal flow (Research lens)
```
Research brief
├─ facts/citations ──────────────► all scorers (unchanged)
├─ Market block (NEW) ───────────► Primary / Validation / CCs
│ (no Phase-7 rewrite)
└─ gaps/unverifiable ────────────► Audit watches for overclaim
Dropped edges (were low-signal):
Research ══re-summary══► Market narrative ══► Reasoning
CC-C ══correlate≈0══► CC average
```
---
## Constraint checklist
| Constraint | Status |
|---|---|
| ~15% reduction | **15.4% seats**; ~1418% COGS; ~1115% latency |
| No model >2 roles | Fable Team+Audit=2; all others ≤1 |
| Audit survives | Phase 10' Fable, fenced |
| Research survives | Phase 1 Grok, expanded |
| Minimum 2 cross-checks | Sol + DeepSeek |
| Vendor ≤33% | After Legal→Qwen: Anthropic 3/11 = 27.3% |
---
## Implementation note (one PR)
1. Spec §1.1 / §1.2 / §2.3 / §3.2 / §3.3 / §4.1 / §6.3 — role counts must all read **11 seats / 10 phases**
2. Cross-section reconciliation (lesson from v2.0→v2.1 role expansion) before any re-review
3. Cache-bust deploy `judge-pool-spec.md?v=2.2`
4. Do **not** build until 35 real proposals still clear the ≥0.5 panel-vs-solo gate on the *reduced* panel
---
## Bottom line
**ONE consolidation:** *Grounding-Redundancy Cut* — delete Market-Reality as a serial phase (it re-summarizes Research) and delete the third cross-check (empirically zero incremental spread); fold market into the Research brief; keep DeepSeek as CC-B; move Legal to Qwen for vendor-cap compliance.
**15.4% seats, Research + Audit intact, 2 blind cross-checks retained, specialists preserved.**