- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs - disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md - clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy) - projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture - proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review - docs/super-search/firecrawl-provider-strategy.md - updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io - .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
334 lines
15 KiB
Markdown
334 lines
15 KiB
Markdown
# VerdictTank Pipeline Consolidation Recommendation
|
|
## Lens: Reasoning / Logic Overlap Detection (Contradiction Detector)
|
|
**Author:** Kimi K2.6 (Reasoning-Verification specialist)
|
|
**Spec:** judge-pool-spec.md v2.1
|
|
**Target:** ~15% reduction
|
|
**Date:** 2026-08-12
|
|
|
|
---
|
|
|
|
## Executive finding
|
|
|
|
**ONE recommendation:** Collapse **Phase 10 (Reasoning-Verification)** and **Phase 11 (Audit Agent)** into a single post-scoring synthesis seat: **Synthesis and Integrity Gate**.
|
|
|
|
These two roles ask the same meta-question with different labels:
|
|
|
|
| Role | Stated question | Actual job |
|
|
|---|---|---|
|
|
| Phase 10 Reasoning-Verification | Where do judges contradict? | Read all score sets, find disagreements, synthesize |
|
|
| Phase 11 Audit Agent | What did the panel miss / groupthink? | Read finished review, find blind spots / groupthink, adjust confidence |
|
|
|
|
Both are **non-scoring meta-reviewers of the panel's own output**. Neither sees the raw proposal as primary input in a unique way that the other does not eventually consume. Both produce a **post-hoc integrity adjustment**, not an independent domain score. Under a contradiction-detection lens, they are statistically and logically collinear seats.
|
|
|
|
---
|
|
|
|
## 1. Changes
|
|
|
|
### Collapse
|
|
|
|
| Remove | Absorb into |
|
|
|---|---|
|
|
| Phase 10: Reasoning-Verification (standalone) | **New Phase 10: Synthesis and Integrity Gate** |
|
|
| Phase 11: Audit Agent (standalone) | (same merged seat) |
|
|
|
|
### New merged role definition
|
|
|
|
| Field | Value |
|
|
|---|---|
|
|
| **Phase** | 10 (final) |
|
|
| **Role** | **Synthesis and Integrity Gate** |
|
|
| **Model** | Kimi K2.6 (Moonshot) — keeps contradiction specialty as primary |
|
|
| **Fallback 1** | Claude Fable 5 (Anthropic) — former Audit primary |
|
|
| **Fallback 2** | GPT-5.6 Sol (OpenAI) — former Reasoning F2 |
|
|
| **Scores?** | No (preserves non-scoring fence) |
|
|
| **Input** | ALL score sets from phases 2-9 + assembled draft review |
|
|
| **Output** | (a) Contradiction / alignment report, (b) synthesized dimension scores via section 4.1 formula, (c) blind-spot / groupthink notes, (d) single confidence adjustment +/-0.5 |
|
|
|
|
### Prompt contract (must stay explicit)
|
|
|
|
Merged prompt has **two ordered sections**, not a vague do-both:
|
|
|
|
1. **Contradiction pass** (former Phase 10): pairwise/cluster disagreement map across Primary, Validation, Cross-Check avg, Financial, Team, Market, Execution, Legal. Emit synthesized scores.
|
|
2. **Integrity pass** (former Phase 11): given the synthesis above, scan for groupthink (panel sigma too low), missed domain blind spots, and apply **one** confidence adjustment (+/-0.5).
|
|
|
|
Single model, single call, single output schema. No second serial hop.
|
|
|
|
### Spec section edits required
|
|
|
|
| Section | Edit |
|
|
|---|---|
|
|
| section 1.1 Pipeline Phases | Drop row 11; rewrite Phase 10 as Synthesis and Integrity Gate; phase count **11 to 10** |
|
|
| section 1.2 Enterprise Lineup | **13 seats to 11 seats** (12 unique models still available in pool; one fewer active seat) |
|
|
| section 1.2 Model double-ups | Remove Fable 5: Team + Audit double-up. Fable becomes Reasoning/Audit **fallback only** (Team primary only) |
|
|
| section 2.3 Fallback table | Delete Audit Agent row; replace Reasoning-Verification row with merged role chain above |
|
|
| section 3.1 Hard constraints | Replace "Audit reads FINISHED review only" with "Synthesis and Integrity Gate is post-scoring, read-only, cannot open new domain scores" |
|
|
| section 3.3 Low-complexity Enterprise | Already drops Audit — update text to drop CC-C, Team, Legal (Audit no longer separate) |
|
|
| section 4.1 Scoring algorithm | Merge the two non-scoring bullets into one: Synthesis and Integrity Gate produces contradiction report + synthesized scores + confidence +/-0.5 |
|
|
| section 4.2 Cost ceiling rule | Swap Tier A to B for non-scoring roles first (Research, Synthesis) — drop separate Audit |
|
|
| section 4.4 Latency | Enterprise sequential phases lose one hop — recalibrate target downward slightly (e.g. 180s to ~165s aspirational; keep 300s max) |
|
|
| section 6.3 Feature matrix | Replace Audit Agent (Phase 11) row with Synthesis and Integrity Gate (includes audit) — still Enterprise/WL only if desired |
|
|
|
|
### What this is NOT
|
|
|
|
- Not dropping Cross-Check C (that is diversity, not logic-overlap).
|
|
- Not merging Primary with Validation (Validation is deliberately contaminated by Primary; different question).
|
|
- Not merging Market with Execution (different input scopes and vertical triggers).
|
|
- Not touching Research (unique tool surface; never scores).
|
|
|
|
---
|
|
|
|
## 2. Reduction %
|
|
|
|
### Seat / phase accounting (Enterprise standard panel)
|
|
|
|
| Metric | Before (v2.1) | After | Delta |
|
|
|---|---|---|---|
|
|
| Pipeline phases | 11 | **10** | -1 phase (-9.1%) |
|
|
| Active seats | 13 | **11** | -2 seats (-15.4%) |
|
|
| Serial post-score hops | 2 (P10 then P11) | **1** | -50% of meta-tail latency |
|
|
| Non-scoring roles | 3 (Research, Reasoning, Audit) | **2** (Research, Synthesis) | -33% meta overhead |
|
|
| Unique models in default lineup | 11 distinct (Sonnet x2, Fable x2) | **11 distinct** (Sonnet still x2 Validation+Legal; Fable Team-only) | pool size unchanged |
|
|
| Anthropic share of active seats | 4/13 = 30.8% | **3/11 = 27.3%** | more headroom under 33% cap |
|
|
|
|
### Cost reduction (directional)
|
|
|
|
- Eliminate one full Tier-A (or Tier-A-fallback) completion on every Enterprise review.
|
|
- v2.1 cited about $1.30-2.10 per standard full review — expect **~12-18% COGS cut** on the judge tail (one fewer expensive synthesis call), landing near **~15%** all-in when weighted by token volume of P10+P11 vs earlier short scoring calls.
|
|
- Latency: remove one sequential dependency (~15-40s depending on model) from the critical path after phase 9.
|
|
|
|
**Headline reduction: ~15%** (seats -15.4%; cost band -12-18%; phases -9%).
|
|
|
|
---
|
|
|
|
## 3. What survives
|
|
|
|
### Intact domain + scoring spine
|
|
|
|
```
|
|
1 Research Agent (Grok 4.5) - grounding, never scores
|
|
2 Primary Reviewer (Claude Opus 5) - 10-dimension critique
|
|
3 Validation Reviewer (Claude Sonnet 5) - challenges Primary
|
|
4a Cross-Check A (GPT-5.6 Sol) - blind re-score
|
|
4b Cross-Check B (Gemini Pro Latest) - blind re-score
|
|
4c Cross-Check C (DeepSeek V4 Pro) - blind re-score
|
|
5 Financial Integrity (MiniMax-M3) - cap table / unit econ
|
|
6 Team/Founder Assessment (Claude Fable 5) - founder-market fit
|
|
7 Market-Reality (Qwen3.7 Plus) - TAM / competitive / non-Western
|
|
8 Execution-Feasibility (GPT-5.6 Terra) - ops / timeline realism
|
|
9 Legal/Regulatory (Claude Sonnet 5) - multi-jurisdiction risk
|
|
10 Synthesis and Integrity Gate (Kimi K2.6) - contradictions + audit + +/-0.5
|
|
```
|
|
|
|
### Preserved invariants
|
|
|
|
- Research never scores
|
|
- Three blind cross-checks, isolated from Primary/Validation
|
|
- Domain specialists (Financial, Team, Market, Execution, Legal) remain **separate scoring questions**
|
|
- Section 4.1 weighted aggregation formula unchanged
|
|
- Cross-vendor fallback rule unchanged (merged chain still cross-vendor at every hop)
|
|
- Vendor cap still satisfied (Anthropic drops to 27.3%)
|
|
- Build validation gate (section 5.5) unchanged — still require >=0.5 panel advantage
|
|
- Fenced post-score mode retained (merged seat cannot reopen domain scores)
|
|
|
|
### Why these roles were NOT collapsed
|
|
|
|
| Pair | Why keep separate |
|
|
|---|---|
|
|
| Primary vs Validation | Validation is **score-conditioned** (sees Primary). Different epistemic job than independent re-score. |
|
|
| Validation vs Cross-Checks | CCs are blind; Validation is not. Collapsing destroys the blind control. |
|
|
| Cross-Check A/B/C | Same question by design — but **vendor diversity is the product**. Overlap is intentional ensemble, not accidental duplication. Do not collapse under this lens. |
|
|
| Market vs Execution | Different section 3.2 triggers, different evidence (competitive density vs ops timeline). |
|
|
| Financial vs Execution | Cap-table math is not shipping realism. |
|
|
| Legal vs Audit | Legal is domain-scoring on statutes; Audit was meta. Legal stays; Audit folds into Synthesis. |
|
|
| Team vs Audit (same model Fable) | Team scores founders; Audit was meta. After merge, Fable stops double-duty — **removes same-model Team-to-Audit soft contamination path**. |
|
|
|
|
---
|
|
|
|
## 4. Risks
|
|
|
|
| Risk | Severity | Mitigation |
|
|
|---|---|---|
|
|
| **Single-call attention split** — model skimps on contradiction map OR groupthink scan | Medium | Enforce two-section JSON schema; reject output missing either contradictions[] or integrity.confidence_delta. |
|
|
| **Loss of independent second meta-reader** — two weak meta-passes can catch what one strong pass misses | Medium | Keep dual-section prompt; optionally run Fable fallback as **shadow audit** on 10% of Enterprise reviews for 90 days (not on critical path). |
|
|
| **K2.6 empty-output / reasoning-token burn** (section 4.5) now single-points the entire meta-tail | High | Raise min tokens to 768-1024 for merged seat; on empty go immediate F1 (Fable). Circuit breaker already in section 4.2. |
|
|
| **Confidence +/-0.5 applied by same model that synthesized scores** — self-grading bias | Medium | Constrain: confidence delta may only move on documented groupthink (panel sigma under threshold) or explicit missed-claim list; ban free-form "I feel +/-0.5". |
|
|
| **Pro tier marketing** currently sells no Audit; matrix row rename may confuse | Low | Feature matrix: Synthesis and Integrity (Enterprise) — capability retained, seat count honest. |
|
|
| **Low-complexity Enterprise drop list** referenced Audit explicitly | Low | Update section 3.3 text only. |
|
|
| **Outcome-tracking granularity** — cannot attribute T+90 errors to Reasoning vs Audit separately | Low | Accept; meta-tail was always under-identified. Track merged seat as one accuracy series. |
|
|
| **Over-collapse temptation** — if 15% feels good, next PR kills CC-C or Validation | Process | This lens only authorizes **logic-collinear non-scoring** merges. Ensemble scorers are out of scope. |
|
|
|
|
### Explicit non-risks
|
|
|
|
- Vendor diversity of **scoring** panel unchanged (still 8 vendors available; scoring seats still multi-lab).
|
|
- No new same-vendor Primary/meta self-audit (Kimi is not Anthropic Primary/Validation).
|
|
- Fable no longer audits after sitting Team — **contamination risk decreases**.
|
|
|
|
---
|
|
|
|
## 5. New diagram
|
|
|
|
### Before (v2.1) — 11 phases / 13 seats
|
|
|
|
```
|
|
[Proposal]
|
|
|
|
|
v
|
|
+-----------------+
|
|
| 1 Research | Grok 4.5 (no score)
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 2 Primary | Opus 5 (scores)
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 3 Validation | Sonnet 5 (scores, sees Primary)
|
|
+--------+--------+
|
|
v
|
|
+----+----+
|
|
v v v
|
|
+----+ +----+ +----+
|
|
|4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores)
|
|
+--+-+ +--+-+ +--+-+
|
|
+-----+-----+
|
|
v
|
|
+-----------------+
|
|
| 5 Financial | MiniMax-M3
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 6 Team/Founder | Fable 5
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 7 Market | Qwen3.7 Plus
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 8 Execution | Terra
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 9 Legal | Sonnet 5
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
|10 Reasoning | Kimi K2.6 (contradictions + synth)
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
|11 Audit | Fable 5 (groupthink + +/-0.5)
|
|
+--------+--------+
|
|
v
|
|
[Final Report]
|
|
```
|
|
|
|
### After (recommended) — 10 phases / 11 seats (~15% seat cut)
|
|
|
|
```
|
|
[Proposal]
|
|
|
|
|
v
|
|
+-----------------+
|
|
| 1 Research | Grok 4.5 (no score)
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 2 Primary | Opus 5 (scores)
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 3 Validation | Sonnet 5 (scores, sees Primary)
|
|
+--------+--------+
|
|
v
|
|
+----+----+
|
|
v v v
|
|
+----+ +----+ +----+
|
|
|4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores)
|
|
+--+-+ +--+-+ +--+-+
|
|
+-----+-----+
|
|
v
|
|
+-----------------+
|
|
| 5 Financial | MiniMax-M3
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 6 Team/Founder | Fable 5 (Team only — no Audit double)
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 7 Market | Qwen3.7 Plus
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 8 Execution | Terra
|
|
+--------+--------+
|
|
v
|
|
+-----------------+
|
|
| 9 Legal | Sonnet 5
|
|
+--------+--------+
|
|
v
|
|
+-------------------------------------+
|
|
|10 Synthesis and Integrity Gate | Kimi K2.6
|
|
| - contradiction / alignment map |
|
|
| - section 4.1 synthesized scores |
|
|
| - groupthink / blind-spot notes |
|
|
| - confidence delta +/-0.5 (gated) |
|
|
| F1: Fable 5 · F2: Sol |
|
|
+------------------+------------------+
|
|
v
|
|
[Final Report]
|
|
```
|
|
|
|
### Overlap evidence (why this pair, not another)
|
|
|
|
```
|
|
INPUT SCOPE
|
|
proposal scores finished-review
|
|
Primary ### . .
|
|
Validation ## ### .
|
|
Cross-Checks ### . . intentional ensemble (KEEP x3)
|
|
Financial ## ## .
|
|
Team ## ## .
|
|
Market ## ## .
|
|
Execution ## ## .
|
|
Legal ## # .
|
|
Reasoning (P10) . ### ## ]
|
|
Audit (P11) . ## ### ] SAME META QUESTION
|
|
Research tools . . (unique — KEEP)
|
|
|
|
Overlap score P10 intersect P11:
|
|
question similarity ~0.85 (contradiction equiv missed-agreement/groupthink)
|
|
input Jaccard ~0.80 (both post-domain, panel-output-primary)
|
|
output collinearity ~0.75 (both emit integrity adjustment, not domain scores)
|
|
-> COLLAPSE CANDIDATE under reasoning lens
|
|
```
|
|
|
|
### Opus empirical echo (from v2.1 review results)
|
|
|
|
Opus 5 panel-vs-single test already showed cross-checks can be **statistically tight** (market dimension spread 0.00 across three CCs). That finding pressures **ensemble scorer** count — out of scope for *this* lens. The safer 15% cut is the **duplicated meta-tail**, not the diversity scorers the product sells. Killing Audit+Reasoning redundancy reduces COGS without touching the scoring majority (still 9 scoring judges + Research + 1 meta).
|
|
|
|
**Scoring judges retained: 9** (Primary, Validation, 3xCC, Financial, Team, Market, Execution, Legal)
|
|
**Meta retained: 1** (merged) + Research
|
|
|
|
---
|
|
|
|
## Decision line
|
|
|
|
> **Collapse Reasoning-Verification and Audit Agent into one Synthesis and Integrity Gate (Kimi K2.6).**
|
|
> **-2 seats (-15.4%), -1 serial phase, ~15% meta-tail cost, cleaner vendor math, removes Fable Team-to-Audit double-duty.**
|
|
> Domain ensemble untouched. Ship as v2.2 delta on top of v2.1 blocker fixes.
|
|
|
|
---
|
|
|
|
## Implementation checklist (for conductor)
|
|
|
|
1. Patch judge-pool-spec.md sections listed in section 1
|
|
2. Bump version header to v2.2; status remains FINAL DRAFT pending 3-5 proposal gate
|
|
3. Update feature matrix + latency table
|
|
4. Rewrite section 2.3 fallback row; delete Audit row
|
|
5. Add output JSON schema for two-section merged seat
|
|
6. Re-run vendor % assertion (expect 3/11 Anthropic)
|
|
7. Do **not** change Pro 7-judge composition in this PR (Enterprise matrix row rename only)
|
|
8. Optional: 90-day 10% shadow Fable integrity pass off critical path
|