Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs - disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md - clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy) - projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture - proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review - docs/super-search/firecrawl-provider-strategy.md - updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io - .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
This commit is contained in:
@@ -0,0 +1,333 @@
|
||||
# VerdictTank Pipeline Consolidation Recommendation
|
||||
## Lens: Reasoning / Logic Overlap Detection (Contradiction Detector)
|
||||
**Author:** Kimi K2.6 (Reasoning-Verification specialist)
|
||||
**Spec:** judge-pool-spec.md v2.1
|
||||
**Target:** ~15% reduction
|
||||
**Date:** 2026-08-12
|
||||
|
||||
---
|
||||
|
||||
## Executive finding
|
||||
|
||||
**ONE recommendation:** Collapse **Phase 10 (Reasoning-Verification)** and **Phase 11 (Audit Agent)** into a single post-scoring synthesis seat: **Synthesis and Integrity Gate**.
|
||||
|
||||
These two roles ask the same meta-question with different labels:
|
||||
|
||||
| Role | Stated question | Actual job |
|
||||
|---|---|---|
|
||||
| Phase 10 Reasoning-Verification | Where do judges contradict? | Read all score sets, find disagreements, synthesize |
|
||||
| Phase 11 Audit Agent | What did the panel miss / groupthink? | Read finished review, find blind spots / groupthink, adjust confidence |
|
||||
|
||||
Both are **non-scoring meta-reviewers of the panel's own output**. Neither sees the raw proposal as primary input in a unique way that the other does not eventually consume. Both produce a **post-hoc integrity adjustment**, not an independent domain score. Under a contradiction-detection lens, they are statistically and logically collinear seats.
|
||||
|
||||
---
|
||||
|
||||
## 1. Changes
|
||||
|
||||
### Collapse
|
||||
|
||||
| Remove | Absorb into |
|
||||
|---|---|
|
||||
| Phase 10: Reasoning-Verification (standalone) | **New Phase 10: Synthesis and Integrity Gate** |
|
||||
| Phase 11: Audit Agent (standalone) | (same merged seat) |
|
||||
|
||||
### New merged role definition
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| **Phase** | 10 (final) |
|
||||
| **Role** | **Synthesis and Integrity Gate** |
|
||||
| **Model** | Kimi K2.6 (Moonshot) — keeps contradiction specialty as primary |
|
||||
| **Fallback 1** | Claude Fable 5 (Anthropic) — former Audit primary |
|
||||
| **Fallback 2** | GPT-5.6 Sol (OpenAI) — former Reasoning F2 |
|
||||
| **Scores?** | No (preserves non-scoring fence) |
|
||||
| **Input** | ALL score sets from phases 2-9 + assembled draft review |
|
||||
| **Output** | (a) Contradiction / alignment report, (b) synthesized dimension scores via section 4.1 formula, (c) blind-spot / groupthink notes, (d) single confidence adjustment +/-0.5 |
|
||||
|
||||
### Prompt contract (must stay explicit)
|
||||
|
||||
Merged prompt has **two ordered sections**, not a vague do-both:
|
||||
|
||||
1. **Contradiction pass** (former Phase 10): pairwise/cluster disagreement map across Primary, Validation, Cross-Check avg, Financial, Team, Market, Execution, Legal. Emit synthesized scores.
|
||||
2. **Integrity pass** (former Phase 11): given the synthesis above, scan for groupthink (panel sigma too low), missed domain blind spots, and apply **one** confidence adjustment (+/-0.5).
|
||||
|
||||
Single model, single call, single output schema. No second serial hop.
|
||||
|
||||
### Spec section edits required
|
||||
|
||||
| Section | Edit |
|
||||
|---|---|
|
||||
| section 1.1 Pipeline Phases | Drop row 11; rewrite Phase 10 as Synthesis and Integrity Gate; phase count **11 to 10** |
|
||||
| section 1.2 Enterprise Lineup | **13 seats to 11 seats** (12 unique models still available in pool; one fewer active seat) |
|
||||
| section 1.2 Model double-ups | Remove Fable 5: Team + Audit double-up. Fable becomes Reasoning/Audit **fallback only** (Team primary only) |
|
||||
| section 2.3 Fallback table | Delete Audit Agent row; replace Reasoning-Verification row with merged role chain above |
|
||||
| section 3.1 Hard constraints | Replace "Audit reads FINISHED review only" with "Synthesis and Integrity Gate is post-scoring, read-only, cannot open new domain scores" |
|
||||
| section 3.3 Low-complexity Enterprise | Already drops Audit — update text to drop CC-C, Team, Legal (Audit no longer separate) |
|
||||
| section 4.1 Scoring algorithm | Merge the two non-scoring bullets into one: Synthesis and Integrity Gate produces contradiction report + synthesized scores + confidence +/-0.5 |
|
||||
| section 4.2 Cost ceiling rule | Swap Tier A to B for non-scoring roles first (Research, Synthesis) — drop separate Audit |
|
||||
| section 4.4 Latency | Enterprise sequential phases lose one hop — recalibrate target downward slightly (e.g. 180s to ~165s aspirational; keep 300s max) |
|
||||
| section 6.3 Feature matrix | Replace Audit Agent (Phase 11) row with Synthesis and Integrity Gate (includes audit) — still Enterprise/WL only if desired |
|
||||
|
||||
### What this is NOT
|
||||
|
||||
- Not dropping Cross-Check C (that is diversity, not logic-overlap).
|
||||
- Not merging Primary with Validation (Validation is deliberately contaminated by Primary; different question).
|
||||
- Not merging Market with Execution (different input scopes and vertical triggers).
|
||||
- Not touching Research (unique tool surface; never scores).
|
||||
|
||||
---
|
||||
|
||||
## 2. Reduction %
|
||||
|
||||
### Seat / phase accounting (Enterprise standard panel)
|
||||
|
||||
| Metric | Before (v2.1) | After | Delta |
|
||||
|---|---|---|---|
|
||||
| Pipeline phases | 11 | **10** | -1 phase (-9.1%) |
|
||||
| Active seats | 13 | **11** | -2 seats (-15.4%) |
|
||||
| Serial post-score hops | 2 (P10 then P11) | **1** | -50% of meta-tail latency |
|
||||
| Non-scoring roles | 3 (Research, Reasoning, Audit) | **2** (Research, Synthesis) | -33% meta overhead |
|
||||
| Unique models in default lineup | 11 distinct (Sonnet x2, Fable x2) | **11 distinct** (Sonnet still x2 Validation+Legal; Fable Team-only) | pool size unchanged |
|
||||
| Anthropic share of active seats | 4/13 = 30.8% | **3/11 = 27.3%** | more headroom under 33% cap |
|
||||
|
||||
### Cost reduction (directional)
|
||||
|
||||
- Eliminate one full Tier-A (or Tier-A-fallback) completion on every Enterprise review.
|
||||
- v2.1 cited about $1.30-2.10 per standard full review — expect **~12-18% COGS cut** on the judge tail (one fewer expensive synthesis call), landing near **~15%** all-in when weighted by token volume of P10+P11 vs earlier short scoring calls.
|
||||
- Latency: remove one sequential dependency (~15-40s depending on model) from the critical path after phase 9.
|
||||
|
||||
**Headline reduction: ~15%** (seats -15.4%; cost band -12-18%; phases -9%).
|
||||
|
||||
---
|
||||
|
||||
## 3. What survives
|
||||
|
||||
### Intact domain + scoring spine
|
||||
|
||||
```
|
||||
1 Research Agent (Grok 4.5) - grounding, never scores
|
||||
2 Primary Reviewer (Claude Opus 5) - 10-dimension critique
|
||||
3 Validation Reviewer (Claude Sonnet 5) - challenges Primary
|
||||
4a Cross-Check A (GPT-5.6 Sol) - blind re-score
|
||||
4b Cross-Check B (Gemini Pro Latest) - blind re-score
|
||||
4c Cross-Check C (DeepSeek V4 Pro) - blind re-score
|
||||
5 Financial Integrity (MiniMax-M3) - cap table / unit econ
|
||||
6 Team/Founder Assessment (Claude Fable 5) - founder-market fit
|
||||
7 Market-Reality (Qwen3.7 Plus) - TAM / competitive / non-Western
|
||||
8 Execution-Feasibility (GPT-5.6 Terra) - ops / timeline realism
|
||||
9 Legal/Regulatory (Claude Sonnet 5) - multi-jurisdiction risk
|
||||
10 Synthesis and Integrity Gate (Kimi K2.6) - contradictions + audit + +/-0.5
|
||||
```
|
||||
|
||||
### Preserved invariants
|
||||
|
||||
- Research never scores
|
||||
- Three blind cross-checks, isolated from Primary/Validation
|
||||
- Domain specialists (Financial, Team, Market, Execution, Legal) remain **separate scoring questions**
|
||||
- Section 4.1 weighted aggregation formula unchanged
|
||||
- Cross-vendor fallback rule unchanged (merged chain still cross-vendor at every hop)
|
||||
- Vendor cap still satisfied (Anthropic drops to 27.3%)
|
||||
- Build validation gate (section 5.5) unchanged — still require >=0.5 panel advantage
|
||||
- Fenced post-score mode retained (merged seat cannot reopen domain scores)
|
||||
|
||||
### Why these roles were NOT collapsed
|
||||
|
||||
| Pair | Why keep separate |
|
||||
|---|---|
|
||||
| Primary vs Validation | Validation is **score-conditioned** (sees Primary). Different epistemic job than independent re-score. |
|
||||
| Validation vs Cross-Checks | CCs are blind; Validation is not. Collapsing destroys the blind control. |
|
||||
| Cross-Check A/B/C | Same question by design — but **vendor diversity is the product**. Overlap is intentional ensemble, not accidental duplication. Do not collapse under this lens. |
|
||||
| Market vs Execution | Different section 3.2 triggers, different evidence (competitive density vs ops timeline). |
|
||||
| Financial vs Execution | Cap-table math is not shipping realism. |
|
||||
| Legal vs Audit | Legal is domain-scoring on statutes; Audit was meta. Legal stays; Audit folds into Synthesis. |
|
||||
| Team vs Audit (same model Fable) | Team scores founders; Audit was meta. After merge, Fable stops double-duty — **removes same-model Team-to-Audit soft contamination path**. |
|
||||
|
||||
---
|
||||
|
||||
## 4. Risks
|
||||
|
||||
| Risk | Severity | Mitigation |
|
||||
|---|---|---|
|
||||
| **Single-call attention split** — model skimps on contradiction map OR groupthink scan | Medium | Enforce two-section JSON schema; reject output missing either contradictions[] or integrity.confidence_delta. |
|
||||
| **Loss of independent second meta-reader** — two weak meta-passes can catch what one strong pass misses | Medium | Keep dual-section prompt; optionally run Fable fallback as **shadow audit** on 10% of Enterprise reviews for 90 days (not on critical path). |
|
||||
| **K2.6 empty-output / reasoning-token burn** (section 4.5) now single-points the entire meta-tail | High | Raise min tokens to 768-1024 for merged seat; on empty go immediate F1 (Fable). Circuit breaker already in section 4.2. |
|
||||
| **Confidence +/-0.5 applied by same model that synthesized scores** — self-grading bias | Medium | Constrain: confidence delta may only move on documented groupthink (panel sigma under threshold) or explicit missed-claim list; ban free-form "I feel +/-0.5". |
|
||||
| **Pro tier marketing** currently sells no Audit; matrix row rename may confuse | Low | Feature matrix: Synthesis and Integrity (Enterprise) — capability retained, seat count honest. |
|
||||
| **Low-complexity Enterprise drop list** referenced Audit explicitly | Low | Update section 3.3 text only. |
|
||||
| **Outcome-tracking granularity** — cannot attribute T+90 errors to Reasoning vs Audit separately | Low | Accept; meta-tail was always under-identified. Track merged seat as one accuracy series. |
|
||||
| **Over-collapse temptation** — if 15% feels good, next PR kills CC-C or Validation | Process | This lens only authorizes **logic-collinear non-scoring** merges. Ensemble scorers are out of scope. |
|
||||
|
||||
### Explicit non-risks
|
||||
|
||||
- Vendor diversity of **scoring** panel unchanged (still 8 vendors available; scoring seats still multi-lab).
|
||||
- No new same-vendor Primary/meta self-audit (Kimi is not Anthropic Primary/Validation).
|
||||
- Fable no longer audits after sitting Team — **contamination risk decreases**.
|
||||
|
||||
---
|
||||
|
||||
## 5. New diagram
|
||||
|
||||
### Before (v2.1) — 11 phases / 13 seats
|
||||
|
||||
```
|
||||
[Proposal]
|
||||
|
|
||||
v
|
||||
+-----------------+
|
||||
| 1 Research | Grok 4.5 (no score)
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 2 Primary | Opus 5 (scores)
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 3 Validation | Sonnet 5 (scores, sees Primary)
|
||||
+--------+--------+
|
||||
v
|
||||
+----+----+
|
||||
v v v
|
||||
+----+ +----+ +----+
|
||||
|4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores)
|
||||
+--+-+ +--+-+ +--+-+
|
||||
+-----+-----+
|
||||
v
|
||||
+-----------------+
|
||||
| 5 Financial | MiniMax-M3
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 6 Team/Founder | Fable 5
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 7 Market | Qwen3.7 Plus
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 8 Execution | Terra
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 9 Legal | Sonnet 5
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
|10 Reasoning | Kimi K2.6 (contradictions + synth)
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
|11 Audit | Fable 5 (groupthink + +/-0.5)
|
||||
+--------+--------+
|
||||
v
|
||||
[Final Report]
|
||||
```
|
||||
|
||||
### After (recommended) — 10 phases / 11 seats (~15% seat cut)
|
||||
|
||||
```
|
||||
[Proposal]
|
||||
|
|
||||
v
|
||||
+-----------------+
|
||||
| 1 Research | Grok 4.5 (no score)
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 2 Primary | Opus 5 (scores)
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 3 Validation | Sonnet 5 (scores, sees Primary)
|
||||
+--------+--------+
|
||||
v
|
||||
+----+----+
|
||||
v v v
|
||||
+----+ +----+ +----+
|
||||
|4a | |4b | |4c | Sol / Gemini / DeepSeek (blind scores)
|
||||
+--+-+ +--+-+ +--+-+
|
||||
+-----+-----+
|
||||
v
|
||||
+-----------------+
|
||||
| 5 Financial | MiniMax-M3
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 6 Team/Founder | Fable 5 (Team only — no Audit double)
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 7 Market | Qwen3.7 Plus
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 8 Execution | Terra
|
||||
+--------+--------+
|
||||
v
|
||||
+-----------------+
|
||||
| 9 Legal | Sonnet 5
|
||||
+--------+--------+
|
||||
v
|
||||
+-------------------------------------+
|
||||
|10 Synthesis and Integrity Gate | Kimi K2.6
|
||||
| - contradiction / alignment map |
|
||||
| - section 4.1 synthesized scores |
|
||||
| - groupthink / blind-spot notes |
|
||||
| - confidence delta +/-0.5 (gated) |
|
||||
| F1: Fable 5 · F2: Sol |
|
||||
+------------------+------------------+
|
||||
v
|
||||
[Final Report]
|
||||
```
|
||||
|
||||
### Overlap evidence (why this pair, not another)
|
||||
|
||||
```
|
||||
INPUT SCOPE
|
||||
proposal scores finished-review
|
||||
Primary ### . .
|
||||
Validation ## ### .
|
||||
Cross-Checks ### . . intentional ensemble (KEEP x3)
|
||||
Financial ## ## .
|
||||
Team ## ## .
|
||||
Market ## ## .
|
||||
Execution ## ## .
|
||||
Legal ## # .
|
||||
Reasoning (P10) . ### ## ]
|
||||
Audit (P11) . ## ### ] SAME META QUESTION
|
||||
Research tools . . (unique — KEEP)
|
||||
|
||||
Overlap score P10 intersect P11:
|
||||
question similarity ~0.85 (contradiction equiv missed-agreement/groupthink)
|
||||
input Jaccard ~0.80 (both post-domain, panel-output-primary)
|
||||
output collinearity ~0.75 (both emit integrity adjustment, not domain scores)
|
||||
-> COLLAPSE CANDIDATE under reasoning lens
|
||||
```
|
||||
|
||||
### Opus empirical echo (from v2.1 review results)
|
||||
|
||||
Opus 5 panel-vs-single test already showed cross-checks can be **statistically tight** (market dimension spread 0.00 across three CCs). That finding pressures **ensemble scorer** count — out of scope for *this* lens. The safer 15% cut is the **duplicated meta-tail**, not the diversity scorers the product sells. Killing Audit+Reasoning redundancy reduces COGS without touching the scoring majority (still 9 scoring judges + Research + 1 meta).
|
||||
|
||||
**Scoring judges retained: 9** (Primary, Validation, 3xCC, Financial, Team, Market, Execution, Legal)
|
||||
**Meta retained: 1** (merged) + Research
|
||||
|
||||
---
|
||||
|
||||
## Decision line
|
||||
|
||||
> **Collapse Reasoning-Verification and Audit Agent into one Synthesis and Integrity Gate (Kimi K2.6).**
|
||||
> **-2 seats (-15.4%), -1 serial phase, ~15% meta-tail cost, cleaner vendor math, removes Fable Team-to-Audit double-duty.**
|
||||
> Domain ensemble untouched. Ship as v2.2 delta on top of v2.1 blocker fixes.
|
||||
|
||||
---
|
||||
|
||||
## Implementation checklist (for conductor)
|
||||
|
||||
1. Patch judge-pool-spec.md sections listed in section 1
|
||||
2. Bump version header to v2.2; status remains FINAL DRAFT pending 3-5 proposal gate
|
||||
3. Update feature matrix + latency table
|
||||
4. Rewrite section 2.3 fallback row; delete Audit row
|
||||
5. Add output JSON schema for two-section merged seat
|
||||
6. Re-run vendor % assertion (expect 3/11 Anthropic)
|
||||
7. Do **not** change Pro 7-judge composition in this PR (Enterprise matrix row rename only)
|
||||
8. Optional: 90-day 10% shadow Fable integrity pass off critical path
|
||||
Reference in New Issue
Block a user