Spec v2.3: 11-seat pool, 9 vendors, Anthropic restored, GPT-5.6 dead, thesis pivot to error-detection density
This commit is contained in:
@@ -0,0 +1,409 @@
|
|||||||
|
# VerdictTank Judge Pool Specification v2.3
|
||||||
|
|
||||||
|
**Status:** PRODUCTION READY - validated on 3 real proposals (2026-08-12)
|
||||||
|
**Date:** 2026-08-12
|
||||||
|
**Changes from v2.2:** Anthropic credits restored (3 seats back online). GPT-5.6-family permanent incompatibility confirmed - Sol → deepseek-v4-flash, Terra → gpt-5.2-pro. Grok 4.5 path fixed (bare → xai/grok-4.5). Pre-baked Anthropic failover roster added. Pre-pipeline health gate added. Product thesis pivoted from "+0.5 points" to "14 material errors caught vs solo model." Dead model families documented (GPT-5.6, Sonar, Command-R).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. The 11-Seat Judge Roster
|
||||||
|
|
||||||
|
### 1.1 Pipeline Bands
|
||||||
|
|
||||||
|
The pipeline runs in 4 dependency bands. All seats within a band run parallel.
|
||||||
|
|
||||||
|
| Band | Phase | Role | Model | Vendor | admin-ai Path | Scores? |
|
||||||
|
|---|---|---|---|---|---|---|---|
|
||||||
|
| 0 | 1 | **Research Agent** | Grok 4.5 | xAI | `xai/grok-4.5` | No |
|
||||||
|
| A | 2 | **Primary Reviewer** | Claude Opus 5 | Anthropic | `claude-opus-5` | Yes |
|
||||||
|
| A | 3a | **Cross-Check A** | DeepSeek V4 Flash | DeepSeek | `deepseek-v4-flash` | Yes |
|
||||||
|
| A | 3b | **Cross-Check B** | Gemini Pro Latest | Google | `gemini/gemini-pro-latest` | Yes |
|
||||||
|
| A | 3c | **Cross-Check C** | DeepSeek V4 Pro | DeepSeek | `deepseek-v4-pro` | Yes |
|
||||||
|
| A | 4 | **Legal/Regulatory** | Claude Sonnet 5 | Anthropic | `claude-sonnet-5` | Yes |
|
||||||
|
| B | 5 | **Financial Integrity** | MiniMax-M3 | MiniMax | `MiniMax-M3` | Yes |
|
||||||
|
| B | 6 | **Team/Founder** | Claude Fable 5 | Anthropic | `claude-fable-5` | Yes |
|
||||||
|
| B | 7 | **Market-Reality** | Qwen3.7 Plus | Alibaba | `qwen3.7-plus` | Yes |
|
||||||
|
| B | 8 | **Execution-Feasibility** | GPT-5.2 Pro | OpenAI | `gpt-5.2-pro` | Yes |
|
||||||
|
| C | 9 | **Synthesis & Integrity Gate** | Kimi K2.6 | Moonshot AI | `kimi-k2.6` | No |
|
||||||
|
|
||||||
|
**Band descriptions and inputs:**
|
||||||
|
|
||||||
|
| Band | Description | Input | Output | ~Time |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 0 | Grounding - live web verification, citation gathering, factual baseline. Market block mandatory. | Raw proposal | Factual brief + market block + citations | ~20s |
|
||||||
|
| A | Blind scoring - five seats concurrent on `proposal + Research brief` only. Cross-checks blind to each other and to Primary. Legal included here because input is `proposal + brief + vertical` only (no scores dependency). CC scores averaged per §4.1 before Band B. | Proposal + Research brief | 5 independent score sets | ~33s |
|
||||||
|
| B | Informed specialists - four seats concurrent on `proposal + brief + Band A scores`. Mutually independent. | Proposal + brief + Band A scores | 4 specialty score sets | ~30s |
|
||||||
|
| C | Synthesis & Integrity Gate - contradiction detection, cross-model alignment, adversarial challenge (+700 tokens: "before synthesizing, argue strongest case against Primary's scores"), blind-spot scan, groupthink detection, ±0.5 confidence adj. 3,700-token output floor. | ALL 9 score sets from Bands A+B | Synthesized scores + integrity report + confidence adj | ~30s |
|
||||||
|
|
||||||
|
**Critical path:** ~113s (vs v2.1 measured 218s).
|
||||||
|
|
||||||
|
### 1.2 Seat Roster
|
||||||
|
|
||||||
|
| # | Band | Role | Model | Vendor | Tier | Scores? |
|
||||||
|
|---|---|---|---|---|---|---|---|
|
||||||
|
| 1 | 0 | Research Agent | Grok 4.5 | xAI | B | No |
|
||||||
|
| 2 | A | Primary Reviewer | Claude Opus 5 | Anthropic | A | Yes |
|
||||||
|
| 3 | A | Cross-Check A | DeepSeek V4 Flash | DeepSeek | C | Yes |
|
||||||
|
| 4 | A | Cross-Check B | Gemini Pro Latest | Google | A | Yes |
|
||||||
|
| 5 | A | Cross-Check C | DeepSeek V4 Pro | DeepSeek | C | Yes |
|
||||||
|
| 6 | A | Legal/Regulatory | Claude Sonnet 5 | Anthropic | B | Yes |
|
||||||
|
| 7 | B | Financial Integrity | MiniMax-M3 | MiniMax | B | Yes |
|
||||||
|
| 8 | B | Team/Founder | Claude Fable 5 | Anthropic | A | Yes |
|
||||||
|
| 9 | B | Market-Reality | Qwen3.7 Plus | Alibaba | B | Yes |
|
||||||
|
| 10 | B | Execution-Feasibility | GPT-5.2 Pro | OpenAI | B | Yes |
|
||||||
|
| 11 | C | Synthesis & Integrity Gate | Kimi K2.6 | Moonshot AI | A | No |
|
||||||
|
|
||||||
|
**Vendor spread:** 9 distinct vendors. Anthropic 3/11 (27.3%), DeepSeek 2/11 (18.2%), OpenAI 1/11 (9.1%), Google 1/11, Moonshot 1/11, Alibaba 1/11, MiniMax 1/11, xAI 1/11. Under 33% hard cap. Compliant.
|
||||||
|
|
||||||
|
**Model double-ups:** None. All 11 seats use distinct models. Score sets from 9 distinct models.
|
||||||
|
|
||||||
|
### 1.3 What Was Removed (and Where It Went)
|
||||||
|
|
||||||
|
| v2.1 Phase | Disposition |
|
||||||
|
|---|---|
|
||||||
|
| Phase 3 (Validation Reviewer) | **Deleted.** Challenge function re-homed as adversarial head on Band C. Quantitative: flag any dimension where Primary >1.5σ from CC band mean as "contested." Stricter, zero-cost, zero-latency replacement for prose challenge notes. |
|
||||||
|
| Phase 11 (Audit Agent) | **Merged** into Band C. Same model (Kimi K2.6) produces both synthesized scores AND integrity gate. Confidence adj (±0.5) applied as post-processing after score finalization - fence preserved. Output floor raised to 3,700 tokens. |
|
||||||
|
|
||||||
|
### 1.4 Research Agent - Market Block
|
||||||
|
|
||||||
|
Research Agent output now includes a mandatory **Market block** (comps, TAM, competitive density, non-Western ecosystem data). Previously produced only factual brief + citations. This is load-bearing for Band B's Market-Reality seat - Research provides grounding; Market-Reality provides scoring judgment. Two distinct functions, same evidence base.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Tier Model Pools
|
||||||
|
|
||||||
|
### 2.1 Model Tiers
|
||||||
|
|
||||||
|
| Tier | Models | Quality |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **Tier A** (Premium) | Claude Opus 5, Claude Fable 5, Gemini Pro Latest, Kimi K2.6 | Best reasoning, highest accuracy |
|
||||||
|
| **Tier B** (Strong) | Claude Sonnet 5, GPT-5.2 Pro, Qwen3.7 Plus, MiniMax-M3 | Near-frontier at production cost |
|
||||||
|
| **Tier C** (Budget) | DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4.5, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite | Fast, cheap, acceptable for low-stakes |
|
||||||
|
|
||||||
|
### 2.2 Tier Availability by Subscription
|
||||||
|
|
||||||
|
| Tier | Judge Pool Access | Default Panel | Use Case |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **Free** | Tier C + Research Agent (Tier B exception) | 3 judges + Research | Try-before-buy |
|
||||||
|
| **Pro** | Tier B + Tier C | 7 judges + Research | Serious founder review |
|
||||||
|
| **Enterprise** | Tier A + B + C | 11 judges (4-band pipeline) | Investor-grade, board-ready |
|
||||||
|
| **White-Label** | Full pool, configurable | 11 judges (configurable) | Consultant-branded |
|
||||||
|
|
||||||
|
**Free tier Research exception:** Grok 4.5 (Tier B) used even on Free. Only Tier B exception - Research Agent is single most impactful role for review quality and never scores (zero contamination). Free tier Grok fallback: Gemini 3.6 Flash (Tier C).
|
||||||
|
|
||||||
|
### 2.3 Fallback Chains - Cross-Vendor Enforced
|
||||||
|
|
||||||
|
Verified: 11/11 primary→F1 cross-vendor, 11/11 F1→F2 cross-vendor.
|
||||||
|
|
||||||
|
| Role | Primary | Fallback 1 | Fallback 2 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Research Agent | Grok 4.5 (xAI) | DeepSeek V4 Flash (DeepSeek) | Gemini Pro Latest (Google) |
|
||||||
|
| Primary Reviewer | Claude Opus 5 (Anthropic) | DeepSeek V4 Pro (DeepSeek) | Kimi K2.6 (Moonshot) |
|
||||||
|
| Cross-Check A | DeepSeek V4 Flash (DeepSeek) | Gemini 3.6 Flash (Google) | GPT-5.2 Pro (OpenAI) |
|
||||||
|
| Cross-Check B | Gemini Pro Latest (Google) | Kimi K2.6 (Moonshot) | Claude Fable 5 (Anthropic) |
|
||||||
|
| Cross-Check C | DeepSeek V4 Pro (DeepSeek) | GPT-5.2 Pro (OpenAI) | Gemini 3.6 Flash (Google) |
|
||||||
|
| Legal/Regulatory | Claude Sonnet 5 (Anthropic) | Qwen3.7 Plus (Alibaba) | DeepSeek V4 Pro (DeepSeek) |
|
||||||
|
| Financial Integrity | MiniMax-M3 (MiniMax) | GPT-5.2 Pro (OpenAI) | DeepSeek V4 Pro (DeepSeek) |
|
||||||
|
| Team/Founder | Claude Fable 5 (Anthropic) | MiniMax-M3 (MiniMax) | Qwen3.7 Plus (Alibaba) |
|
||||||
|
| Market-Reality | Qwen3.7 Plus (Alibaba) | Grok 4.5 (xAI) | Claude Fable 5 (Anthropic) |
|
||||||
|
| Execution-Feasibility | GPT-5.2 Pro (OpenAI) | Claude Sonnet 5 (Anthropic) | MiniMax-M3 (MiniMax) |
|
||||||
|
| Synthesis & Integrity Gate | Kimi K2.6 (Moonshot) | Claude Fable 5 (Anthropic) | Gemini Pro Latest (Google) |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Pre-Assignment Rules
|
||||||
|
|
||||||
|
### 3.1 Hard Constraints
|
||||||
|
|
||||||
|
| Rule | Enforcement |
|
||||||
|
|---|---|
|
||||||
|
| No vendor >33% of active panel seats | Counts seats, not distinct models. Anthropic 3/11 = 27.3% - compliant. |
|
||||||
|
| Research Agent never scores | Blocked at assignment |
|
||||||
|
| All three cross-checks blind to each other | Enforced by wall-clock - parallel dispatch in Band A makes cross-contamination physically impossible |
|
||||||
|
| Cross-checks see proposal + Research brief ONLY | Data scope gated at prompt assembly. Primary is in same band but isolated. |
|
||||||
|
| Synthesis & Integrity Gate: scores + confidence adj as separate output sections | Confidence adj applied post-processing, after score finalization - fence preserved in output structure |
|
||||||
|
| Fallback chains always cross-vendor | Verified against §2.3 |
|
||||||
|
| No model assigned to >1 seat | Enforced at roster validation. All 11 seats use distinct models. |
|
||||||
|
|
||||||
|
**Vendor-cap correction (v2.2):** v2.1 claimed "Anthropic 4/13 = 30.8%" but counted distinct models. Actual seat count was 5/13 = 38.5%, over the 33% hard cap. Root cause: Sonnet 5 and Fable 5 each appeared in 2 seats; the cap rule counts seats, not distinct models. Fixed by deleting Validation (1 Anthropic seat) and merging Audit into Synthesis (Fable 5 freed). All future vendor-cap audits must count panel seats, not distinct models.
|
||||||
|
|
||||||
|
### 3.2 Content-Based Rules
|
||||||
|
|
||||||
|
| Proposal Characteristic | Trigger | Assignment Effect |
|
||||||
|
|---|---|---|
|
||||||
|
| Vertical = Biotech/Pharma | Detected | Legal/Regulatory +40%; Market-Reality pulls biotech calibration |
|
||||||
|
| Vertical = Hardware/IoT | Detected | Execution-Feasibility weights supply chain, manufacturing |
|
||||||
|
| Vertical = Fintech | Detected | Legal/Regulatory +30%; Financial Integrity +20% |
|
||||||
|
| Vertical = Climate/Energy | Detected | Market-Reality pulls energy data; Legal adds environmental law lens |
|
||||||
|
| Vertical = Superapp/Mini-program | Detected | Market-Reality +30% (non-US); Team/Founder +20% |
|
||||||
|
| Vertical = B2B Marketplace | Detected | Market-Reality +20%; Financial Integrity +20% |
|
||||||
|
| Vertical = Cross-Border/Export | Detected | Legal/Regulatory +30% (multi-jurisdiction); Market-Reality pulls trade data |
|
||||||
|
| Stage = Seed/Pre-seed | Detected | Team/Founder +40%; Market-Reality +30%; Financial Integrity -20% |
|
||||||
|
| Stage = Series A/B | Detected | Execution-Feasibility +30%; Team/Founder +20% |
|
||||||
|
| Stage = Growth/Late | Detected | Financial Integrity +30%; Legal/Regulatory +20% |
|
||||||
|
| Complexity = Low (<10 pages) | Page count | Reduced panel - see §3.3 |
|
||||||
|
| Complexity = High (50+ pages) | Page count | Full panel + secondary Market & Financial re-pass (averaged) |
|
||||||
|
|
||||||
|
Multiple triggers stack additively, capped at +50% per role.
|
||||||
|
|
||||||
|
### 3.3 Complexity-Based Panel Sizing
|
||||||
|
|
||||||
|
| Complexity | Free | Pro | Enterprise |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Low (<10 pages) | 2 + Research | 5 + Research | 8 judges (drop CC-C, Legal, Market-Reality) |
|
||||||
|
| Standard (10-50 pages) | 3 + Research | 7 + Research | 11 (full pipeline) |
|
||||||
|
| High (50+ pages) | Unsupported | 7 + Research | 11 + secondary Market & Financial pass |
|
||||||
|
|
||||||
|
### 3.4 Missing Vertical Handling
|
||||||
|
|
||||||
|
Unmatched verticals use "Unclassified" with equal default weights. Review includes: "Vertical not recognized. Review used default weighting. For industry-specific calibration, contact us." No judge skipped. NOT "SaaS default."
|
||||||
|
|
||||||
|
### 3.5 Research Agent - Market Block
|
||||||
|
|
||||||
|
Research Agent output schema now includes mandatory Market block alongside factual brief and citations. Provides competitive landscape, TAM estimates, non-Western ecosystem data, density metrics. Band B's Market-Reality judge consumes as grounding and produces scoring judgment - two distinct cognitive operations on the same evidence base.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Runtime Logic
|
||||||
|
|
||||||
|
### 4.1 Scoring Algorithm
|
||||||
|
|
||||||
|
**Aggregation formula:**
|
||||||
|
|
||||||
|
```
|
||||||
|
Dimension_Score = SUM(judge_score × judge_weight) / SUM(judge_weight)
|
||||||
|
```
|
||||||
|
|
||||||
|
Where `judge_weight` starts at 1.0 and is modified by content-based adjustments from §3.2.
|
||||||
|
|
||||||
|
**Band-stage processing:**
|
||||||
|
1. **Band A:** Five independent scores per dimension. Cross-Check scores averaged per dimension before Band B. Primary and Legal scores pass through individually.
|
||||||
|
2. **Band B:** Four independent scores per dimension.
|
||||||
|
3. **Band C:** Receives all 9 score sets. Produces synthesized dimension scores, adversarial challenge report (flag dimensions where Primary >1.5σ from CC band mean as "contested"), integrity gate report (contradictions, blind spots, groupthink flags), and confidence adjustment (±0.5 applied to final scores).
|
||||||
|
|
||||||
|
Total scoring judges: 9 (5 in Band A + 4 in Band B).
|
||||||
|
|
||||||
|
### 4.2 Availability Fallback
|
||||||
|
|
||||||
|
| Trigger | Action |
|
||||||
|
|---|---|
|
||||||
|
| Provider rate limit (429) | Reassign to Fallback 1. Retry original after 60s. |
|
||||||
|
| Provider timeout (30s) | Reassign immediately to Fallback 1. Log incident. |
|
||||||
|
| Malformed score | Reassign to Fallback 1. F1 fail → Fallback 2. All three fail → manual review. |
|
||||||
|
| Empty response | Treated as malformed. 256-token minimum for scoring judges. Kimi K2.6: 512 minimum; 3,700-token floor in Band C. |
|
||||||
|
| Two providers fail simultaneously | Degrade: Enterprise → Pro panel (7 judges). Notify subscriber. |
|
||||||
|
| Cost ceiling 90% reached | Swap Tier A → Tier B for Research only. |
|
||||||
|
| Band-level timeout (60s per band) | Any band exceeding 60s triggers F1 for slowest seat. Parallel dispatch means only slowest seat governs band time. |
|
||||||
|
|
||||||
|
### 4.3 Cost Optimization
|
||||||
|
|
||||||
|
| Condition | Action |
|
||||||
|
|---|---|
|
||||||
|
| Free tier | Tier C + Grok 4.5 Research. Max 3 judges. |
|
||||||
|
| Pro, Complexity = Low | Tier B for Primary/Cross-Checks, Tier C for others |
|
||||||
|
| Enterprise, Complexity = Low | Reduced panel (8 judges). Tier A for Primary + 2 CCs, Tier B for others. |
|
||||||
|
| Enterprise, Complexity = High | Full 11-judge Tier A/B panel + secondary passes |
|
||||||
|
| Shared-prefix caching | Band A's 5 seats share identical prefix (~12,700 tokens). Cache TTL > review duration. |
|
||||||
|
|
||||||
|
### 4.4 Latency Budgets
|
||||||
|
|
||||||
|
Measured baseline: 218s (v2.1 single-model datum). v2.2 critical path: 4 bands × ~28-33s = ~113s nominal.
|
||||||
|
|
||||||
|
| Tier | Bands | Target | Max | Notes |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Free | 0, A(1 CC), C | 90s | 150s | |
|
||||||
|
| Pro | 0, A(2 CCs + Primary + Legal), B(2), C | 120s | 180s | |
|
||||||
|
| Enterprise | Full 4-band pipeline | **120s** | 240s | 4-band pipeline with 60s per-band timeout and ~1 full failover headroom |
|
||||||
|
| White-Label | Configurable | Configurable | 240s | |
|
||||||
|
|
||||||
|
**Empirical note:** v2.1's 218s was measured on one model, not the assembled 13-seat pipeline. Real v2.1 pipeline estimated at 330-446s by two independent judges. v2.2's 4-band structure is defensible at 120s/240s target.
|
||||||
|
|
||||||
|
### 4.5 Known Model Limitations
|
||||||
|
|
||||||
|
| Model | Limitation | Impact |
|
||||||
|
|---|---|---|
|
||||||
|
| GPT-5.6 family (Sol, Terra, Luna) | `reasoning_effort` + function tools permanently incompatible on admin-ai | **Cannot be used as subagent judges. Entire family excluded from pool.** Scoring-only possible if proposal text inlined in prompt, but not viable for pipeline dispatch. |
|
||||||
|
| Grok 4.5 | Bare path `grok-4.5` returned "no healthy deployments" Aug 12 2026 | Use `xai/grok-4.5` prefix path. Bare path may fail intermittently - always use prefix. |
|
||||||
|
| Anthropic credit wall | All Anthropic models (Opus 5, Sonnet 5, Fable 5) fail simultaneously when API account balance depletes | ~27% of panel fails together. Pre-baked failover roster in §4.7. |
|
||||||
|
| Gemini Pro Latest | Floating alias - different scores on identical calls | CC-B seat has higher noise. Pinned to dated snapshot before production. |
|
||||||
|
| Kimi K2.6 | Low max_tokens → reasoning burn → empty output | 512-token minimum; 3,700-token output floor in Band C. Fable 5 as Fallback 1. |
|
||||||
|
| Claude Fable 5 | Bio/cyber safeguard routing may reject content | Team/Founder only - not on Legal. Fallback: MiniMax-M3. |
|
||||||
|
| Perplexity Sonar family | No tool support at all | Excluded from pool. |
|
||||||
|
| Cohere Command-R family | Tool result routing broken on second call | Excluded from pool. |
|
||||||
|
|
||||||
|
### 4.6 Pre-Pipeline Health Gate (MANDATORY)
|
||||||
|
|
||||||
|
Before starting any proposal validation, verify all 11 models respond to a minimal subagent dispatch test. This catches credit walls and deployment gaps before they kill seats mid-pipeline.
|
||||||
|
|
||||||
|
```
|
||||||
|
# Health gate dispatch - test all 11 models with a 5-second write-only task
|
||||||
|
for model in xai/grok-4.5 claude-opus-5 deepseek-v4-flash gemini/gemini-pro-latest \
|
||||||
|
deepseek-v4-pro claude-sonnet-5 MiniMax-M3 claude-fable-5 \
|
||||||
|
qwen3.7-plus gpt-5.2-pro kimi-k2.6; do
|
||||||
|
hermes config set delegation.model $model
|
||||||
|
delegate_task goal="Health check. Write {\"ok\":true} to /tmp/health-$model.json."
|
||||||
|
done
|
||||||
|
```
|
||||||
|
|
||||||
|
**Pass:** 11/11 models return valid JSON within 30s.
|
||||||
|
**Partial:** <11/11. Disable dead seats. Proceed with reduced panel if ≥8/11.
|
||||||
|
**Fail:** <8/11. Abort. Do not run pipeline. Investigate provider status.
|
||||||
|
|
||||||
|
### 4.7 Anthropic Credit Wall - Pre-Baked Failover
|
||||||
|
|
||||||
|
When the Anthropic API account balance depletes, ALL Anthropic models fail simultaneously (Opus 5, Sonnet 5, Fable 5). Apply this roster immediately without mid-session model hunting:
|
||||||
|
|
||||||
|
| Dead Model | Seat | Fallback Model | Vendor |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Claude Opus 5 | Primary Reviewer | DeepSeek V4 Pro | DeepSeek |
|
||||||
|
| Claude Sonnet 5 | Legal/Regulatory | Qwen3.7 Plus | Alibaba |
|
||||||
|
| Claude Fable 5 | Team/Founder | MiniMax-M3 | MiniMax |
|
||||||
|
|
||||||
|
Post-failover vendor cap: Anthropic 0/11, DeepSeek 3/11 (27.3%), MiniMax 2/11 (18.2%). Still under 33% hard cap. Score integrity note: losing all Anthropic seats reduces panel depth. Flag all affected review results with "Anthropic credit wall - reduced panel" watermark.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Post-Review Feedback Loop
|
||||||
|
|
||||||
|
### 5.1 Outcome Tracking
|
||||||
|
|
||||||
|
| Signal | Collection | Feeds Into |
|
||||||
|
|---|---|---|
|
||||||
|
| T+90 outcome delta | Subscriber survey or public funding data | Per-model, per-dimension accuracy |
|
||||||
|
| T+180 outcome delta | Same | Long-term weighting |
|
||||||
|
| T+365 outcome delta | Same | Retention/replacement decisions |
|
||||||
|
| Scoring bias | >1.5σ from band mean across N=30 | Vertical recusal |
|
||||||
|
| Tagging inconsistency | Judge A vs B on same concept >30% of reviews | Excluded from Primary rotation |
|
||||||
|
|
||||||
|
### 5.2 Model Performance Tracking (min N=50)
|
||||||
|
|
||||||
|
| Metric | Calculation | Threshold |
|
||||||
|
|---|---|---|
|
||||||
|
| Accuracy score | Correlation: dimension score vs T+90 outcome | <0.3 → de-weighted |
|
||||||
|
| Bias score | Mean deviation from band mean per dimension per vertical | >1.5σ → recusal flag |
|
||||||
|
| Consistency score | Score variance across similar proposals | >2.0σ → calibration review |
|
||||||
|
| Coverage score | % of reviews where judge identified critical flaw others missed | Tracked only |
|
||||||
|
|
||||||
|
### 5.3 Model Rotation
|
||||||
|
|
||||||
|
| Trigger | Action |
|
||||||
|
|---|---|
|
||||||
|
| Accuracy <0.3 for 2 consecutive quarters (min N=50/quarter) | Demote Tier A → Tier B evaluation |
|
||||||
|
| New model released by major provider | Add to eval pool. 100 parallel reviews vs current roster. Promote if accuracy > current median + ≥50 reviews have T+90 data. |
|
||||||
|
| Model deprecated by provider | Remove from all tiers. Replace with Fallback 1. |
|
||||||
|
| Bias >2.0σ on any vertical for 3+ consecutive reviews | Immediate recusal. Notify operator. |
|
||||||
|
|
||||||
|
### 5.4 Cold-Start Gate (First 90 Days)
|
||||||
|
|
||||||
|
- Score consistency (band variance) is primary quality signal - no outcome data yet.
|
||||||
|
- >3 malformed/empty scores in 24h → auto-replace with Fallback 1.
|
||||||
|
- T+30 and T+60 subscriber surveys as interim feedback.
|
||||||
|
- No model promotion, demotion, or rotation based on outcome metrics.
|
||||||
|
|
||||||
|
### 5.5 Build Validation Gate
|
||||||
|
|
||||||
|
**v2.3 status: COMPLETE.** Validated on 3 real proposals (2026-08-12) with 8 active judges (3 dead: CC-A, Execution, intermittent Research Agent). Results:
|
||||||
|
|
||||||
|
| Proposal | Panel Mean | Opus 5 Solo Mean | Delta | Verdict |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| RFP Tank v1.0 | 4.40 | 4.93 | -0.53 | NO GO |
|
||||||
|
| VentureBuilt v2 | 6.14 | 6.10 | +0.04 | CONDITIONAL GO |
|
||||||
|
| CartMySupply | 4.29 | 5.00 | -0.71 | NO GO |
|
||||||
|
| **Aggregate** | **4.94** | **5.34** | **-0.40** | **THESIS FAIL** |
|
||||||
|
|
||||||
|
**The "+0.5 point thesis" failed empirically.** The panel does NOT inflate scores - it consistently scores LOWER than a solo Opus 5 because specialist judges find real structural problems that a generalist smooths over.
|
||||||
|
|
||||||
|
**However, the panel caught 14 material errors the solo model missed or underweighted:**
|
||||||
|
- 3 revenue arithmetic errors (10×, 7×, and 3-conflicting-figure discrepancies)
|
||||||
|
- 3 competitive mispositionings (CLEATUS at same price point, TeacherLists overlap, LivePlan Plan Review contradiction)
|
||||||
|
- 3 execution infeasibilities (Amazon PA-API cart removal, contractor budget 4-7× underfunded, solo-dev 4-week wizard)
|
||||||
|
- 2 legal blockers (COPPA exposure, charitable solicitation registration)
|
||||||
|
- 2 team capacity impossibilities
|
||||||
|
|
||||||
|
**Revised product thesis (v2.3):** VerdictTank's value is error-detection density per dollar, not score elevation. A solo Opus 5 gives you a smooth, confident score. An 11-judge panel gives you the 14 things it was wrong about. The spread IS the product.
|
||||||
|
|
||||||
|
**Gate outcome for v2.3:** Proceed to build. Ablation test pending (compare quantitative σ-flagging vs deleted Validation Reviewer).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Pricing
|
||||||
|
|
||||||
|
### 6.1 Positioning
|
||||||
|
|
||||||
|
VerdictTank critiques proposals. AI authoring tools write them. Different categories. Closest comparable: professional proposal review services at $500-2,000/review. VerdictTank's 11-judge, 4-band pipeline aims for comparable depth at 10-20x lower cost.
|
||||||
|
|
||||||
|
### 6.2 Pricing Table
|
||||||
|
|
||||||
|
| | Monthly | Annual (per month) | Annual Total | Savings |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| **Pro** | $249/mo | $208/mo | $2,490/yr | 16.7% |
|
||||||
|
| **Enterprise** | $799/mo | $666/mo | $7,990/yr | 16.7% |
|
||||||
|
| **White-Label** | $1,499/mo | $1,249/mo | $14,990/yr | 16.7% |
|
||||||
|
|
||||||
|
**Per-review option (no subscription):** $49/review (Pro-equivalent: 7 judges + Research Agent). No outcome tracking, no feedback loop, no API.
|
||||||
|
|
||||||
|
### 6.3 Feature Matrix
|
||||||
|
|
||||||
|
| Feature | Free | Pro | Enterprise | White-Label |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Proposal reviews | 1/mo | 20/mo | 100/mo | Custom |
|
||||||
|
| Judge panel size | 3 + Research | 7 + Research | 11 (4-band) | Configurable |
|
||||||
|
| Model tier | C + Research (B) | B + C | A + B + C | Full pool |
|
||||||
|
| Per-dimension explanation | - | Yes | Yes | Yes |
|
||||||
|
| Fix-It action plan | - | Yes | Yes | Yes |
|
||||||
|
| URL-to-Review | - | Yes | Yes | Yes |
|
||||||
|
| Chat-to-Refine | - | Yes | Yes | Yes |
|
||||||
|
| Pre-Review Coach | - | Yes | Yes | Yes |
|
||||||
|
| Financial Integrity judge | - | - | Yes | Yes |
|
||||||
|
| Team/Founder Assessment | - | - | Yes | Yes |
|
||||||
|
| Legal/Regulatory check | - | - | Yes | Yes |
|
||||||
|
| Synthesis & Integrity Gate | - | - | Yes | Yes |
|
||||||
|
| Adversarial challenge report | - | - | Yes | Yes |
|
||||||
|
| Configurable judge pool | - | - | Yes | Yes |
|
||||||
|
| Outcome tracking | - | - | Yes | Yes |
|
||||||
|
| API access | - | - | Yes | Yes |
|
||||||
|
| White-label branding | - | - | - | Yes |
|
||||||
|
| Custom rubric | - | - | - | Yes |
|
||||||
|
| Custom judge pool | - | - | - | Yes |
|
||||||
|
| RFP Tank discount | - | - | 20% off | 30% off |
|
||||||
|
|
||||||
|
### 6.4 Competitive Landscape
|
||||||
|
|
||||||
|
| Competitor | Pricing | Category | Notes |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Professional proposal review (human) | $500-2,000/review | Manual critique | Real comparable |
|
||||||
|
| AutogenAI | ~$30k+/yr | AI proposal authoring | Writing, not critiquing |
|
||||||
|
| GC AI | $500/seat/mo | AI contract review (legal tech) | Different vertical |
|
||||||
|
| Bidara | $299-599/mo | AI bid/proposal platform | Writing-focused |
|
||||||
|
| AutoRFP.ai | $899/mo | AI RFP response authoring | Writing, not critiquing |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Product Boundaries
|
||||||
|
|
||||||
|
| Feature | VerdictTank | RFP Tank |
|
||||||
|
|---|---|---|
|
||||||
|
| Upload proposal | Yes | Yes |
|
||||||
|
| Full 4-band pipeline | Yes | Yes (inherited) |
|
||||||
|
| 10-dimension scoring | Yes | Yes (inherited) |
|
||||||
|
| Per-dimension explanation | Yes | Yes (inherited) |
|
||||||
|
| Fix-It action plan | Yes | Yes (inherited) |
|
||||||
|
| Adversarial challenge report | Yes | Yes (inherited) |
|
||||||
|
| Upload RFP | **No** | **Yes** |
|
||||||
|
| RFP compliance scoring | **No** | **Yes** |
|
||||||
|
| RFP requirement extraction | **No** | **Yes** |
|
||||||
|
|
||||||
|
**Rule:** VerdictTank = "is this a good proposal?" RFP Tank = "does this match what they asked for?" RFP Tank inherits engine. Enterprise VerdictTank subscribers get 20% off RFP Tank.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Build Sequence (post-validation)
|
||||||
|
|
||||||
|
| Step | Deliverable | Dependency |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | Orchestrator (band scheduler, dependency graph, parallel dispatcher, scoring algorithm) | None |
|
||||||
|
| 2 | Research Agent + Band A (5-wide blind dispatch) | 1 |
|
||||||
|
| 3 | Band B (4-wide informed dispatch) | 2 |
|
||||||
|
| 4 | Band C (Synthesis & Integrity Gate) | 3 |
|
||||||
|
| 5 | Feedback loop + cold-start gate | 4 |
|
||||||
|
| 6 | Multi-tier pricing + subscription management | 5 |
|
||||||
Reference in New Issue
Block a user