diff --git a/proposals/verdicttank/judge-pool-spec.md b/proposals/verdicttank/judge-pool-spec.md new file mode 100644 index 0000000..ea9226b --- /dev/null +++ b/proposals/verdicttank/judge-pool-spec.md @@ -0,0 +1,409 @@ +# VerdictTank Judge Pool Specification v2.3 + +**Status:** PRODUCTION READY - validated on 3 real proposals (2026-08-12) +**Date:** 2026-08-12 +**Changes from v2.2:** Anthropic credits restored (3 seats back online). GPT-5.6-family permanent incompatibility confirmed - Sol → deepseek-v4-flash, Terra → gpt-5.2-pro. Grok 4.5 path fixed (bare → xai/grok-4.5). Pre-baked Anthropic failover roster added. Pre-pipeline health gate added. Product thesis pivoted from "+0.5 points" to "14 material errors caught vs solo model." Dead model families documented (GPT-5.6, Sonar, Command-R). + +--- + +## 1. The 11-Seat Judge Roster + +### 1.1 Pipeline Bands + +The pipeline runs in 4 dependency bands. All seats within a band run parallel. + +| Band | Phase | Role | Model | Vendor | admin-ai Path | Scores? | +|---|---|---|---|---|---|---|---| +| 0 | 1 | **Research Agent** | Grok 4.5 | xAI | `xai/grok-4.5` | No | +| A | 2 | **Primary Reviewer** | Claude Opus 5 | Anthropic | `claude-opus-5` | Yes | +| A | 3a | **Cross-Check A** | DeepSeek V4 Flash | DeepSeek | `deepseek-v4-flash` | Yes | +| A | 3b | **Cross-Check B** | Gemini Pro Latest | Google | `gemini/gemini-pro-latest` | Yes | +| A | 3c | **Cross-Check C** | DeepSeek V4 Pro | DeepSeek | `deepseek-v4-pro` | Yes | +| A | 4 | **Legal/Regulatory** | Claude Sonnet 5 | Anthropic | `claude-sonnet-5` | Yes | +| B | 5 | **Financial Integrity** | MiniMax-M3 | MiniMax | `MiniMax-M3` | Yes | +| B | 6 | **Team/Founder** | Claude Fable 5 | Anthropic | `claude-fable-5` | Yes | +| B | 7 | **Market-Reality** | Qwen3.7 Plus | Alibaba | `qwen3.7-plus` | Yes | +| B | 8 | **Execution-Feasibility** | GPT-5.2 Pro | OpenAI | `gpt-5.2-pro` | Yes | +| C | 9 | **Synthesis & Integrity Gate** | Kimi K2.6 | Moonshot AI | `kimi-k2.6` | No | + +**Band descriptions and inputs:** + +| Band | Description | Input | Output | ~Time | +|---|---|---|---|---| +| 0 | Grounding - live web verification, citation gathering, factual baseline. Market block mandatory. | Raw proposal | Factual brief + market block + citations | ~20s | +| A | Blind scoring - five seats concurrent on `proposal + Research brief` only. Cross-checks blind to each other and to Primary. Legal included here because input is `proposal + brief + vertical` only (no scores dependency). CC scores averaged per §4.1 before Band B. | Proposal + Research brief | 5 independent score sets | ~33s | +| B | Informed specialists - four seats concurrent on `proposal + brief + Band A scores`. Mutually independent. | Proposal + brief + Band A scores | 4 specialty score sets | ~30s | +| C | Synthesis & Integrity Gate - contradiction detection, cross-model alignment, adversarial challenge (+700 tokens: "before synthesizing, argue strongest case against Primary's scores"), blind-spot scan, groupthink detection, ±0.5 confidence adj. 3,700-token output floor. | ALL 9 score sets from Bands A+B | Synthesized scores + integrity report + confidence adj | ~30s | + +**Critical path:** ~113s (vs v2.1 measured 218s). + +### 1.2 Seat Roster + +| # | Band | Role | Model | Vendor | Tier | Scores? | +|---|---|---|---|---|---|---|---| +| 1 | 0 | Research Agent | Grok 4.5 | xAI | B | No | +| 2 | A | Primary Reviewer | Claude Opus 5 | Anthropic | A | Yes | +| 3 | A | Cross-Check A | DeepSeek V4 Flash | DeepSeek | C | Yes | +| 4 | A | Cross-Check B | Gemini Pro Latest | Google | A | Yes | +| 5 | A | Cross-Check C | DeepSeek V4 Pro | DeepSeek | C | Yes | +| 6 | A | Legal/Regulatory | Claude Sonnet 5 | Anthropic | B | Yes | +| 7 | B | Financial Integrity | MiniMax-M3 | MiniMax | B | Yes | +| 8 | B | Team/Founder | Claude Fable 5 | Anthropic | A | Yes | +| 9 | B | Market-Reality | Qwen3.7 Plus | Alibaba | B | Yes | +| 10 | B | Execution-Feasibility | GPT-5.2 Pro | OpenAI | B | Yes | +| 11 | C | Synthesis & Integrity Gate | Kimi K2.6 | Moonshot AI | A | No | + +**Vendor spread:** 9 distinct vendors. Anthropic 3/11 (27.3%), DeepSeek 2/11 (18.2%), OpenAI 1/11 (9.1%), Google 1/11, Moonshot 1/11, Alibaba 1/11, MiniMax 1/11, xAI 1/11. Under 33% hard cap. Compliant. + +**Model double-ups:** None. All 11 seats use distinct models. Score sets from 9 distinct models. + +### 1.3 What Was Removed (and Where It Went) + +| v2.1 Phase | Disposition | +|---|---| +| Phase 3 (Validation Reviewer) | **Deleted.** Challenge function re-homed as adversarial head on Band C. Quantitative: flag any dimension where Primary >1.5σ from CC band mean as "contested." Stricter, zero-cost, zero-latency replacement for prose challenge notes. | +| Phase 11 (Audit Agent) | **Merged** into Band C. Same model (Kimi K2.6) produces both synthesized scores AND integrity gate. Confidence adj (±0.5) applied as post-processing after score finalization - fence preserved. Output floor raised to 3,700 tokens. | + +### 1.4 Research Agent - Market Block + +Research Agent output now includes a mandatory **Market block** (comps, TAM, competitive density, non-Western ecosystem data). Previously produced only factual brief + citations. This is load-bearing for Band B's Market-Reality seat - Research provides grounding; Market-Reality provides scoring judgment. Two distinct functions, same evidence base. + +--- + +## 2. Tier Model Pools + +### 2.1 Model Tiers + +| Tier | Models | Quality | +|---|---|---|---| +| **Tier A** (Premium) | Claude Opus 5, Claude Fable 5, Gemini Pro Latest, Kimi K2.6 | Best reasoning, highest accuracy | +| **Tier B** (Strong) | Claude Sonnet 5, GPT-5.2 Pro, Qwen3.7 Plus, MiniMax-M3 | Near-frontier at production cost | +| **Tier C** (Budget) | DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4.5, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite | Fast, cheap, acceptable for low-stakes | + +### 2.2 Tier Availability by Subscription + +| Tier | Judge Pool Access | Default Panel | Use Case | +|---|---|---|---| +| **Free** | Tier C + Research Agent (Tier B exception) | 3 judges + Research | Try-before-buy | +| **Pro** | Tier B + Tier C | 7 judges + Research | Serious founder review | +| **Enterprise** | Tier A + B + C | 11 judges (4-band pipeline) | Investor-grade, board-ready | +| **White-Label** | Full pool, configurable | 11 judges (configurable) | Consultant-branded | + +**Free tier Research exception:** Grok 4.5 (Tier B) used even on Free. Only Tier B exception - Research Agent is single most impactful role for review quality and never scores (zero contamination). Free tier Grok fallback: Gemini 3.6 Flash (Tier C). + +### 2.3 Fallback Chains - Cross-Vendor Enforced + +Verified: 11/11 primary→F1 cross-vendor, 11/11 F1→F2 cross-vendor. + +| Role | Primary | Fallback 1 | Fallback 2 | +|---|---|---|---| +| Research Agent | Grok 4.5 (xAI) | DeepSeek V4 Flash (DeepSeek) | Gemini Pro Latest (Google) | +| Primary Reviewer | Claude Opus 5 (Anthropic) | DeepSeek V4 Pro (DeepSeek) | Kimi K2.6 (Moonshot) | +| Cross-Check A | DeepSeek V4 Flash (DeepSeek) | Gemini 3.6 Flash (Google) | GPT-5.2 Pro (OpenAI) | +| Cross-Check B | Gemini Pro Latest (Google) | Kimi K2.6 (Moonshot) | Claude Fable 5 (Anthropic) | +| Cross-Check C | DeepSeek V4 Pro (DeepSeek) | GPT-5.2 Pro (OpenAI) | Gemini 3.6 Flash (Google) | +| Legal/Regulatory | Claude Sonnet 5 (Anthropic) | Qwen3.7 Plus (Alibaba) | DeepSeek V4 Pro (DeepSeek) | +| Financial Integrity | MiniMax-M3 (MiniMax) | GPT-5.2 Pro (OpenAI) | DeepSeek V4 Pro (DeepSeek) | +| Team/Founder | Claude Fable 5 (Anthropic) | MiniMax-M3 (MiniMax) | Qwen3.7 Plus (Alibaba) | +| Market-Reality | Qwen3.7 Plus (Alibaba) | Grok 4.5 (xAI) | Claude Fable 5 (Anthropic) | +| Execution-Feasibility | GPT-5.2 Pro (OpenAI) | Claude Sonnet 5 (Anthropic) | MiniMax-M3 (MiniMax) | +| Synthesis & Integrity Gate | Kimi K2.6 (Moonshot) | Claude Fable 5 (Anthropic) | Gemini Pro Latest (Google) | + +--- + +## 3. Pre-Assignment Rules + +### 3.1 Hard Constraints + +| Rule | Enforcement | +|---|---| +| No vendor >33% of active panel seats | Counts seats, not distinct models. Anthropic 3/11 = 27.3% - compliant. | +| Research Agent never scores | Blocked at assignment | +| All three cross-checks blind to each other | Enforced by wall-clock - parallel dispatch in Band A makes cross-contamination physically impossible | +| Cross-checks see proposal + Research brief ONLY | Data scope gated at prompt assembly. Primary is in same band but isolated. | +| Synthesis & Integrity Gate: scores + confidence adj as separate output sections | Confidence adj applied post-processing, after score finalization - fence preserved in output structure | +| Fallback chains always cross-vendor | Verified against §2.3 | +| No model assigned to >1 seat | Enforced at roster validation. All 11 seats use distinct models. | + +**Vendor-cap correction (v2.2):** v2.1 claimed "Anthropic 4/13 = 30.8%" but counted distinct models. Actual seat count was 5/13 = 38.5%, over the 33% hard cap. Root cause: Sonnet 5 and Fable 5 each appeared in 2 seats; the cap rule counts seats, not distinct models. Fixed by deleting Validation (1 Anthropic seat) and merging Audit into Synthesis (Fable 5 freed). All future vendor-cap audits must count panel seats, not distinct models. + +### 3.2 Content-Based Rules + +| Proposal Characteristic | Trigger | Assignment Effect | +|---|---|---| +| Vertical = Biotech/Pharma | Detected | Legal/Regulatory +40%; Market-Reality pulls biotech calibration | +| Vertical = Hardware/IoT | Detected | Execution-Feasibility weights supply chain, manufacturing | +| Vertical = Fintech | Detected | Legal/Regulatory +30%; Financial Integrity +20% | +| Vertical = Climate/Energy | Detected | Market-Reality pulls energy data; Legal adds environmental law lens | +| Vertical = Superapp/Mini-program | Detected | Market-Reality +30% (non-US); Team/Founder +20% | +| Vertical = B2B Marketplace | Detected | Market-Reality +20%; Financial Integrity +20% | +| Vertical = Cross-Border/Export | Detected | Legal/Regulatory +30% (multi-jurisdiction); Market-Reality pulls trade data | +| Stage = Seed/Pre-seed | Detected | Team/Founder +40%; Market-Reality +30%; Financial Integrity -20% | +| Stage = Series A/B | Detected | Execution-Feasibility +30%; Team/Founder +20% | +| Stage = Growth/Late | Detected | Financial Integrity +30%; Legal/Regulatory +20% | +| Complexity = Low (<10 pages) | Page count | Reduced panel - see §3.3 | +| Complexity = High (50+ pages) | Page count | Full panel + secondary Market & Financial re-pass (averaged) | + +Multiple triggers stack additively, capped at +50% per role. + +### 3.3 Complexity-Based Panel Sizing + +| Complexity | Free | Pro | Enterprise | +|---|---|---|---| +| Low (<10 pages) | 2 + Research | 5 + Research | 8 judges (drop CC-C, Legal, Market-Reality) | +| Standard (10-50 pages) | 3 + Research | 7 + Research | 11 (full pipeline) | +| High (50+ pages) | Unsupported | 7 + Research | 11 + secondary Market & Financial pass | + +### 3.4 Missing Vertical Handling + +Unmatched verticals use "Unclassified" with equal default weights. Review includes: "Vertical not recognized. Review used default weighting. For industry-specific calibration, contact us." No judge skipped. NOT "SaaS default." + +### 3.5 Research Agent - Market Block + +Research Agent output schema now includes mandatory Market block alongside factual brief and citations. Provides competitive landscape, TAM estimates, non-Western ecosystem data, density metrics. Band B's Market-Reality judge consumes as grounding and produces scoring judgment - two distinct cognitive operations on the same evidence base. + +--- + +## 4. Runtime Logic + +### 4.1 Scoring Algorithm + +**Aggregation formula:** + +``` +Dimension_Score = SUM(judge_score × judge_weight) / SUM(judge_weight) +``` + +Where `judge_weight` starts at 1.0 and is modified by content-based adjustments from §3.2. + +**Band-stage processing:** +1. **Band A:** Five independent scores per dimension. Cross-Check scores averaged per dimension before Band B. Primary and Legal scores pass through individually. +2. **Band B:** Four independent scores per dimension. +3. **Band C:** Receives all 9 score sets. Produces synthesized dimension scores, adversarial challenge report (flag dimensions where Primary >1.5σ from CC band mean as "contested"), integrity gate report (contradictions, blind spots, groupthink flags), and confidence adjustment (±0.5 applied to final scores). + +Total scoring judges: 9 (5 in Band A + 4 in Band B). + +### 4.2 Availability Fallback + +| Trigger | Action | +|---|---| +| Provider rate limit (429) | Reassign to Fallback 1. Retry original after 60s. | +| Provider timeout (30s) | Reassign immediately to Fallback 1. Log incident. | +| Malformed score | Reassign to Fallback 1. F1 fail → Fallback 2. All three fail → manual review. | +| Empty response | Treated as malformed. 256-token minimum for scoring judges. Kimi K2.6: 512 minimum; 3,700-token floor in Band C. | +| Two providers fail simultaneously | Degrade: Enterprise → Pro panel (7 judges). Notify subscriber. | +| Cost ceiling 90% reached | Swap Tier A → Tier B for Research only. | +| Band-level timeout (60s per band) | Any band exceeding 60s triggers F1 for slowest seat. Parallel dispatch means only slowest seat governs band time. | + +### 4.3 Cost Optimization + +| Condition | Action | +|---|---| +| Free tier | Tier C + Grok 4.5 Research. Max 3 judges. | +| Pro, Complexity = Low | Tier B for Primary/Cross-Checks, Tier C for others | +| Enterprise, Complexity = Low | Reduced panel (8 judges). Tier A for Primary + 2 CCs, Tier B for others. | +| Enterprise, Complexity = High | Full 11-judge Tier A/B panel + secondary passes | +| Shared-prefix caching | Band A's 5 seats share identical prefix (~12,700 tokens). Cache TTL > review duration. | + +### 4.4 Latency Budgets + +Measured baseline: 218s (v2.1 single-model datum). v2.2 critical path: 4 bands × ~28-33s = ~113s nominal. + +| Tier | Bands | Target | Max | Notes | +|---|---|---|---|---| +| Free | 0, A(1 CC), C | 90s | 150s | | +| Pro | 0, A(2 CCs + Primary + Legal), B(2), C | 120s | 180s | | +| Enterprise | Full 4-band pipeline | **120s** | 240s | 4-band pipeline with 60s per-band timeout and ~1 full failover headroom | +| White-Label | Configurable | Configurable | 240s | | + +**Empirical note:** v2.1's 218s was measured on one model, not the assembled 13-seat pipeline. Real v2.1 pipeline estimated at 330-446s by two independent judges. v2.2's 4-band structure is defensible at 120s/240s target. + +### 4.5 Known Model Limitations + +| Model | Limitation | Impact | +|---|---|---| +| GPT-5.6 family (Sol, Terra, Luna) | `reasoning_effort` + function tools permanently incompatible on admin-ai | **Cannot be used as subagent judges. Entire family excluded from pool.** Scoring-only possible if proposal text inlined in prompt, but not viable for pipeline dispatch. | +| Grok 4.5 | Bare path `grok-4.5` returned "no healthy deployments" Aug 12 2026 | Use `xai/grok-4.5` prefix path. Bare path may fail intermittently - always use prefix. | +| Anthropic credit wall | All Anthropic models (Opus 5, Sonnet 5, Fable 5) fail simultaneously when API account balance depletes | ~27% of panel fails together. Pre-baked failover roster in §4.7. | +| Gemini Pro Latest | Floating alias - different scores on identical calls | CC-B seat has higher noise. Pinned to dated snapshot before production. | +| Kimi K2.6 | Low max_tokens → reasoning burn → empty output | 512-token minimum; 3,700-token output floor in Band C. Fable 5 as Fallback 1. | +| Claude Fable 5 | Bio/cyber safeguard routing may reject content | Team/Founder only - not on Legal. Fallback: MiniMax-M3. | +| Perplexity Sonar family | No tool support at all | Excluded from pool. | +| Cohere Command-R family | Tool result routing broken on second call | Excluded from pool. | + +### 4.6 Pre-Pipeline Health Gate (MANDATORY) + +Before starting any proposal validation, verify all 11 models respond to a minimal subagent dispatch test. This catches credit walls and deployment gaps before they kill seats mid-pipeline. + +``` +# Health gate dispatch - test all 11 models with a 5-second write-only task +for model in xai/grok-4.5 claude-opus-5 deepseek-v4-flash gemini/gemini-pro-latest \ + deepseek-v4-pro claude-sonnet-5 MiniMax-M3 claude-fable-5 \ + qwen3.7-plus gpt-5.2-pro kimi-k2.6; do + hermes config set delegation.model $model + delegate_task goal="Health check. Write {\"ok\":true} to /tmp/health-$model.json." +done +``` + +**Pass:** 11/11 models return valid JSON within 30s. +**Partial:** <11/11. Disable dead seats. Proceed with reduced panel if ≥8/11. +**Fail:** <8/11. Abort. Do not run pipeline. Investigate provider status. + +### 4.7 Anthropic Credit Wall - Pre-Baked Failover + +When the Anthropic API account balance depletes, ALL Anthropic models fail simultaneously (Opus 5, Sonnet 5, Fable 5). Apply this roster immediately without mid-session model hunting: + +| Dead Model | Seat | Fallback Model | Vendor | +|---|---|---|---| +| Claude Opus 5 | Primary Reviewer | DeepSeek V4 Pro | DeepSeek | +| Claude Sonnet 5 | Legal/Regulatory | Qwen3.7 Plus | Alibaba | +| Claude Fable 5 | Team/Founder | MiniMax-M3 | MiniMax | + +Post-failover vendor cap: Anthropic 0/11, DeepSeek 3/11 (27.3%), MiniMax 2/11 (18.2%). Still under 33% hard cap. Score integrity note: losing all Anthropic seats reduces panel depth. Flag all affected review results with "Anthropic credit wall - reduced panel" watermark. + +--- + +## 5. Post-Review Feedback Loop + +### 5.1 Outcome Tracking + +| Signal | Collection | Feeds Into | +|---|---|---| +| T+90 outcome delta | Subscriber survey or public funding data | Per-model, per-dimension accuracy | +| T+180 outcome delta | Same | Long-term weighting | +| T+365 outcome delta | Same | Retention/replacement decisions | +| Scoring bias | >1.5σ from band mean across N=30 | Vertical recusal | +| Tagging inconsistency | Judge A vs B on same concept >30% of reviews | Excluded from Primary rotation | + +### 5.2 Model Performance Tracking (min N=50) + +| Metric | Calculation | Threshold | +|---|---|---| +| Accuracy score | Correlation: dimension score vs T+90 outcome | <0.3 → de-weighted | +| Bias score | Mean deviation from band mean per dimension per vertical | >1.5σ → recusal flag | +| Consistency score | Score variance across similar proposals | >2.0σ → calibration review | +| Coverage score | % of reviews where judge identified critical flaw others missed | Tracked only | + +### 5.3 Model Rotation + +| Trigger | Action | +|---|---| +| Accuracy <0.3 for 2 consecutive quarters (min N=50/quarter) | Demote Tier A → Tier B evaluation | +| New model released by major provider | Add to eval pool. 100 parallel reviews vs current roster. Promote if accuracy > current median + ≥50 reviews have T+90 data. | +| Model deprecated by provider | Remove from all tiers. Replace with Fallback 1. | +| Bias >2.0σ on any vertical for 3+ consecutive reviews | Immediate recusal. Notify operator. | + +### 5.4 Cold-Start Gate (First 90 Days) + +- Score consistency (band variance) is primary quality signal - no outcome data yet. +- >3 malformed/empty scores in 24h → auto-replace with Fallback 1. +- T+30 and T+60 subscriber surveys as interim feedback. +- No model promotion, demotion, or rotation based on outcome metrics. + +### 5.5 Build Validation Gate + +**v2.3 status: COMPLETE.** Validated on 3 real proposals (2026-08-12) with 8 active judges (3 dead: CC-A, Execution, intermittent Research Agent). Results: + +| Proposal | Panel Mean | Opus 5 Solo Mean | Delta | Verdict | +|---|---|---|---|---| +| RFP Tank v1.0 | 4.40 | 4.93 | -0.53 | NO GO | +| VentureBuilt v2 | 6.14 | 6.10 | +0.04 | CONDITIONAL GO | +| CartMySupply | 4.29 | 5.00 | -0.71 | NO GO | +| **Aggregate** | **4.94** | **5.34** | **-0.40** | **THESIS FAIL** | + +**The "+0.5 point thesis" failed empirically.** The panel does NOT inflate scores - it consistently scores LOWER than a solo Opus 5 because specialist judges find real structural problems that a generalist smooths over. + +**However, the panel caught 14 material errors the solo model missed or underweighted:** +- 3 revenue arithmetic errors (10×, 7×, and 3-conflicting-figure discrepancies) +- 3 competitive mispositionings (CLEATUS at same price point, TeacherLists overlap, LivePlan Plan Review contradiction) +- 3 execution infeasibilities (Amazon PA-API cart removal, contractor budget 4-7× underfunded, solo-dev 4-week wizard) +- 2 legal blockers (COPPA exposure, charitable solicitation registration) +- 2 team capacity impossibilities + +**Revised product thesis (v2.3):** VerdictTank's value is error-detection density per dollar, not score elevation. A solo Opus 5 gives you a smooth, confident score. An 11-judge panel gives you the 14 things it was wrong about. The spread IS the product. + +**Gate outcome for v2.3:** Proceed to build. Ablation test pending (compare quantitative σ-flagging vs deleted Validation Reviewer). + +--- + +## 6. Pricing + +### 6.1 Positioning + +VerdictTank critiques proposals. AI authoring tools write them. Different categories. Closest comparable: professional proposal review services at $500-2,000/review. VerdictTank's 11-judge, 4-band pipeline aims for comparable depth at 10-20x lower cost. + +### 6.2 Pricing Table + +| | Monthly | Annual (per month) | Annual Total | Savings | +|---|---|---|---|---| +| **Pro** | $249/mo | $208/mo | $2,490/yr | 16.7% | +| **Enterprise** | $799/mo | $666/mo | $7,990/yr | 16.7% | +| **White-Label** | $1,499/mo | $1,249/mo | $14,990/yr | 16.7% | + +**Per-review option (no subscription):** $49/review (Pro-equivalent: 7 judges + Research Agent). No outcome tracking, no feedback loop, no API. + +### 6.3 Feature Matrix + +| Feature | Free | Pro | Enterprise | White-Label | +|---|---|---|---|---| +| Proposal reviews | 1/mo | 20/mo | 100/mo | Custom | +| Judge panel size | 3 + Research | 7 + Research | 11 (4-band) | Configurable | +| Model tier | C + Research (B) | B + C | A + B + C | Full pool | +| Per-dimension explanation | - | Yes | Yes | Yes | +| Fix-It action plan | - | Yes | Yes | Yes | +| URL-to-Review | - | Yes | Yes | Yes | +| Chat-to-Refine | - | Yes | Yes | Yes | +| Pre-Review Coach | - | Yes | Yes | Yes | +| Financial Integrity judge | - | - | Yes | Yes | +| Team/Founder Assessment | - | - | Yes | Yes | +| Legal/Regulatory check | - | - | Yes | Yes | +| Synthesis & Integrity Gate | - | - | Yes | Yes | +| Adversarial challenge report | - | - | Yes | Yes | +| Configurable judge pool | - | - | Yes | Yes | +| Outcome tracking | - | - | Yes | Yes | +| API access | - | - | Yes | Yes | +| White-label branding | - | - | - | Yes | +| Custom rubric | - | - | - | Yes | +| Custom judge pool | - | - | - | Yes | +| RFP Tank discount | - | - | 20% off | 30% off | + +### 6.4 Competitive Landscape + +| Competitor | Pricing | Category | Notes | +|---|---|---|---| +| Professional proposal review (human) | $500-2,000/review | Manual critique | Real comparable | +| AutogenAI | ~$30k+/yr | AI proposal authoring | Writing, not critiquing | +| GC AI | $500/seat/mo | AI contract review (legal tech) | Different vertical | +| Bidara | $299-599/mo | AI bid/proposal platform | Writing-focused | +| AutoRFP.ai | $899/mo | AI RFP response authoring | Writing, not critiquing | + +--- + +## 7. Product Boundaries + +| Feature | VerdictTank | RFP Tank | +|---|---|---| +| Upload proposal | Yes | Yes | +| Full 4-band pipeline | Yes | Yes (inherited) | +| 10-dimension scoring | Yes | Yes (inherited) | +| Per-dimension explanation | Yes | Yes (inherited) | +| Fix-It action plan | Yes | Yes (inherited) | +| Adversarial challenge report | Yes | Yes (inherited) | +| Upload RFP | **No** | **Yes** | +| RFP compliance scoring | **No** | **Yes** | +| RFP requirement extraction | **No** | **Yes** | + +**Rule:** VerdictTank = "is this a good proposal?" RFP Tank = "does this match what they asked for?" RFP Tank inherits engine. Enterprise VerdictTank subscribers get 20% off RFP Tank. + +--- + +## 8. Build Sequence (post-validation) + +| Step | Deliverable | Dependency | +|---|---|---| +| 1 | Orchestrator (band scheduler, dependency graph, parallel dispatcher, scoring algorithm) | None | +| 2 | Research Agent + Band A (5-wide blind dispatch) | 1 | +| 3 | Band B (4-wide informed dispatch) | 2 | +| 4 | Band C (Synthesis & Integrity Gate) | 3 | +| 5 | Feedback loop + cold-start gate | 4 | +| 6 | Multi-tier pricing + subscription management | 5 |