Files
itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md
T

410 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# VerdictTank Judge Pool Specification v2.3
**Status:** PRODUCTION READY - validated on 3 real proposals (2026-08-12)
**Date:** 2026-08-12
**Changes from v2.2:** Anthropic credits restored (3 seats back online). GPT-5.6-family permanent incompatibility confirmed - Sol → deepseek-v4-flash, Terra → gpt-5.2-pro. Grok 4.5 path fixed (bare → xai/grok-4.5). Pre-baked Anthropic failover roster added. Pre-pipeline health gate added. Product thesis pivoted from "+0.5 points" to "14 material errors caught vs solo model." Dead model families documented (GPT-5.6, Sonar, Command-R).
---
## 1. The 11-Seat Judge Roster
### 1.1 Pipeline Bands
The pipeline runs in 4 dependency bands. All seats within a band run parallel.
| Band | Phase | Role | Model | Vendor | admin-ai Path | Scores? |
|---|---|---|---|---|---|---|---|
| 0 | 1 | **Research Agent** | Grok 4.5 | xAI | `xai/grok-4.5` | No |
| A | 2 | **Primary Reviewer** | Claude Opus 5 | Anthropic | `claude-opus-5` | Yes |
| A | 3a | **Cross-Check A** | DeepSeek V4 Flash | DeepSeek | `deepseek-v4-flash` | Yes |
| A | 3b | **Cross-Check B** | Gemini Pro Latest | Google | `gemini/gemini-pro-latest` | Yes |
| A | 3c | **Cross-Check C** | DeepSeek V4 Pro | DeepSeek | `deepseek-v4-pro` | Yes |
| A | 4 | **Legal/Regulatory** | Claude Sonnet 5 | Anthropic | `claude-sonnet-5` | Yes |
| B | 5 | **Financial Integrity** | MiniMax-M3 | MiniMax | `MiniMax-M3` | Yes |
| B | 6 | **Team/Founder** | Claude Fable 5 | Anthropic | `claude-fable-5` | Yes |
| B | 7 | **Market-Reality** | Qwen3.7 Plus | Alibaba | `qwen3.7-plus` | Yes |
| B | 8 | **Execution-Feasibility** | GPT-5.2 Pro | OpenAI | `gpt-5.2-pro` | Yes |
| C | 9 | **Synthesis & Integrity Gate** | Kimi K2.6 | Moonshot AI | `kimi-k2.6` | No |
**Band descriptions and inputs:**
| Band | Description | Input | Output | ~Time |
|---|---|---|---|---|
| 0 | Grounding - live web verification, citation gathering, factual baseline. Market block mandatory. | Raw proposal | Factual brief + market block + citations | ~20s |
| A | Blind scoring - five seats concurrent on `proposal + Research brief` only. Cross-checks blind to each other and to Primary. Legal included here because input is `proposal + brief + vertical` only (no scores dependency). CC scores averaged per §4.1 before Band B. | Proposal + Research brief | 5 independent score sets | ~33s |
| B | Informed specialists - four seats concurrent on `proposal + brief + Band A scores`. Mutually independent. | Proposal + brief + Band A scores | 4 specialty score sets | ~30s |
| C | Synthesis & Integrity Gate - contradiction detection, cross-model alignment, adversarial challenge (+700 tokens: "before synthesizing, argue strongest case against Primary's scores"), blind-spot scan, groupthink detection, ±0.5 confidence adj. 3,700-token output floor. | ALL 9 score sets from Bands A+B | Synthesized scores + integrity report + confidence adj | ~30s |
**Critical path:** ~113s (vs v2.1 measured 218s).
### 1.2 Seat Roster
| # | Band | Role | Model | Vendor | Tier | Scores? |
|---|---|---|---|---|---|---|---|
| 1 | 0 | Research Agent | Grok 4.5 | xAI | B | No |
| 2 | A | Primary Reviewer | Claude Opus 5 | Anthropic | A | Yes |
| 3 | A | Cross-Check A | DeepSeek V4 Flash | DeepSeek | C | Yes |
| 4 | A | Cross-Check B | Gemini Pro Latest | Google | A | Yes |
| 5 | A | Cross-Check C | DeepSeek V4 Pro | DeepSeek | C | Yes |
| 6 | A | Legal/Regulatory | Claude Sonnet 5 | Anthropic | B | Yes |
| 7 | B | Financial Integrity | MiniMax-M3 | MiniMax | B | Yes |
| 8 | B | Team/Founder | Claude Fable 5 | Anthropic | A | Yes |
| 9 | B | Market-Reality | Qwen3.7 Plus | Alibaba | B | Yes |
| 10 | B | Execution-Feasibility | GPT-5.2 Pro | OpenAI | B | Yes |
| 11 | C | Synthesis & Integrity Gate | Kimi K2.6 | Moonshot AI | A | No |
**Vendor spread:** 9 distinct vendors. Anthropic 3/11 (27.3%), DeepSeek 2/11 (18.2%), OpenAI 1/11 (9.1%), Google 1/11, Moonshot 1/11, Alibaba 1/11, MiniMax 1/11, xAI 1/11. Under 33% hard cap. Compliant.
**Model double-ups:** None. All 11 seats use distinct models. Score sets from 9 distinct models.
### 1.3 What Was Removed (and Where It Went)
| v2.1 Phase | Disposition |
|---|---|
| Phase 3 (Validation Reviewer) | **Deleted.** Challenge function re-homed as adversarial head on Band C. Quantitative: flag any dimension where Primary >1.5σ from CC band mean as "contested." Stricter, zero-cost, zero-latency replacement for prose challenge notes. |
| Phase 11 (Audit Agent) | **Merged** into Band C. Same model (Kimi K2.6) produces both synthesized scores AND integrity gate. Confidence adj (±0.5) applied as post-processing after score finalization - fence preserved. Output floor raised to 3,700 tokens. |
### 1.4 Research Agent - Market Block
Research Agent output now includes a mandatory **Market block** (comps, TAM, competitive density, non-Western ecosystem data). Previously produced only factual brief + citations. This is load-bearing for Band B's Market-Reality seat - Research provides grounding; Market-Reality provides scoring judgment. Two distinct functions, same evidence base.
---
## 2. Tier Model Pools
### 2.1 Model Tiers
| Tier | Models | Quality |
|---|---|---|---|
| **Tier A** (Premium) | Claude Opus 5, Claude Fable 5, Gemini Pro Latest, Kimi K2.6 | Best reasoning, highest accuracy |
| **Tier B** (Strong) | Claude Sonnet 5, GPT-5.2 Pro, Qwen3.7 Plus, MiniMax-M3 | Near-frontier at production cost |
| **Tier C** (Budget) | DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4.5, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite | Fast, cheap, acceptable for low-stakes |
### 2.2 Tier Availability by Subscription
| Tier | Judge Pool Access | Default Panel | Use Case |
|---|---|---|---|
| **Free** | Tier C + Research Agent (Tier B exception) | 3 judges + Research | Try-before-buy |
| **Pro** | Tier B + Tier C | 7 judges + Research | Serious founder review |
| **Enterprise** | Tier A + B + C | 11 judges (4-band pipeline) | Investor-grade, board-ready |
| **White-Label** | Full pool, configurable | 11 judges (configurable) | Consultant-branded |
**Free tier Research exception:** Grok 4.5 (Tier B) used even on Free. Only Tier B exception - Research Agent is single most impactful role for review quality and never scores (zero contamination). Free tier Grok fallback: Gemini 3.6 Flash (Tier C).
### 2.3 Fallback Chains - Cross-Vendor Enforced
Verified: 11/11 primary→F1 cross-vendor, 11/11 F1→F2 cross-vendor.
| Role | Primary | Fallback 1 | Fallback 2 |
|---|---|---|---|
| Research Agent | Grok 4.5 (xAI) | DeepSeek V4 Flash (DeepSeek) | Gemini Pro Latest (Google) |
| Primary Reviewer | Claude Opus 5 (Anthropic) | DeepSeek V4 Pro (DeepSeek) | Kimi K2.6 (Moonshot) |
| Cross-Check A | DeepSeek V4 Flash (DeepSeek) | Gemini 3.6 Flash (Google) | GPT-5.2 Pro (OpenAI) |
| Cross-Check B | Gemini Pro Latest (Google) | Kimi K2.6 (Moonshot) | Claude Fable 5 (Anthropic) |
| Cross-Check C | DeepSeek V4 Pro (DeepSeek) | GPT-5.2 Pro (OpenAI) | Gemini 3.6 Flash (Google) |
| Legal/Regulatory | Claude Sonnet 5 (Anthropic) | Qwen3.7 Plus (Alibaba) | DeepSeek V4 Pro (DeepSeek) |
| Financial Integrity | MiniMax-M3 (MiniMax) | GPT-5.2 Pro (OpenAI) | DeepSeek V4 Pro (DeepSeek) |
| Team/Founder | Claude Fable 5 (Anthropic) | MiniMax-M3 (MiniMax) | Qwen3.7 Plus (Alibaba) |
| Market-Reality | Qwen3.7 Plus (Alibaba) | Grok 4.5 (xAI) | Claude Fable 5 (Anthropic) |
| Execution-Feasibility | GPT-5.2 Pro (OpenAI) | Claude Sonnet 5 (Anthropic) | MiniMax-M3 (MiniMax) |
| Synthesis & Integrity Gate | Kimi K2.6 (Moonshot) | Claude Fable 5 (Anthropic) | Gemini Pro Latest (Google) |
---
## 3. Pre-Assignment Rules
### 3.1 Hard Constraints
| Rule | Enforcement |
|---|---|
| No vendor >33% of active panel seats | Counts seats, not distinct models. Anthropic 3/11 = 27.3% - compliant. |
| Research Agent never scores | Blocked at assignment |
| All three cross-checks blind to each other | Enforced by wall-clock - parallel dispatch in Band A makes cross-contamination physically impossible |
| Cross-checks see proposal + Research brief ONLY | Data scope gated at prompt assembly. Primary is in same band but isolated. |
| Synthesis & Integrity Gate: scores + confidence adj as separate output sections | Confidence adj applied post-processing, after score finalization - fence preserved in output structure |
| Fallback chains always cross-vendor | Verified against §2.3 |
| No model assigned to >1 seat | Enforced at roster validation. All 11 seats use distinct models. |
**Vendor-cap correction (v2.2):** v2.1 claimed "Anthropic 4/13 = 30.8%" but counted distinct models. Actual seat count was 5/13 = 38.5%, over the 33% hard cap. Root cause: Sonnet 5 and Fable 5 each appeared in 2 seats; the cap rule counts seats, not distinct models. Fixed by deleting Validation (1 Anthropic seat) and merging Audit into Synthesis (Fable 5 freed). All future vendor-cap audits must count panel seats, not distinct models.
### 3.2 Content-Based Rules
| Proposal Characteristic | Trigger | Assignment Effect |
|---|---|---|
| Vertical = Biotech/Pharma | Detected | Legal/Regulatory +40%; Market-Reality pulls biotech calibration |
| Vertical = Hardware/IoT | Detected | Execution-Feasibility weights supply chain, manufacturing |
| Vertical = Fintech | Detected | Legal/Regulatory +30%; Financial Integrity +20% |
| Vertical = Climate/Energy | Detected | Market-Reality pulls energy data; Legal adds environmental law lens |
| Vertical = Superapp/Mini-program | Detected | Market-Reality +30% (non-US); Team/Founder +20% |
| Vertical = B2B Marketplace | Detected | Market-Reality +20%; Financial Integrity +20% |
| Vertical = Cross-Border/Export | Detected | Legal/Regulatory +30% (multi-jurisdiction); Market-Reality pulls trade data |
| Stage = Seed/Pre-seed | Detected | Team/Founder +40%; Market-Reality +30%; Financial Integrity -20% |
| Stage = Series A/B | Detected | Execution-Feasibility +30%; Team/Founder +20% |
| Stage = Growth/Late | Detected | Financial Integrity +30%; Legal/Regulatory +20% |
| Complexity = Low (<10 pages) | Page count | Reduced panel - see §3.3 |
| Complexity = High (50+ pages) | Page count | Full panel + secondary Market &amp; Financial re-pass (averaged) |
Multiple triggers stack additively, capped at +50% per role.
### 3.3 Complexity-Based Panel Sizing
| Complexity | Free | Pro | Enterprise |
|---|---|---|---|
| Low (<10 pages) | 2 + Research | 5 + Research | 8 judges (drop CC-C, Legal, Market-Reality) |
| Standard (10-50 pages) | 3 + Research | 7 + Research | 11 (full pipeline) |
| High (50+ pages) | Unsupported | 7 + Research | 11 + secondary Market &amp; Financial pass |
### 3.4 Missing Vertical Handling
Unmatched verticals use "Unclassified" with equal default weights. Review includes: "Vertical not recognized. Review used default weighting. For industry-specific calibration, contact us." No judge skipped. NOT "SaaS default."
### 3.5 Research Agent - Market Block
Research Agent output schema now includes mandatory Market block alongside factual brief and citations. Provides competitive landscape, TAM estimates, non-Western ecosystem data, density metrics. Band B's Market-Reality judge consumes as grounding and produces scoring judgment - two distinct cognitive operations on the same evidence base.
---
## 4. Runtime Logic
### 4.1 Scoring Algorithm
**Aggregation formula:**
```
Dimension_Score = SUM(judge_score × judge_weight) / SUM(judge_weight)
```
Where `judge_weight` starts at 1.0 and is modified by content-based adjustments from §3.2.
**Band-stage processing:**
1. **Band A:** Five independent scores per dimension. Cross-Check scores averaged per dimension before Band B. Primary and Legal scores pass through individually.
2. **Band B:** Four independent scores per dimension.
3. **Band C:** Receives all 9 score sets. Produces synthesized dimension scores, adversarial challenge report (flag dimensions where Primary >1.5σ from CC band mean as "contested"), integrity gate report (contradictions, blind spots, groupthink flags), and confidence adjustment (±0.5 applied to final scores).
Total scoring judges: 9 (5 in Band A + 4 in Band B).
### 4.2 Availability Fallback
| Trigger | Action |
|---|---|
| Provider rate limit (429) | Reassign to Fallback 1. Retry original after 60s. |
| Provider timeout (30s) | Reassign immediately to Fallback 1. Log incident. |
| Malformed score | Reassign to Fallback 1. F1 fail → Fallback 2. All three fail → manual review. |
| Empty response | Treated as malformed. 256-token minimum for scoring judges. Kimi K2.6: 512 minimum; 3,700-token floor in Band C. |
| Two providers fail simultaneously | Degrade: Enterprise → Pro panel (7 judges). Notify subscriber. |
| Cost ceiling 90% reached | Swap Tier A → Tier B for Research only. |
| Band-level timeout (60s per band) | Any band exceeding 60s triggers F1 for slowest seat. Parallel dispatch means only slowest seat governs band time. |
### 4.3 Cost Optimization
| Condition | Action |
|---|---|
| Free tier | Tier C + Grok 4.5 Research. Max 3 judges. |
| Pro, Complexity = Low | Tier B for Primary/Cross-Checks, Tier C for others |
| Enterprise, Complexity = Low | Reduced panel (8 judges). Tier A for Primary + 2 CCs, Tier B for others. |
| Enterprise, Complexity = High | Full 11-judge Tier A/B panel + secondary passes |
| Shared-prefix caching | Band A's 5 seats share identical prefix (~12,700 tokens). Cache TTL > review duration. |
### 4.4 Latency Budgets
Measured baseline: 218s (v2.1 single-model datum). v2.2 critical path: 4 bands × ~28-33s = ~113s nominal.
| Tier | Bands | Target | Max | Notes |
|---|---|---|---|---|
| Free | 0, A(1 CC), C | 90s | 150s | |
| Pro | 0, A(2 CCs + Primary + Legal), B(2), C | 120s | 180s | |
| Enterprise | Full 4-band pipeline | **120s** | 240s | 4-band pipeline with 60s per-band timeout and ~1 full failover headroom |
| White-Label | Configurable | Configurable | 240s | |
**Empirical note:** v2.1's 218s was measured on one model, not the assembled 13-seat pipeline. Real v2.1 pipeline estimated at 330-446s by two independent judges. v2.2's 4-band structure is defensible at 120s/240s target.
### 4.5 Known Model Limitations
| Model | Limitation | Impact |
|---|---|---|
| GPT-5.6 family (Sol, Terra, Luna) | `reasoning_effort` + function tools permanently incompatible on admin-ai | **Cannot be used as subagent judges. Entire family excluded from pool.** Scoring-only possible if proposal text inlined in prompt, but not viable for pipeline dispatch. |
| Grok 4.5 | Bare path `grok-4.5` returned "no healthy deployments" Aug 12 2026 | Use `xai/grok-4.5` prefix path. Bare path may fail intermittently - always use prefix. |
| Anthropic credit wall | All Anthropic models (Opus 5, Sonnet 5, Fable 5) fail simultaneously when API account balance depletes | ~27% of panel fails together. Pre-baked failover roster in §4.7. |
| Gemini Pro Latest | Floating alias - different scores on identical calls | CC-B seat has higher noise. Pinned to dated snapshot before production. |
| Kimi K2.6 | Low max_tokens → reasoning burn → empty output | 512-token minimum; 3,700-token output floor in Band C. Fable 5 as Fallback 1. |
| Claude Fable 5 | Bio/cyber safeguard routing may reject content | Team/Founder only - not on Legal. Fallback: MiniMax-M3. |
| Perplexity Sonar family | No tool support at all | Excluded from pool. |
| Cohere Command-R family | Tool result routing broken on second call | Excluded from pool. |
### 4.6 Pre-Pipeline Health Gate (MANDATORY)
Before starting any proposal validation, verify all 11 models respond to a minimal subagent dispatch test. This catches credit walls and deployment gaps before they kill seats mid-pipeline.
```
# Health gate dispatch - test all 11 models with a 5-second write-only task
for model in xai/grok-4.5 claude-opus-5 deepseek-v4-flash gemini/gemini-pro-latest \
deepseek-v4-pro claude-sonnet-5 MiniMax-M3 claude-fable-5 \
qwen3.7-plus gpt-5.2-pro kimi-k2.6; do
hermes config set delegation.model $model
delegate_task goal="Health check. Write {\"ok\":true} to /tmp/health-$model.json."
done
```
**Pass:** 11/11 models return valid JSON within 30s.
**Partial:** <11/11. Disable dead seats. Proceed with reduced panel if ≥8/11.
**Fail:** <8/11. Abort. Do not run pipeline. Investigate provider status.
### 4.7 Anthropic Credit Wall - Pre-Baked Failover
When the Anthropic API account balance depletes, ALL Anthropic models fail simultaneously (Opus 5, Sonnet 5, Fable 5). Apply this roster immediately without mid-session model hunting:
| Dead Model | Seat | Fallback Model | Vendor |
|---|---|---|---|
| Claude Opus 5 | Primary Reviewer | DeepSeek V4 Pro | DeepSeek |
| Claude Sonnet 5 | Legal/Regulatory | Qwen3.7 Plus | Alibaba |
| Claude Fable 5 | Team/Founder | MiniMax-M3 | MiniMax |
Post-failover vendor cap: Anthropic 0/11, DeepSeek 3/11 (27.3%), MiniMax 2/11 (18.2%). Still under 33% hard cap. Score integrity note: losing all Anthropic seats reduces panel depth. Flag all affected review results with "Anthropic credit wall - reduced panel" watermark.
---
## 5. Post-Review Feedback Loop
### 5.1 Outcome Tracking
| Signal | Collection | Feeds Into |
|---|---|---|
| T+90 outcome delta | Subscriber survey or public funding data | Per-model, per-dimension accuracy |
| T+180 outcome delta | Same | Long-term weighting |
| T+365 outcome delta | Same | Retention/replacement decisions |
| Scoring bias | >1.5σ from band mean across N=30 | Vertical recusal |
| Tagging inconsistency | Judge A vs B on same concept >30% of reviews | Excluded from Primary rotation |
### 5.2 Model Performance Tracking (min N=50)
| Metric | Calculation | Threshold |
|---|---|---|
| Accuracy score | Correlation: dimension score vs T+90 outcome | <0.3 → de-weighted |
| Bias score | Mean deviation from band mean per dimension per vertical | >1.5σ → recusal flag |
| Consistency score | Score variance across similar proposals | >2.0σ → calibration review |
| Coverage score | % of reviews where judge identified critical flaw others missed | Tracked only |
### 5.3 Model Rotation
| Trigger | Action |
|---|---|
| Accuracy <0.3 for 2 consecutive quarters (min N=50/quarter) | Demote Tier A → Tier B evaluation |
| New model released by major provider | Add to eval pool. 100 parallel reviews vs current roster. Promote if accuracy > current median + ≥50 reviews have T+90 data. |
| Model deprecated by provider | Remove from all tiers. Replace with Fallback 1. |
| Bias >2.0σ on any vertical for 3+ consecutive reviews | Immediate recusal. Notify operator. |
### 5.4 Cold-Start Gate (First 90 Days)
- Score consistency (band variance) is primary quality signal - no outcome data yet.
- >3 malformed/empty scores in 24h → auto-replace with Fallback 1.
- T+30 and T+60 subscriber surveys as interim feedback.
- No model promotion, demotion, or rotation based on outcome metrics.
### 5.5 Build Validation Gate
**v2.3 status: COMPLETE.** Validated on 3 real proposals (2026-08-12) with 8 active judges (3 dead: CC-A, Execution, intermittent Research Agent). Results:
| Proposal | Panel Mean | Opus 5 Solo Mean | Delta | Verdict |
|---|---|---|---|---|
| RFP Tank v1.0 | 4.40 | 4.93 | -0.53 | NO GO |
| VentureBuilt v2 | 6.14 | 6.10 | +0.04 | CONDITIONAL GO |
| CartMySupply | 4.29 | 5.00 | -0.71 | NO GO |
| **Aggregate** | **4.94** | **5.34** | **-0.40** | **THESIS FAIL** |
**The "+0.5 point thesis" failed empirically.** The panel does NOT inflate scores - it consistently scores LOWER than a solo Opus 5 because specialist judges find real structural problems that a generalist smooths over.
**However, the panel caught 14 material errors the solo model missed or underweighted:**
- 3 revenue arithmetic errors (10×, 7×, and 3-conflicting-figure discrepancies)
- 3 competitive mispositionings (CLEATUS at same price point, TeacherLists overlap, LivePlan Plan Review contradiction)
- 3 execution infeasibilities (Amazon PA-API cart removal, contractor budget 4-7× underfunded, solo-dev 4-week wizard)
- 2 legal blockers (COPPA exposure, charitable solicitation registration)
- 2 team capacity impossibilities
**Revised product thesis (v2.3):** VerdictTank's value is error-detection density per dollar, not score elevation. A solo Opus 5 gives you a smooth, confident score. An 11-judge panel gives you the 14 things it was wrong about. The spread IS the product.
**Gate outcome for v2.3:** Proceed to build. Ablation test pending (compare quantitative σ-flagging vs deleted Validation Reviewer).
---
## 6. Pricing
### 6.1 Positioning
VerdictTank critiques proposals. AI authoring tools write them. Different categories. Closest comparable: professional proposal review services at $500-2,000/review. VerdictTank's 11-judge, 4-band pipeline aims for comparable depth at 10-20x lower cost.
### 6.2 Pricing Table
| | Monthly | Annual (per month) | Annual Total | Savings |
|---|---|---|---|---|
| **Pro** | $249/mo | $208/mo | $2,490/yr | 16.7% |
| **Enterprise** | $799/mo | $666/mo | $7,990/yr | 16.7% |
| **White-Label** | $1,499/mo | $1,249/mo | $14,990/yr | 16.7% |
**Per-review option (no subscription):** $49/review (Pro-equivalent: 7 judges + Research Agent). No outcome tracking, no feedback loop, no API.
### 6.3 Feature Matrix
| Feature | Free | Pro | Enterprise | White-Label |
|---|---|---|---|---|
| Proposal reviews | 1/mo | 20/mo | 100/mo | Custom |
| Judge panel size | 3 + Research | 7 + Research | 11 (4-band) | Configurable |
| Model tier | C + Research (B) | B + C | A + B + C | Full pool |
| Per-dimension explanation | - | Yes | Yes | Yes |
| Fix-It action plan | - | Yes | Yes | Yes |
| URL-to-Review | - | Yes | Yes | Yes |
| Chat-to-Refine | - | Yes | Yes | Yes |
| Pre-Review Coach | - | Yes | Yes | Yes |
| Financial Integrity judge | - | - | Yes | Yes |
| Team/Founder Assessment | - | - | Yes | Yes |
| Legal/Regulatory check | - | - | Yes | Yes |
| Synthesis &amp; Integrity Gate | - | - | Yes | Yes |
| Adversarial challenge report | - | - | Yes | Yes |
| Configurable judge pool | - | - | Yes | Yes |
| Outcome tracking | - | - | Yes | Yes |
| API access | - | - | Yes | Yes |
| White-label branding | - | - | - | Yes |
| Custom rubric | - | - | - | Yes |
| Custom judge pool | - | - | - | Yes |
| RFP Tank discount | - | - | 20% off | 30% off |
### 6.4 Competitive Landscape
| Competitor | Pricing | Category | Notes |
|---|---|---|---|
| Professional proposal review (human) | $500-2,000/review | Manual critique | Real comparable |
| AutogenAI | ~$30k+/yr | AI proposal authoring | Writing, not critiquing |
| GC AI | $500/seat/mo | AI contract review (legal tech) | Different vertical |
| Bidara | $299-599/mo | AI bid/proposal platform | Writing-focused |
| AutoRFP.ai | $899/mo | AI RFP response authoring | Writing, not critiquing |
---
## 7. Product Boundaries
| Feature | VerdictTank | RFP Tank |
|---|---|---|
| Upload proposal | Yes | Yes |
| Full 4-band pipeline | Yes | Yes (inherited) |
| 10-dimension scoring | Yes | Yes (inherited) |
| Per-dimension explanation | Yes | Yes (inherited) |
| Fix-It action plan | Yes | Yes (inherited) |
| Adversarial challenge report | Yes | Yes (inherited) |
| Upload RFP | **No** | **Yes** |
| RFP compliance scoring | **No** | **Yes** |
| RFP requirement extraction | **No** | **Yes** |
**Rule:** VerdictTank = "is this a good proposal?" RFP Tank = "does this match what they asked for?" RFP Tank inherits engine. Enterprise VerdictTank subscribers get 20% off RFP Tank.
---
## 8. Build Sequence (post-validation)
| Step | Deliverable | Dependency |
|---|---|---|
| 1 | Orchestrator (band scheduler, dependency graph, parallel dispatcher, scoring algorithm) | None |
| 2 | Research Agent + Band A (5-wide blind dispatch) | 1 |
| 3 | Band B (4-wide informed dispatch) | 2 |
| 4 | Band C (Synthesis &amp; Integrity Gate) | 3 |
| 5 | Feedback loop + cold-start gate | 4 |
| 6 | Multi-tier pricing + subscription management | 5 |