14 blind errors your solo model missed
A solo frontier model gives you a smooth, confident score. A 6-judge panel gives you the 14 things it was wrong about. VerdictTank does not sell you a higher number. It sells you the errors that number was hiding.
01The Thesis Changed, Because the Data Said So
VerdictTank v4.0 was sold on a claim we could not defend: that a multimodel panel produces a better score than a single strong model. On 2026-08-12 we ran that claim against real proposals and it failed. What we found instead is a stronger product.
What failed
The score-elevation thesis. Across three real proposals the panel mean came in 0.40 points below the solo baseline. The panel did not lift scores. On two of three proposals it pushed them down. We are publishing that result rather than burying it, because the reason it happened is the product.
What worked
Error detection density. The same panel run surfaced 14 material errors that the solo baseline missed or underweighted: revenue arithmetic that was wrong by a factor of seven, a funded direct competitor the solo pass never named, a launch-blocking compliance cost larger than projected first-year revenue. None of those show up as a score. All of them decide whether the proposal wins.
You buy it to find the $60K compliance hole and the broken contractor budget before you ship.
Why the spread is the signal
A single model scoring alone produces low variance. It reads the document once, forms one coherent opinion, and every dimension it emits is downstream of that opinion. The result feels authoritative precisely because nothing inside it disagrees.
A 6-seat panel of five different vendors cannot produce that coherence, and the incoherence is diagnostic. When a Financial Integrity judge scores a proposal 8.2 while an Execution Feasibility judge scores the same document 2.8, that 5.4-point spread is not noise. It is a precise statement: the money works, the delivery plan does not. A solo model averages that tension away into a single confident 6.1 and tells you nothing actionable.
02How We Found This: Four Generations and a Self-Review
VerdictTank is a proposal review engine, not a proposal writer. It ingests a finished document and returns scored dimensions plus a ranked list of concrete Fix-It items. The pipeline was not designed in the abstract. It was hardened across four architectural generations, and then it was pointed at its own proposals.
Single-Model Scorer
v2 established the core insight: a proposal has two independent quality axes. Narrative quality (clarity, structure, persuasion) and compliance quality (does it actually answer the scored requirements). One model scored both from one prompt. It proved the concept and exposed the flaw: the axes bled together. A beautifully written section that missed a mandatory requirement scored too high, because the same reasoning pass that admired the prose also graded the compliance.
Separated Scoring Passes
v3 split scoring into two independent passes with two purpose-built prompts. The narrative pass never sees the compliance rubric. The compliance pass never rewards eloquence. This is the decision that makes the dual score trustworthy: the two numbers can now disagree, and their disagreement carries information. A 9/10 narrative next to a 4/10 compliance is a proposal about to lose.
Multimodel Adversarial Review
A single model scoring in isolation is confidently wrong at a predictable rate. v4 introduced a multimodel pipeline: a fast model produces first-pass scores and Fix-It candidates, then a stronger model reviews that output adversarially, challenging every deduction and confirming each Fix-It maps to real proposal text. Scores stopped drifting between runs. This is the generation that made the output defensible.
Specialist Panel and the Integrity Gate
v5 replaces the adversarial pair with a 6-seat specialist panel across five vendors, and adds a synthesis seat whose only job is to compute panel statistics, flag scores more than 1.5 standard deviations from the mean, and reconcile the verdict against the evidence. The output is no longer a number. It is a number, a spread, an outlier list, and a ranked set of material errors with the judge that caught each one.
The self-review that broke the old thesis
Before selling a review engine we ran the engine on our own work. We assembled the panel and scored three real proposals, in full, with the same prompts and rubric a paying customer would get. One of the three was VerdictTank's own sibling product. The panel returned a NO GO on it.
That run executed on the v2.2 roster. It cost roughly $150 in inference across the three proposals and returned 8 reporting seats. Two seats were lost to a provider credit wall hit mid-run, one to a model family that could not be dispatched at all, and a meaningful share of the spend went to retries against models that were already dead. The incomplete panel is why v2.3 of the judge pool spec introduces a pre-flight health gate, a pre-baked failover roster, and a consolidated 6-seat roster, all covered in section 5. The results below are what those 8 v2.2 seats actually produced. We report them as measured and do not extrapolate them onto the 6-seat roster.
03Validation Run: 3 Real Proposals, 8 Reporting Seats
Every figure in this section comes from the 2026-08-12 validation run. Nothing is modeled, projected, or illustrative. The solo baseline is Claude Opus 5 scoring the same documents against the same 10-dimension rubric.
Panel mean vs solo baseline
| Proposal | Panel mean | Solo baseline | Delta | Panel spread | Solo spread | Verdict |
|---|---|---|---|---|---|---|
| RFP Tank v1.0 | 4.40 | 4.93 | -0.53 | 3.7 | 1.4 | NO GO |
| VentureBuilt v2 | 6.14 | 6.10 | +0.04 | 5.4 | 2.4 | CONDITIONAL GO |
| CartMySupply | 4.29 | 5.00 | -0.71 | 2.6 | 1.8 | NO GO |
| Aggregate | 4.94 | 5.34 | -0.40 | Panel spread exceeded solo spread on all 3 | THESIS FAIL | |
Thesis under test: the panel must show a greater than 0.5 point advantage over the solo mean to justify premium pricing. Result: FAIL on all three proposals and FAIL in aggregate. Panel composition for this run was 8 reporting judges (4 Band A, 4 Band B) on the v2.2 roster. The v2.3 roster documented in section 5 consolidates to 6 seats.
Why the panel scored lower
The panel does not elevate scores. It sharpens error detection, and error detection on a flawed document moves the number down. All three proposals contained severe cross-cutting defects that additional specialist scrutiny exposed more precisely: fatal execution gaps, competitive mispositioning, and legal blockers. The solo baseline was directionally correct on all three. The panel added precision, not points.
That is the entire finding, and it inverts the sales pitch. If your proposal is sound, the panel will roughly agree with a good solo model and cost you more. If your proposal has a hole in it, the panel finds the hole and the solo model does not. You are not buying a score. You are buying the probability that a specific, expensive, named mistake gets caught before an evaluator or an investor finds it for you.
Specialist divergence, measured
| Observation | Evidence from the run | What it means |
|---|---|---|
| Generalist seats run optimistic | Gemini Pro scored RFP Tank 6.7 as a Band A generalist and 3.0 as the Band B Market specialist. Same model, same document, 3.7 points apart. | Band A generalist scoring without specialist cross-check is systematically over-optimistic. The role, not the model, drives the score. |
| Role divergence beats model divergence | DeepSeek V4 Pro scored VentureBuilt 6.4 as Cross-Check C and 2.8 as Execution Feasibility. A 3.6 point split inside one vendor. | Panel diversity is not primarily about buying different vendors. It is about buying different questions. |
| One seat can flip a verdict | Remove the 2.8 Execution score from VentureBuilt and the panel averages 6.6, reading as a clean GO. With it, the verdict is CONDITIONAL GO with a named contractor-budget fix. | The lowest score in the panel is frequently the only one doing work. Averaging is what a solo model already does. |
| Tight clustering is also a signal | CartMySupply produced zero outliers beyond 1.5 sigma and the tightest spread of the three (sigma 0.89). | Unanimity across independent vendors on a low score is a far stronger NO GO than one model's low score. |
04The 14 Errors: Every One Named
This is the product. Fourteen material errors the 8-judge panel caught that the solo baseline missed or underweighted, grouped by failure class. Each is a real finding from the 2026-08-12 run against a real document.
RFP Tank v1.0 · panel 4.40 vs solo 4.93
| Class | Error the panel caught | Caught by |
|---|---|---|
| Revenue | Three mutually inconsistent Year-1 revenue figures inside one document: $1.2M, $1.361M, and $372K. Plus a 22% MRR ramp inconsistency the narrative never reconciles. | Financial Integrity |
| Competitive | CLEATUS is a real, funded competitor at $4M seed with public product-led pricing of $39 to $250/mo, occupying the identical quadrant. The proposal does not name it. The Band A generalist seat actually cited CLEATUS pricing as a positive signal. | Market Reality |
| Competitive | GovEagle pricing referenced at a 15x inconsistency against the proposal's own comparison table. | Market Reality |
| Execution | Five of seven features marked TO BUILD at HIGH effort. The real-time Compliance Copilot alone needs 2 to 3 developers for 8 to 12 weeks. The plan allocates 4 weeks, solo. | Execution Feasibility |
| Team | A solo founder shipping a 7-feature AI SaaS in 10 weeks, with hiring contingent on revenue that requires the product to already exist. A closed loop with no entry point. | Team / Founder |
| Legal | No privacy policy and no terms of service, against FAR and CUI exposure, with ITAR implications on German-hosted infrastructure. | Legal / Regulatory |
Panel verdict: NO GO. Estimated rework 40+ hours. Recommendation is to cut scope to two features, extend to 20 weeks, hire a second developer before month one, rebuild the financial model, and address CLEATUS directly.
VentureBuilt v2 · panel 6.14 vs solo 6.10
| Class | Error the panel caught | Caught by |
|---|---|---|
| Execution | Contractor budget broken by a factor of 4 to 7. The stated $1,500/mo implies $11 to $22 per hour against a market rate of $75 to $100. At real rates that budget buys 105 to 140 hours and leaves roughly 800 hours on the founder. | Execution Feasibility |
| Revenue | Year 2 stated on a run-rate basis rather than recognized revenue. Restated correctly, the healthy scenario loses roughly $11K to $18K. | Financial Integrity |
| Competitive | The uniqueness claim is contradicted by shipping products. LivePlan Plan Review and IdeaProof already occupy the space. | Market Reality |
| Team | 37 engagements plus 950 hours plus an MSP day job. The three commitments cannot coexist in one calendar. | Team / Founder |
Panel verdict: CONDITIONAL GO with six named conditions. Estimated rework 15 to 20 hours. This is the case that most clearly shows the value: the panel mean (6.14) and the solo mean (6.10) are statistically tied, so on score alone the panel added nothing. What it added was a bimodal split, Financial 8.2 against Execution 2.8, and the four errors above.
CartMySupply · panel 4.29 vs solo 5.00
| Class | Error the panel caught | Caught by |
|---|---|---|
| Revenue | The $2.7M headline is wrong by 7x to 10x against the proposal's own inputs, which compute to $269K. Stripe fees understated by roughly $11K per year. CAC absent entirely. | Financial Integrity |
| Competitive | TeacherLists already solves the identical problem, free, across 2 million lists. Target ships native School List Assist. No technical moat is claimed or demonstrable. | Market Reality |
| Execution | Amazon PA-API 5 removed Cart API support. Target has no self-serve multi-item cart API. Walmart requires separate catalog matching. The core mechanic of the product does not have a supported integration path at any of the three named retailers. | Execution Feasibility |
| Legal | Charitable solicitation registration required in 40+ states at $30K to $75K, plus COPPA exposure and FTC penalty risk. Compliance cost of $60K to $150K exceeds projected Year-1 revenue of $3K to $14K by an order of magnitude. | Legal / Regulatory |
Panel verdict: NO GO, unanimous, zero outliers, tightest spread of the three. Estimated rework 60+ hours. The build estimate of 116 hours was independently judged 4x to 10x too low.
05Judge Pool v2.3: 6 Seats, 5 Vendors, Zero Double-Ups
The panel that produced the validation data ran 8 reporting seats on the v2.2 roster, after two seats hit a provider credit wall mid-run and one model family could not be dispatched at all. v2.3 is the spec written in response to that failure. It cuts the roster to 6 seats, keeps every distinct question the validation run proved was load-bearing, and drops the redundant generalist cross-checks that contributed correlated opinions rather than new findings. Full detail lives in the judge pool specification v2.3.
The roster
| Band | Seat | Model | Vendor | Scores |
|---|---|---|---|---|
| 0 | Research Agent | Grok 4.5 | xAI | No |
| A | Primary Reviewer | Claude Opus 5 | Anthropic | Yes |
| A | Cross-Check | DeepSeek V4 Pro | DeepSeek | Yes |
| A | Legal + Compliance | Claude Sonnet 5 | Anthropic | Yes |
| B | Financial + Market | MiniMax-M3 | MiniMax | Yes |
| C | Synthesis + Gate | Kimi K2.6 | Moonshot AI | No |
Five distinct vendors across six seats. Four scoring seats, four distinct scoring models. Zero model double-ups: no single model occupies two scoring seats, which is the constraint that keeps correlated failure out of the panel mean. Maximum vendor concentration is Anthropic at 2 of 6 (33%), inside the 40% ceiling. Every other vendor holds exactly one seat. The Research Agent and the Synthesis seat do not score, so the panel mean is computed from four independent specialist verdicts across four vendors.
The 6-seat flow
The Free tier runs a reduced 4-seat version of this flow: Research Agent, Primary Reviewer, Legal + Compliance, Synthesis + Gate. It drops Cross-Check and Financial + Market. That configuration still catches legal and compliance blockers, which was the single highest-value error class in validation, so a free review proves the concept on the errors that matter most without carrying the full panel cost.
What changed in v2.3
Pre-flight health gate
Before any scoring begins, the orchestrator pings every rostered model with a 5-second probe and writes the result to a per-model health file. Any model returning HTTP 400, HTTP 429, or a no-healthy-deployments error is swapped for its pre-assigned failover before a single scoring call is spent. The 2026-08-12 run burned roughly 12 dispatches discovering dead models at runtime. That failure mode is now closed.
Pre-baked failover roster
Every seat carries a named failover from a different vendor, resolved at gate time rather than improvised mid-run. Failover selection preserves both the vendor-diversity ceiling and the no-double-up rule, so a degraded panel is still a valid panel rather than an accidentally correlated one.
Credit-wall resilience
The provider credit exhaustion that cost two seats mid-run is now detected at the gate and treated as an availability failure, not an error. Anthropic capacity has been restored and the Primary Reviewer seat runs Claude Opus 5 as specified.
Permanent exclusions
One frontier model family proved structurally incapable of running as a panel seat under our orchestration and is permanently excluded from the roster, not merely deprioritized. Excluded models cannot be selected as a failover target either.
Latency
v2.3 targets a critical path of roughly 113 seconds, against 218 seconds measured on the v2.1 architecture. The improvement comes from band parallelism: Band A and Band B seats execute concurrently rather than sequentially, and the Synthesis seat is the only stage that must wait for all scoring seats to return.
The ten scored dimensions
Every scoring seat rates the proposal 1 to 10 on the same ten dimensions, so panel spread is computed dimension by dimension and not only in aggregate:
The Synthesis and Integrity Gate
The Band C seat never scores. It reads all scoring output and performs a fixed checklist: verify score arithmetic, compute panel means and per-dimension spread, compute the delta against the solo baseline, flag every score more than 1.5 standard deviations from the panel mean with a written rationale, and confirm the verdict follows from panel evidence rather than from the Primary Reviewer alone. That gate is what turns four scored opinions into one auditable report.
06Worked Example: VentureBuilt v2, Where the Score Said Nothing
This is the clearest case in the validation set, because it is the one where score elevation delivered exactly zero and error detection delivered everything. Real scores from the 2026-08-12 run.
Median 6.45 · standard deviation 1.758 · spread 5.4 (min 2.8, max 8.2)
Spread 2.4 · delta +0.04 · statistically tied with the panel
Individual seat scores
| Seat | Model | Score | Sigma from mean | Flag |
|---|---|---|---|---|
| Band B · Financial Integrity | MiniMax | 8.2 | +1.17 | Within 1.5 sigma |
| Band B · Market Reality | Gemini | 7.7 | +0.89 | Within 1.5 sigma |
| Band A · Primary Reviewer | Opus 5 | 6.9 | +0.43 | Within 1.5 sigma |
| Band A · Legal / Regulatory | Qwen | 6.5 | +0.21 | Within 1.5 sigma |
| Band A · Cross-Check C | DeepSeek V4 Pro | 6.4 | +0.15 | Within 1.5 sigma |
| Band A · Cross-Check B | Gemini | 6.2 | +0.04 | Within 1.5 sigma |
| Band B · Team / Founder | Kimi | 4.4 | -0.99 | Within 1.5 sigma |
| Band B · Execution Feasibility | DeepSeek V4 Pro | 2.8 | -1.91 | OUTLIER |
What the outlier actually found
The 2.8 was not a grumpy model. The Integrity Gate challenged it at 1.91 sigma and it survived the challenge on evidence: a contractor budget broken 4x to 7x, an architecture that regressed from v1 with no schema and no API contract, and a founder workload of 37 engagements plus 950 hours alongside an MSP day job. The Team seat (4.4) and the Band A generalists (6.2 to 6.9) all acknowledged the same workload problem. They weighted it less severely. The Execution specialist is the only seat that forced it into the verdict.
Fix-It items, ranked by materiality
-
CriticalExecutionFix the contractor budget or cut the scope. $1,500/mo buys 105 to 140 hours at market rates, not the volume the plan assumes. Raise to roughly $7,500/mo or reduce scope to fit the hours actually purchased.
-
CriticalFinancialRestate Year 2 on a recognized-revenue basis. On run-rate the year looks healthy. On recognized revenue it loses roughly $11K to $18K. Present both.
-
HighCompetitiveWithdraw or qualify the uniqueness claim. LivePlan Plan Review and IdeaProof already ship in this space. Reposition on a defensible axis.
-
HighTeamName the contractor and the sourcing plan before Phase 2. A budget line with no named person is not a capacity plan.
-
MediumGo-to-marketMap the Year 1 to Year 2 GTM bridge. Eight net-new signups per month appear in the model with no acquisition mechanism behind them.
-
MediumLegalComplete data protection and trademark clearance before Phase 0 to 1.
Panel verdict: CONDITIONAL GO. Estimated rework 15 to 20 hours. The solo baseline returned a 6.10 and none of the six conditions above.
07Pricing: Priced Per Error Found, Not Per Point Gained
Four tiers with declared review quantities. No asterisks, no fair-use clauses, no metered surprises. Every tier states exactly how many reviews it includes and exactly what an extra review costs. The pricing logic follows the revised thesis directly: a panel run is worth what a caught error is worth, and a caught error is worth far more than a point of score.
Free
- 1 review per month
- 4-seat reduced panel
- Top 3 Fix-It items
- Panel score and spread
- Catches legal and compliance blockers
Pro
- 5 reviews per month
- Full 6-seat panel
- Full Fix-It list, ranked
- Re-score loop with before and after
- Pre-Review Coach
- Extra reviews $15 each
- Annual billing $207/mo
Enterprise
- 30 reviews per month
- Full 6-seat panel
- Full Fix-It list, ranked
- Re-score loop and Pre-Review Coach
- Branded white-label
- Multi-seat workspaces
- Shared corpus isolation
- Configurable judge pool
- Extra reviews $15 each
- Annual billing $666/mo
White-Label
- 50 reviews per month
- Configurable panel
- Full Fix-It list, ranked
- Re-score loop and Pre-Review Coach
- White-label on your own domain
- Multi-seat workspaces
- Dedicated corpus isolation
- Full custom judge pool
- Reseller model: re-bill $200-500 each
- Extra reviews $10 each
- Annual billing $1,249/mo
Full tier comparison
| Free | Pro | Enterprise | White-Label | |
|---|---|---|---|---|
| Price | Free | $249/mo | $799/mo | $1,499/mo |
| Reviews per month | 1 | 5 | 30 | 50 |
| Overage | Not available | $15/review | $15/review | $10/review |
| Panel | 4-seat reduced | Full 6-seat | Full 6-seat | Configurable |
| Fix-Its | Top 3 | Full, ranked | Full, ranked | Full, ranked |
| Re-score loop | Not included | Yes | Yes | Yes |
| Pre-Review Coach | Not included | Yes | Yes | Yes |
| White-label | Not included | Not included | Branded only | Full domain |
| Workspaces | Not included | Not included | Multi-seat | Multi-seat |
| Corpus isolation | Not included | Not included | Shared | Dedicated |
| Judge pool config | Not included | Not included | Yes | Full custom |
| Reseller model | Not included | Not included | Not included | Re-bill $200-500/ea |
| Annual billing (16.7% off) | Not applicable | $207/mo | $666/mo | $1,249/mo |
Declared quantities only. When a tier is exhausted the customer either buys overage at the published per-review rate or waits for the next cycle. Nothing is throttled silently and no tier promises capacity it cannot cost out, because a panel review has a real marginal cost and pretending otherwise is how usage-based products lose money.
What a review costs us, and why the panel is affordable
The v2.2 validation run cost approximately $150 in inference for three full proposals across eight reporting seats. That figure includes retries against dead models before the health gate existed, which is exactly the waste v2.3 was written to remove. It is the honest anchor, and it is deliberately the worst number we have.
A clean run on the consolidated 6-seat roster, with the pre-flight health gate preventing wasted dispatches and only four seats actually scoring, costs $5.20 per review. That is the number every tier below is built on.
Unit economics at declared quantities
| Tier | Reviews included | COGS at $5.20/review | Revenue | Gross margin |
|---|---|---|---|---|
| Pro | 5 | $26 | $249 | 90% |
| Enterprise | 30 | $156 | $799 | 80% |
| White-Label | 50 | $260 | $1,499 | 83% |
| Overage, Pro and Enterprise | per review | $5.20 | $15.00 | Roughly 3x COGS |
Every declared quantity is margin-positive at full consumption, and so is every overage unit. The $15 overage prices at roughly 3x COGS. The $10 White-Label overage prices at roughly 2x COGS, which is the deliberate discount that makes the reseller math work. There is no consumption pattern inside these tiers that produces a negative unit, which is the whole reason every tier on this page ships a counted quantity and a published overage rate.
The reason a full 6-seat panel fits a $249 tier at five reviews per month is vendor mix and seat discipline. Only the Primary Reviewer runs a premium frontier model. The remaining scoring seats run strong mid-tier models from three different vendors, which is where the error-detection value came from in validation. Consolidating the roster down to six seats removed the redundant generalist cross-checks, not the specialists. Panel diversity is cheaper than panel depth, and diversity is what caught the 14.
Why the value question is not the score question
| Error class | Real example from validation | Cost of missing it |
|---|---|---|
| Legal blocker | Charitable solicitation registration in 40+ states | $30K to $75K of registration, against $3K to $14K of projected revenue |
| Compliance total | Full first-year compliance load on the same proposal | $60K to $150K, exceeding Year-1 revenue by roughly 10x |
| Execution gap | Contractor budget short by 4x to 7x | Roughly 800 unbudgeted founder hours |
| Revenue arithmetic | $2.7M headline against $269K computed from the document's own inputs | Credibility with any investor who checks the math, which is all of them |
| Competitive blind spot | A $4M-seed funded direct rival never named in the document | The first question in the room, unanswered |
A single caught item in the top two rows pays for a decade of the Pro tier. That is the entire pricing argument, and it does not depend on the panel producing a higher score, which it does not. Note that the two highest-value rows are both legal and compliance findings, which is precisely why the Free tier keeps the Legal + Compliance seat.
Positioned against the authoring category
| Comparison | Their price | VerdictTank | Multiple |
|---|---|---|---|
| Pro vs Bidara Starter | $499/mo | $249/mo | 2.0x less |
| Pro vs AutoRFP.ai Scale | $899/mo | $249/mo | 3.6x less |
| Enterprise vs AutogenAI | $30K+/yr custom | $799/mo ($9,588/yr) | 3.1x less annualized |
| Enterprise vs Bidara Starter | $499/mo | $799/mo | 1.6x more |
| Enterprise vs AutoRFP.ai Scale | $899/mo | $799/mo | 1.1x less |
| White-Label vs AutogenAI | $30K+/yr custom | $1,499/mo ($17,988/yr) | 1.7x less annualized |
We are not a proposal team in a box. We are one high-value pass in the workflow. A buyer already spending $499 to $899 per month on an authoring tool should be able to add the error-detection layer. Pricing Pro at $249 is below the GC AI critique seat benchmark at $500/mo, and Enterprise at $799 is a peer price to the authoring tools that feed it while landing 3.1x under an enterprise authoring contract on an annualized basis. White-Label at $1,499 makes resellers whole: 50 included reviews re-billed at $200 to $500 each is $10,000 to $25,000 of tenant revenue against a $1,499 cost.
08Competitive Landscape: Nobody Sells the Errors
Every AI-native player in this space is an authoring tool. They generate drafts. The nearest substitute for what we do is not a competitor product at all. It is a single frontier model and a prompt, and validation showed exactly what that substitute misses.
| Product | Category | Published price | Relationship to VerdictTank |
|---|---|---|---|
| AutogenAI | Enterprise authoring | Custom, sales-led, no self-serve | Complementary. We find the errors in what it writes. |
| Civio | Gov RFP authoring | Custom, sales-led | Complementary. Downstream reviewer. |
| Bidara | Mid-market authoring | $499/mo Starter | Complementary. Transparent pricing, natural comparison anchor. |
| AutoRFP.ai | Response automation | $899/mo Scale | Complementary. Reviews its drafts. |
| DeepRFP | Lean-team authoring | $89/user/mo | Complementary. Lowest per-seat price in the category, natural partner. |
| A solo frontier model | DIY substitute | API cost only | The real competitor. Measured: misses or underweights the material errors a panel catches. |
| VerdictTank | Panel error detection | Free (1/mo) · $249 Pro (5/mo) · $799 Enterprise (30/mo) · $1,499 White-Label (50/mo) | The only 6-seat, 5-vendor review panel with a published integrity gate and declared review quantities |
Competitor prices are vendors' own published rates as of July 2026. Tools without public pricing are shown as sales-led. Every named product was verified to exist and to occupy the authoring category.
Why the DIY substitute is the row that matters
Any buyer sophisticated enough to want proposal review can paste their document into a frontier model and ask for a critique. That is the honest competitive threat, and it is the one we tested against rather than around. The result is in section 3: the solo model returns a defensible, directionally correct score, and it returned none of the six VentureBuilt conditions, none of the CartMySupply compliance exposure, and none of the RFP Tank competitive reality.
1. Incumbents cannot sell honest criticism
Authoring tools sell the promise that they write your proposal. A brutal error list on the output that same tool just produced is a direct admission the generated draft is losing. It is structurally against their interest. We have no draft to defend. The verdict is the product.
2. Review is where the money is decided
Every serious bid already goes through a review gate, the color-team pass organizations run manually by pulling senior staff off billable work. That labor is expensive, slow, inconsistent between reviewers, and unavailable to the solo consultant. The demand is proven by the existence of the manual process.
3. Panel orchestration is a real moat
Five vendors, health gating, pre-baked failover, no model double-ups, and an integrity gate that challenges its own outliers is not a prompt. It is an operations problem, and the 2026-08-12 run is the evidence of what it costs to learn.
4. An empty category sets its own price
A crowded category means the budget line exists and you fight for share. An empty review category means we define the line and set the reference price, while remaining complementary to every authoring tool in the table above.
VerdictTank tells you the fourteen things both of them got wrong.
09Deployment Options
Two supported deployment shapes. Both are managed by IT Pro Partner below the application layer.
Option A · ITPP-INFRA Shared
Runs on existing netcup RS 4000 infrastructure alongside IT Pro Partner operations. Same Wasabi S3 backup pipeline, same Caddy reverse proxy, same monitoring stack (Prometheus and Grafana). Zero new infrastructure cost. Suitable for launch through Series A.
- netcup RS 4000 (app3), Docker Compose
- Wasabi S3 daily backups plus 15-minute sync
- Managed by the IT Pro Partner infrastructure team
Option B · Dedicated
Dedicated netcup or Hetzner instances with a dedicated S3 bucket. Full isolation from ITPP operational infrastructure. Recommended for post-Series A or enterprise white-label deployments requiring independent compliance scope.
- Dedicated netcup RS or Hetzner CPX instances
- Dedicated Wasabi S3 bucket, separate backup schedule
- Managed by IT Pro Partner below the application layer
Shared responsibility: IT Pro Partner manages everything below the application layer (OS, container runtime, networking, backups, monitoring) under both options. The VerdictTank application and its model pipeline are the product team's responsibility.
10Legal, Privacy & Compliance
Trademark clearance, the Minimum Viable Legal framework, the controller and processor role map, the sub-processor training guard, corpus confidentiality, and incident response. Carried forward from v4.0 and updated for the five-vendor panel.
10.1 USPTO trademark clearance: "VerdictTank"
Status: preliminary clearance only. This is not a substitute for a formal search. This assessment was performed with open-web search tools only. USPTO TESS is a JavaScript-rendered application and a static fetch returns only the search shell with no query results. Before any trademark application is filed, a live interactive TESS search or a paid clearance search through a trademark attorney is required.
Open-web common-law search results (performed)
| Search | Result | Assessment |
|---|---|---|
| "VerdictTank" exact, web-wide | Only hit is verdicttank.com itself | No third-party commercial use found |
| "Verdict Tank" space variant | Two incidental unrelated hits, neither a business nor a registered mark | No competing commercial use. Matches are noise. |
| Trademarkia and Justia proxy queries | No results returned | Consistent with no existing registration, but not equivalent to direct TESS |
| Domain: verdicttank.com | Live, owned, serving the product | Primary domain. Confirms operational use in commerce. |
| Domain: rfptank.com | Legacy holding, same naming convention | Defensive only. Retained against a family-of-marks argument. Not a product surface. |
Recommendation
- Before Series A close or any public marketing scale-up, commission a formal USPTO clearance search for Classes 9, 42, 35 and 45.
- File an intent-to-use application for VERDICTTANK as a standard character word mark, Class 42 primary and Class 9 secondary.
- Do not file on the basis of this document alone. It is a preliminary desk review.
10.2 Minimum Viable Legal (MVL) framework
MVL is the internal gate name used in the architecture documents as the precondition for onboarding white-label and enterprise customers. It is not one document. It is five interlocking instruments that must all exist and be internally consistent before the white-label provisioning gate turns green.
| Component | Purpose | Applies to | Status |
|---|---|---|---|
| Terms of Service | Governs the contractual relationship with every direct user | All tiers | Drafting required |
| Privacy Policy | GDPR and CCPA compliant notice of collection and use | All tiers | Drafting required |
| Data Processing Addendum | Article 28 GDPR processor terms | Enterprise, white-label | Hard gate on white-label |
| AI Disclaimer (DISC-001) | Non-removable versioned notice: output is AI opinion, not professional advice | Every scored surface | Engineering spec complete, legal copy needs counsel sign-off |
| Limitation of Liability | Caps aggregate liability at the lesser of $100 or fees paid in the preceding 12 months | All tiers, embedded in ToS | Drafting required |
| Governing law and venue | Recommend Delaware law with Georgia venue, pending confirmation of incorporation state | All tiers | Pending counsel |
| GDPR readiness | Lawful basis mapped per role. Articles 28, 33 and 34. SCCs or IDTA for EU transfers. | Any EU user | Framework mapped, SCC execution pending white-label launch |
| CCPA and CPRA readiness | Service-provider contract terms and a consumer rights workflow | Any California resident | DSAR workflow build pending |
The Free, Pro, Enterprise and White-Label tiers all require Terms of Service, Privacy Policy and the AI Disclaimer at minimum before any paid launch. White-Label additionally requires an executed DPA as a hard provisioning gate.
10.3 Controller and processor role map
| Data flow | Role | Legal basis | Agreements required |
|---|---|---|---|
| Free tier submission and review | Controller | Contract plus legitimate interest | ToS, Privacy Policy |
| Enterprise org admin and org users | Joint controller | Performance of contract | ToS, Enterprise DPA (Art. 26 GDPR) |
| White-label tenant end-users | Processor | Tenant's instructions | DPA, SCCs or IDTA, published sub-processor list |
| Panel model API calls, all five vendors | Controller of the vendor relationship. Each model vendor is a sub-processor. | Legitimate interest | Sub-processor training guard plus a DPA with each vendor |
| Corpus contribution (aggregate scores and structural metadata) | Controller, secondary-use basis | Opt-in consent. Cannot ride on contract or legitimate interest under the purpose limitation principle, Art. 5(1)(b). | Explicit opt-in UI, anonymization pipeline, retention separate from the review record |
10.4 Sub-processor training guard
Model clause for vendor DPAs
Enforcement rule: a vendor without a public, contractually confirmable training opt-out is excluded from the panel roster entirely and cannot be selected as a failover target. The only acceptable path for a non-compliant provider is a customer-side, explicit, revocable opt-in. Never a silent default, and never for corpus-eligible content. Each of the five rostered vendors is audited against this clause before it is eligible for a seat, and the audit is re-run at each roster revision.
10.5 Corpus confidentiality
The corpus is VerdictTank's most valuable long-term asset and its highest confidentiality exposure.
| Data type | Corpus-eligible | Rationale |
|---|---|---|
| Dimension scores and panel spread statistics | Yes, opt-in | Structural, not identifying. Core signal. |
| Structural metadata (vertical, length bucket, revision count, deltas) | Yes, opt-in | Enables content and moat analytics |
| Raw proposal text | Never | Confidential business content plus potential third-party PII |
| Explanation and audit finding text | Never in raw form | Critique text frequently quotes the submission verbatim |
| Chat refinement transcripts | Never as transcript content | Highest incidental-PII risk of any input surface |
Anonymization pipeline
- Source-content exclusion. Raw text fields excluded at the schema and ETL level.
- Structural extraction only. ETL reads scored and aggregated fields, never freeform text.
- Identifier stripping. Review, user and org identifiers replaced with a one-way surrogate key.
- Free-text quarantine. A stricter named-entity pass before any inclusion.
- k-anonymity floor. Public content published only when the cohort exceeds a minimum threshold.
Corpus contribution is off by default and requires explicit, separate opt-in. It is not bundled into ToS acceptance and is revocable at any time from account settings.
10.6 Incident response
Structured around the six functions of NIST CSF 2.0.
| Function | VerdictTank action |
|---|---|
| Govern | Named incident commander. Breach classification criteria documented before any incident. |
| Identify | Asset inventory: transactional database, corpus database, credentials for all five model vendors, white-label tenant segments. |
| Protect | Row-level-security multi-tenant isolation, sanitization gate, sub-processor training guard, per-vendor spend ceilings. |
| Detect | Alerting on anomalous data access, bulk export, and cross-org query attempts. Health-gate telemetry on every panel dispatch. |
| Respond | GDPR: 72-hour notification to the supervisory authority (Art. 33). CCPA: notification without unreasonable delay. |
| Recover | Post-incident review documented. White-label tenants notified per their individual DPA terms. |
Wrong-verdict liability
Scenario: a customer submits a proposal, receives a favorable panel verdict, acts on it, and the verdict was wrong in a way that led to a bad decision. This is primarily a reputational risk. The liability cap bounds legal exposure and does nothing for reputation.
- Legal layer: the AI disclaimer fails closed on every surface, a $100 or 12-months-of-fees liability cap applies, and the terms explicitly instruct users not to rely on AI output for investment decisions.
- Confidence calibration: every verdict ships with panel spread, standard deviation, and outlier flags. A verdict with a 5.4-point spread is a materially different signal from a unanimous one, and the report says so on its face.
- The published FAIL: section 3 of this document is itself part of the defense. We publish the case where our own thesis failed, which is a stronger honesty posture than any disclaimer.
- Incident playbook: do not litigate merits publicly, point to the auditable disclaimer version shown to the user, offer a private re-review, and disclose plus correct any systematic flaw found.
Model provider outage disclosure
- Public status page distinguishing VerdictTank infrastructure incidents from upstream model provider incidents.
- Degraded-mode behavior: a panel that ran short of its full six seats is flagged visibly with the seat count and which roles failed over. We never silently substitute a provider without disclosure. The 2026-08-12 validation run is reported throughout this document at the 8 reporting seats it actually produced on the v2.2 roster, for exactly that reason.
- SLA language: uptime commitments are qualified as dependent on upstream provider availability.