14 blind errors your solo model missed
++ A solo frontier model gives you a smooth, confident score. An 11-judge panel gives you the + 14 things it was wrong about. VerdictTank does not sell you a higher number. It sells you + the errors that number was hiding. +
+ +01The Thesis Changed, Because the Data Said So
++ VerdictTank v4.0 was sold on a claim we could not defend: that a multimodel panel produces a + better score than a single strong model. On 2026-08-12 we ran that claim against real proposals + and it failed. What we found instead is a stronger product. +
+ +What failed
++ The score-elevation thesis. Across three real proposals the panel mean came in + 0.40 points below the solo baseline. The panel did not lift scores. On two + of three proposals it pushed them down. We are publishing that result rather than burying it, + because the reason it happened is the product. +
+What worked
++ Error detection density. The same panel run surfaced 14 material errors that + the solo baseline missed or underweighted: revenue arithmetic that was wrong by a factor of + seven, a funded direct competitor the solo pass never named, a launch-blocking compliance + cost larger than projected first-year revenue. None of those show up as a score. All of them + decide whether the proposal wins. +
++ You buy it to find the $60K compliance hole and the broken contractor budget before you ship. +
Why the spread is the signal
++ A single model scoring alone produces low variance. It reads the document once, forms one + coherent opinion, and every dimension it emits is downstream of that opinion. The result feels + authoritative precisely because nothing inside it disagrees. +
++ An 11-seat panel of nine different vendors cannot produce that coherence, and the incoherence + is diagnostic. When a Financial Integrity judge scores a proposal 8.2 while an + Execution Feasibility judge scores the same document 2.8, that 5.4-point spread is not noise. + It is a precise statement: the money works, the delivery plan does not. A solo model + averages that tension away into a single confident 6.1 and tells you nothing actionable. +
+02How We Found This: Four Generations and a Self-Review
++ VerdictTank is a proposal review engine, not a proposal writer. It ingests a finished document + and returns scored dimensions plus a ranked list of concrete Fix-It items. The pipeline was not + designed in the abstract. It was hardened across four architectural generations, and then it + was pointed at its own proposals. +
+ +Single-Model Scorer
++ v2 established the core insight: a proposal has two independent quality axes. Narrative + quality (clarity, structure, persuasion) and compliance quality (does it actually answer + the scored requirements). One model scored both from one prompt. It proved the concept and + exposed the flaw: the axes bled together. A beautifully written section that missed a + mandatory requirement scored too high, because the same reasoning pass that admired the + prose also graded the compliance. +
+Separated Scoring Passes
++ v3 split scoring into two independent passes with two purpose-built prompts. The narrative + pass never sees the compliance rubric. The compliance pass never rewards eloquence. This is + the decision that makes the dual score trustworthy: the two numbers can now disagree, and + their disagreement carries information. A 9/10 narrative next to a 4/10 compliance is a + proposal about to lose. +
+Multimodel Adversarial Review
++ A single model scoring in isolation is confidently wrong at a predictable rate. v4 + introduced a multimodel pipeline: a fast model produces first-pass scores and Fix-It + candidates, then a stronger model reviews that output adversarially, challenging every + deduction and confirming each Fix-It maps to real proposal text. Scores stopped drifting + between runs. This is the generation that made the output defensible. +
+Specialist Panel and the Integrity Gate
++ v5 replaces the adversarial pair with an 11-seat specialist panel across nine vendors, and + adds a synthesis seat whose only job is to compute panel statistics, flag scores more than + 1.5 standard deviations from the mean, and reconcile the verdict against the evidence. + The output is no longer a number. It is a number, a spread, an outlier list, and a ranked + set of material errors with the judge that caught each one. +
+The self-review that broke the old thesis
++ Before selling a review engine we ran the engine on our own work. We assembled the panel and + scored three real proposals, in full, with the same prompts and rubric a paying customer would + get. One of the three was VerdictTank's own sibling product. The panel returned a NO GO on it. +
++ That run cost roughly $150 in inference and returned 8 of 11 seats. Two seats were lost to a + provider credit wall hit mid-run and one to a model family that could not be dispatched at all. + The incomplete panel is why v2.3 of the judge pool spec now requires a pre-flight health gate + and a pre-baked failover roster, covered in section 5. The results below are what those 8 seats + produced, and we report them at 8 seats rather than extrapolating to 11. +
+ +03Validation Run: 3 Real Proposals, 8 Reporting Seats
++ Every figure in this section comes from the 2026-08-12 validation run. Nothing is modeled, + projected, or illustrative. The solo baseline is Claude Opus 5 scoring the same documents + against the same 10-dimension rubric. +
+ +Panel mean vs solo baseline
+| Proposal | Panel mean | Solo baseline | Delta | Panel spread | Solo spread | Verdict |
|---|---|---|---|---|---|---|
| RFP Tank v1.0 | 4.40 | 4.93 | +-0.53 | 3.7 | 1.4 | +NO GO | +
| VentureBuilt v2 | 6.14 | 6.10 | ++0.04 | 5.4 | 2.4 | +CONDITIONAL GO | +
| CartMySupply | 4.29 | 5.00 | +-0.71 | 2.6 | 1.8 | +NO GO | +
| Aggregate | 4.94 | 5.34 | +-0.40 | Panel spread exceeded solo spread on all 3 | +THESIS FAIL | +|
+ Thesis under test: the panel must show a greater than 0.5 point advantage over the solo mean to + justify premium pricing. Result: FAIL on all three proposals and FAIL in aggregate. Panel + composition for this run was 8 reporting judges (4 Band A, 4 Band B) out of 11 specified seats, + a 73% coverage rate. +
+ +Why the panel scored lower
++ The panel does not elevate scores. It sharpens error detection, and error detection on a flawed + document moves the number down. All three proposals contained severe cross-cutting defects that + additional specialist scrutiny exposed more precisely: fatal execution gaps, competitive + mispositioning, and legal blockers. The solo baseline was directionally correct on all three. + The panel added precision, not points. +
++ That is the entire finding, and it inverts the sales pitch. If your proposal is sound, the panel + will roughly agree with a good solo model and cost you more. If your proposal has a hole in it, + the panel finds the hole and the solo model does not. You are not buying a score. You are buying + the probability that a specific, expensive, named mistake gets caught before an evaluator or an + investor finds it for you. +
+ +Specialist divergence, measured
+| Observation | Evidence from the run | What it means |
|---|---|---|
| Generalist seats run optimistic | +Gemini Pro scored RFP Tank 6.7 as a Band A generalist and 3.0 as the Band B Market specialist. Same model, same document, 3.7 points apart. | +Band A generalist scoring without specialist cross-check is systematically over-optimistic. The role, not the model, drives the score. | +
| Role divergence beats model divergence | +DeepSeek V4 Pro scored VentureBuilt 6.4 as Cross-Check C and 2.8 as Execution Feasibility. A 3.6 point split inside one vendor. | +Panel diversity is not primarily about buying different vendors. It is about buying different questions. | +
| One seat can flip a verdict | +Remove the 2.8 Execution score from VentureBuilt and the panel averages 6.6, reading as a clean GO. With it, the verdict is CONDITIONAL GO with a named contractor-budget fix. | +The lowest score in the panel is frequently the only one doing work. Averaging is what a solo model already does. | +
| Tight clustering is also a signal | +CartMySupply produced zero outliers beyond 1.5 sigma and the tightest spread of the three (sigma 0.89). | +Unanimity across nine vendors on a low score is a far stronger NO GO than one model's low score. | +
04The 14 Errors: Every One Named
++ This is the product. Fourteen material errors the 8-judge panel caught that the solo baseline + missed or underweighted, grouped by failure class. Each is a real finding from the 2026-08-12 + run against a real document. +
+ +RFP Tank v1.0 · panel 4.40 vs solo 4.93
+| Class | Error the panel caught | Caught by |
|---|---|---|
| Revenue | Three mutually inconsistent Year-1 revenue figures inside one document: $1.2M, $1.361M, and $372K. Plus a 22% MRR ramp inconsistency the narrative never reconciles. | Financial Integrity |
| Competitive | CLEATUS is a real, funded competitor at $4M seed with public product-led pricing of $39 to $250/mo, occupying the identical quadrant. The proposal does not name it. The Band A generalist seat actually cited CLEATUS pricing as a positive signal. | Market Reality |
| Competitive | GovEagle pricing referenced at a 15x inconsistency against the proposal's own comparison table. | Market Reality |
| Execution | Five of seven features marked TO BUILD at HIGH effort. The real-time Compliance Copilot alone needs 2 to 3 developers for 8 to 12 weeks. The plan allocates 4 weeks, solo. | Execution Feasibility |
| Team | A solo founder shipping a 7-feature AI SaaS in 10 weeks, with hiring contingent on revenue that requires the product to already exist. A closed loop with no entry point. | Team / Founder |
| Legal | No privacy policy and no terms of service, against FAR and CUI exposure, with ITAR implications on German-hosted infrastructure. | Legal / Regulatory |
+ Panel verdict: NO GO. Estimated rework 40+ hours. Recommendation is to cut scope to two features, + extend to 20 weeks, hire a second developer before month one, rebuild the financial model, and + address CLEATUS directly. +
+ +VentureBuilt v2 · panel 6.14 vs solo 6.10
+| Class | Error the panel caught | Caught by |
|---|---|---|
| Execution | Contractor budget broken by a factor of 4 to 7. The stated $1,500/mo implies $11 to $22 per hour against a market rate of $75 to $100. At real rates that budget buys 105 to 140 hours and leaves roughly 800 hours on the founder. | Execution Feasibility |
| Revenue | Year 2 stated on a run-rate basis rather than recognized revenue. Restated correctly, the healthy scenario loses roughly $11K to $18K. | Financial Integrity |
| Competitive | The uniqueness claim is contradicted by shipping products. LivePlan Plan Review and IdeaProof already occupy the space. | Market Reality |
| Team | 37 engagements plus 950 hours plus an MSP day job. The three commitments cannot coexist in one calendar. | Team / Founder |
+ Panel verdict: CONDITIONAL GO with six named conditions. Estimated rework 15 to 20 hours. This is + the case that most clearly shows the value: the panel mean (6.14) and the solo mean (6.10) are + statistically tied, so on score alone the panel added nothing. What it added was a bimodal split, + Financial 8.2 against Execution 2.8, and the four errors above. +
+ +CartMySupply · panel 4.29 vs solo 5.00
+| Class | Error the panel caught | Caught by |
|---|---|---|
| Revenue | The $2.7M headline is wrong by 7x to 10x against the proposal's own inputs, which compute to $269K. Stripe fees understated by roughly $11K per year. CAC absent entirely. | Financial Integrity |
| Competitive | TeacherLists already solves the identical problem, free, across 2 million lists. Target ships native School List Assist. No technical moat is claimed or demonstrable. | Market Reality |
| Execution | Amazon PA-API 5 removed Cart API support. Target has no self-serve multi-item cart API. Walmart requires separate catalog matching. The core mechanic of the product does not have a supported integration path at any of the three named retailers. | Execution Feasibility |
| Legal | Charitable solicitation registration required in 40+ states at $30K to $75K, plus COPPA exposure and FTC penalty risk. Compliance cost of $60K to $150K exceeds projected Year-1 revenue of $3K to $14K by an order of magnitude. | Legal / Regulatory |
+ Panel verdict: NO GO, unanimous, zero outliers, tightest spread of the three. Estimated rework + 60+ hours. The build estimate of 116 hours was independently judged 4x to 10x too low. +
+ +05Judge Pool v2.3: 11 Seats, 9 Vendors, Zero Double-Ups
++ The panel that produced the validation data ran at 8 of 11 seats because two seats hit a + provider credit wall mid-run and one model family could not be dispatched at all. v2.3 is the + spec written in response to that failure. Full detail lives in the + judge pool specification v2.3. +
+ +The roster
+| Band | Seat | Model | Vendor | Scores |
|---|---|---|---|---|
| 0 | Research Agent | Grok 4.5 | xAI | No |
| A | Primary Reviewer | Claude Opus 5 | Anthropic | Yes |
| A | Cross-Check A | DeepSeek V4 Flash | DeepSeek | Yes |
| A | Cross-Check B | Gemini Pro Latest | Yes | |
| A | Cross-Check C | DeepSeek V4 Pro | DeepSeek | Yes |
| A | Legal / Regulatory | Claude Sonnet 5 | Anthropic | Yes |
| B | Financial Integrity | MiniMax-M3 | MiniMax | Yes |
| B | Team / Founder | Claude Fable 5 | Anthropic | Yes |
| B | Market Reality | Qwen3.7 Plus | Alibaba | Yes |
| B | Execution Feasibility | GPT-5.2 Pro | OpenAI | Yes |
| C | Synthesis & Integrity Gate | Kimi K2.6 | Moonshot | No |
+ Nine distinct vendors across eleven seats. Nine distinct scoring models. Zero model double-ups: + no single model occupies two scoring seats, which is the constraint that keeps correlated + failure out of the panel mean. Maximum vendor concentration is Anthropic at 3 of 11 (27.3%), + comfortably inside the 40% ceiling. DeepSeek holds 2 of 11 (18.2%). Every remaining vendor holds + exactly one seat. +
+ +What changed in v2.3
+Pre-flight health gate
++ Before any scoring begins, the orchestrator pings every rostered model with a 5-second + probe and writes the result to a per-model health file. Any model returning HTTP 400, + HTTP 429, or a no-healthy-deployments error is swapped for its pre-assigned failover before + a single scoring call is spent. The 2026-08-12 run burned roughly 12 dispatches discovering + dead models at runtime. That failure mode is now closed. +
+Pre-baked failover roster
++ Every seat carries a named failover from a different vendor, resolved at gate time rather + than improvised mid-run. Failover selection preserves both the vendor-diversity ceiling and + the no-double-up rule, so a degraded panel is still a valid panel rather than an + accidentally correlated one. +
+Credit-wall resilience
++ The provider credit exhaustion that cost two seats mid-run is now detected at the gate and + treated as an availability failure, not an error. Anthropic capacity has been restored and + the Primary Reviewer seat runs Claude Opus 5 as specified. +
+Permanent exclusions
++ One frontier model family proved structurally incapable of running as a panel seat under + our orchestration and is permanently excluded from the roster, not merely deprioritized. + Excluded models cannot be selected as a failover target either. +
+Latency
++ v2.3 targets a critical path of roughly 113 seconds, against 218 seconds measured on the + v2.1 architecture. The improvement comes from band parallelism: Band A and Band B seats execute + concurrently rather than sequentially, and the Synthesis seat is the only stage that must wait + for all scoring seats to return. +
+ +The ten scored dimensions
++ Every scoring seat rates the proposal 1 to 10 on the same ten dimensions, so panel spread is + computed dimension by dimension and not only in aggregate: +
+The Synthesis and Integrity Gate
++ The Band C seat never scores. It reads all scoring output and performs a fixed checklist: + verify score arithmetic, compute panel means and per-dimension spread, compute the delta against + the solo baseline, flag every score more than 1.5 standard deviations from the panel mean with a + written rationale, and confirm the verdict follows from panel evidence rather than from the + Primary Reviewer alone. That gate is what turns eleven opinions into one auditable report. +
+06Worked Example: VentureBuilt v2, Where the Score Said Nothing
++ This is the clearest case in the validation set, because it is the one where score elevation + delivered exactly zero and error detection delivered everything. Real scores from the + 2026-08-12 run. +
+ +Median 6.45 · standard deviation 1.758 · spread 5.4 (min 2.8, max 8.2)
+Spread 2.4 · delta +0.04 · statistically tied with the panel
+Individual seat scores
+| Seat | Model | Score | Sigma from mean | Flag |
|---|---|---|---|---|
| Band B · Financial Integrity | MiniMax | 8.2 | +1.17 | Within 1.5 sigma |
| Band B · Market Reality | Gemini | 7.7 | +0.89 | Within 1.5 sigma |
| Band A · Primary Reviewer | Opus 5 | 6.9 | +0.43 | Within 1.5 sigma |
| Band A · Legal / Regulatory | Qwen | 6.5 | +0.21 | Within 1.5 sigma |
| Band A · Cross-Check C | DeepSeek V4 Pro | 6.4 | +0.15 | Within 1.5 sigma |
| Band A · Cross-Check B | Gemini | 6.2 | +0.04 | Within 1.5 sigma |
| Band B · Team / Founder | Kimi | 4.4 | -0.99 | Within 1.5 sigma |
| Band B · Execution Feasibility | DeepSeek V4 Pro | 2.8 | -1.91 | OUTLIER | +
What the outlier actually found
++ The 2.8 was not a grumpy model. The Integrity Gate challenged it at 1.91 sigma and it survived + the challenge on evidence: a contractor budget broken 4x to 7x, an architecture that regressed + from v1 with no schema and no API contract, and a founder workload of 37 engagements plus 950 + hours alongside an MSP day job. The Team seat (4.4) and the Band A generalists (6.2 to 6.9) + all acknowledged the same workload problem. They weighted it less severely. The Execution + specialist is the only seat that forced it into the verdict. +
+Fix-It items, ranked by materiality
+-
+
-
+ CriticalExecution+ Fix the contractor budget or cut the scope. $1,500/mo buys 105 to 140 hours + at market rates, not the volume the plan assumes. Raise to roughly $7,500/mo or reduce scope + to fit the hours actually purchased. +
+
-
+ CriticalFinancial+ Restate Year 2 on a recognized-revenue basis. On run-rate the year looks + healthy. On recognized revenue it loses roughly $11K to $18K. Present both. +
+
-
+ HighCompetitive+ Withdraw or qualify the uniqueness claim. LivePlan Plan Review and IdeaProof + already ship in this space. Reposition on a defensible axis. +
+
-
+ HighTeam+ Name the contractor and the sourcing plan before Phase 2. A budget line with + no named person is not a capacity plan. +
+
-
+ MediumGo-to-market+ Map the Year 1 to Year 2 GTM bridge. Eight net-new signups per month appear + in the model with no acquisition mechanism behind them. +
+
-
+ MediumLegal+ Complete data protection and trademark clearance before Phase 0 to 1. +
+
+ Panel verdict: CONDITIONAL GO. Estimated rework 15 to 20 hours. The solo baseline returned a + 6.10 and none of the six conditions above. +
+07Pricing: Priced Per Error Found, Not Per Point Gained
++ Three tiers. The pricing logic follows the revised thesis directly: a panel run is worth what a + caught error is worth, and a caught error is worth far more than a point of score. +
+ +Free
+-
+
- 1 full review +
- Top 3 Fix-It items +
- Panel score and spread +
- Reduced panel size +
Pro
+-
+
- 5 reviews per month +
- Full 11-seat panel +
- Complete Fix-It list, ranked +
- Panel spread and outlier flags +
- Solo-baseline delta comparison +
- Re-score loop with before and after +
Enterprise
+-
+
- Unlimited reviews +
- White-label branding +
- Multi-seat team workspaces +
- Configurable judge pool +
- Corpus isolation and data controls +
- Priority pipeline and support +
What a review costs us, and why the panel is affordable
++ The validation run cost approximately $150 in inference for three full proposals across eight + reporting seats, including retries against dead models before the health gate existed. That + burn is the honest anchor for panel economics: a clean 11-seat run on one proposal, with the + health gate preventing wasted dispatches, sits well inside single-digit dollars. +
++ The reason a full 11-seat panel fits a $79 tier at five reviews per month is vendor mix. Only + a minority of seats run premium frontier models. The specialist Band B seats run strong + mid-tier models from five different vendors, which is where the error-detection value came from + in validation. Panel diversity is cheaper than panel depth, and diversity is what caught the 14. +
+ +Why the value question is not the score question
+| Error class | Real example from validation | Cost of missing it |
|---|---|---|
| Legal blocker | Charitable solicitation registration in 40+ states | $30K to $75K of registration, against $3K to $14K of projected revenue |
| Compliance total | Full first-year compliance load on the same proposal | $60K to $150K, exceeding Year-1 revenue by roughly 10x |
| Execution gap | Contractor budget short by 4x to 7x | Roughly 800 unbudgeted founder hours |
| Revenue arithmetic | $2.7M headline against $269K computed from the document's own inputs | Credibility with any investor who checks the math, which is all of them |
| Competitive blind spot | A $4M-seed funded direct rival never named in the document | The first question in the room, unanswered |
+ A single caught item in the top two rows pays for a decade of the Pro tier. That is the entire + pricing argument, and it does not depend on the panel producing a higher score, which it does + not. +
+ +Positioned against the authoring category
+| Comparison | Their price | VerdictTank | Multiple |
|---|---|---|---|
| Pro vs Bidara Starter | $499/mo | $79/mo | 6.3x cheaper |
| Pro vs AutoRFP.ai Scale | $899/mo | $79/mo | 11.4x cheaper |
| Enterprise vs Bidara Starter | $499/mo | $299/mo | 1.7x cheaper |
| Enterprise vs AutoRFP.ai Scale | $899/mo | $299/mo | 3.0x cheaper |
+ We are not a proposal team in a box. We are one high-value pass in the workflow. A buyer already + spending $499 to $899 per month on an authoring tool should be able to add the error-detection + layer without a second budget conversation. Pricing Pro at $79 makes VerdictTank an add-on + decision rather than a platform decision. +
+08Competitive Landscape: Nobody Sells the Errors
++ Every AI-native player in this space is an authoring tool. They generate drafts. The nearest + substitute for what we do is not a competitor product at all. It is a single frontier model + and a prompt, and validation showed exactly what that substitute misses. +
+ +| Product | Category | Published price | Relationship to VerdictTank |
|---|---|---|---|
| AutogenAI | Enterprise authoring | Custom, sales-led, no self-serve | Complementary. We find the errors in what it writes. |
| Civio | Gov RFP authoring | Custom, sales-led | Complementary. Downstream reviewer. |
| Bidara | Mid-market authoring | $499/mo Starter | Complementary. Transparent pricing, natural comparison anchor. |
| AutoRFP.ai | Response automation | $899/mo Scale | Complementary. Reviews its drafts. |
| DeepRFP | Lean-team authoring | $89/user/mo | Complementary. Lowest per-seat price in the category, natural partner. |
| A solo frontier model | DIY substitute | API cost only | The real competitor. Measured: misses or underweights the material errors a panel catches. |
| VerdictTank | Panel error detection | Free / $79 Pro / $299 Enterprise | The only 11-seat, 9-vendor review panel with a published integrity gate |
+ Competitor prices are vendors' own published rates as of July 2026. Tools without public pricing + are shown as sales-led. Every named product was verified to exist and to occupy the authoring + category. +
+ +Why the DIY substitute is the row that matters
++ Any buyer sophisticated enough to want proposal review can paste their document into a frontier + model and ask for a critique. That is the honest competitive threat, and it is the one we tested + against rather than around. The result is in section 3: the solo model returns a defensible, + directionally correct score, and it returned none of the six VentureBuilt conditions, none of + the CartMySupply compliance exposure, and none of the RFP Tank competitive reality. +
+ +1. Incumbents cannot sell honest criticism
++ Authoring tools sell the promise that they write your proposal. A brutal error list on the + output that same tool just produced is a direct admission the generated draft is losing. It + is structurally against their interest. We have no draft to defend. The verdict is the + product. +
+2. Review is where the money is decided
++ Every serious bid already goes through a review gate, the color-team pass organizations run + manually by pulling senior staff off billable work. That labor is expensive, slow, + inconsistent between reviewers, and unavailable to the solo consultant. The demand is proven + by the existence of the manual process. +
+3. Panel orchestration is a real moat
++ Nine vendors, health gating, pre-baked failover, no model double-ups, and an integrity gate + that challenges its own outliers is not a prompt. It is an operations problem, and the + 2026-08-12 run is the evidence of what it costs to learn. +
+4. An empty category sets its own price
++ A crowded category means the budget line exists and you fight for share. An empty review + category means we define the line and set the reference price, while remaining complementary + to every authoring tool in the table above. +
++ VerdictTank tells you the fourteen things both of them got wrong. +
09Deployment Options
++ Two supported deployment shapes. Both are managed by IT Pro Partner below the application layer. +
+ +Option A · ITPP-INFRA Shared
++ Runs on existing netcup RS 4000 infrastructure alongside IT Pro Partner operations. Same + Wasabi S3 backup pipeline, same Caddy reverse proxy, same monitoring stack (Prometheus and + Grafana). Zero new infrastructure cost. Suitable for launch through Series A. +
+-
+
- netcup RS 4000 (app3), Docker Compose +
- Wasabi S3 daily backups plus 15-minute sync +
- Managed by the IT Pro Partner infrastructure team +
Option B · Dedicated
++ Dedicated netcup or Hetzner instances with a dedicated S3 bucket. Full isolation from ITPP + operational infrastructure. Recommended for post-Series A or enterprise white-label + deployments requiring independent compliance scope. +
+-
+
- Dedicated netcup RS or Hetzner CPX instances +
- Dedicated Wasabi S3 bucket, separate backup schedule +
- Managed by IT Pro Partner below the application layer +
+ Shared responsibility: IT Pro Partner manages everything below the application + layer (OS, container runtime, networking, backups, monitoring) under both options. The + VerdictTank application and its model pipeline are the product team's responsibility. +
+ +10Legal, Privacy & Compliance
++ Trademark clearance, the Minimum Viable Legal framework, the controller and processor role map, + the sub-processor training guard, corpus confidentiality, and incident response. Carried forward + from v4.0 and updated for the nine-vendor panel. +
+ +10.1 USPTO trademark clearance: "VerdictTank"
++ Status: preliminary clearance only. This is not a substitute for a formal search. + This assessment was performed with open-web search tools only. USPTO TESS is a + JavaScript-rendered application and a static fetch returns only the search shell with no query + results. Before any trademark application is filed, a live interactive TESS search or a paid + clearance search through a trademark attorney is required. +
+ +Open-web common-law search results (performed)
+| Search | Result | Assessment |
|---|---|---|
| "VerdictTank" exact, web-wide | Only hit is verdicttank.com itself | No third-party commercial use found |
| "Verdict Tank" space variant | Two incidental unrelated hits, neither a business nor a registered mark | No competing commercial use. Matches are noise. |
| Trademarkia and Justia proxy queries | No results returned | Consistent with no existing registration, but not equivalent to direct TESS |
| Domain: verdicttank.com | Live, owned, serving the product | Primary domain. Confirms operational use in commerce. |
| Domain: rfptank.com | Legacy holding, same naming convention | Defensive only. Retained against a family-of-marks argument. Not a product surface. |
Recommendation
+-
+
- Before Series A close or any public marketing scale-up, commission a formal USPTO clearance search for Classes 9, 42, 35 and 45. +
- File an intent-to-use application for VERDICTTANK as a standard character word mark, Class 42 primary and Class 9 secondary. +
- Do not file on the basis of this document alone. It is a preliminary desk review. +
10.2 Minimum Viable Legal (MVL) framework
++ MVL is the internal gate name used in the architecture documents as the precondition for + onboarding white-label and enterprise customers. It is not one document. It is five interlocking + instruments that must all exist and be internally consistent before the white-label provisioning + gate turns green. +
+| Component | Purpose | Applies to | Status |
|---|---|---|---|
| Terms of Service | Governs the contractual relationship with every direct user | All tiers | Drafting required |
| Privacy Policy | GDPR and CCPA compliant notice of collection and use | All tiers | Drafting required |
| Data Processing Addendum | Article 28 GDPR processor terms | Enterprise, white-label | Hard gate on white-label |
| AI Disclaimer (DISC-001) | Non-removable versioned notice: output is AI opinion, not professional advice | Every scored surface | Engineering spec complete, legal copy needs counsel sign-off |
| Limitation of Liability | Caps aggregate liability at the lesser of $100 or fees paid in the preceding 12 months | All tiers, embedded in ToS | Drafting required |
| Governing law and venue | Recommend Delaware law with Georgia venue, pending confirmation of incorporation state | All tiers | Pending counsel |
| GDPR readiness | Lawful basis mapped per role. Articles 28, 33 and 34. SCCs or IDTA for EU transfers. | Any EU user | Framework mapped, SCC execution pending white-label launch |
| CCPA and CPRA readiness | Service-provider contract terms and a consumer rights workflow | Any California resident | DSAR workflow build pending |
+ The Free, Pro and Enterprise tiers require Terms of Service, Privacy Policy and the AI Disclaimer + at minimum before any paid launch. +
+ +10.3 Controller and processor role map
+| Data flow | Role | Legal basis | Agreements required |
|---|---|---|---|
| Free tier submission and review | Controller | Contract plus legitimate interest | ToS, Privacy Policy |
| Enterprise org admin and org users | Joint controller | Performance of contract | ToS, Enterprise DPA (Art. 26 GDPR) |
| White-label tenant end-users | Processor | Tenant's instructions | DPA, SCCs or IDTA, published sub-processor list |
| Panel model API calls, all nine vendors | Controller of the vendor relationship. Each model vendor is a sub-processor. | Legitimate interest | Sub-processor training guard plus a DPA with each vendor |
| Corpus contribution (aggregate scores and structural metadata) | Controller, secondary-use basis | Opt-in consent. Cannot ride on contract or legitimate interest under the purpose limitation principle, Art. 5(1)(b). | Explicit opt-in UI, anonymization pipeline, retention separate from the review record |
10.4 Sub-processor training guard
+Model clause for vendor DPAs
++ Enforcement rule: a vendor without a public, contractually confirmable training + opt-out is excluded from the panel roster entirely and cannot be selected as a failover target. + The only acceptable path for a non-compliant provider is a customer-side, explicit, revocable + opt-in. Never a silent default, and never for corpus-eligible content. Each of the nine rostered + vendors is audited against this clause before it is eligible for a seat, and the audit is + re-run at each roster revision. +
+ +10.5 Corpus confidentiality
++ The corpus is VerdictTank's most valuable long-term asset and its highest confidentiality + exposure. +
+| Data type | Corpus-eligible | Rationale |
|---|---|---|
| Dimension scores and panel spread statistics | Yes, opt-in | Structural, not identifying. Core signal. |
| Structural metadata (vertical, length bucket, revision count, deltas) | Yes, opt-in | Enables content and moat analytics |
| Raw proposal text | Never | Confidential business content plus potential third-party PII |
| Explanation and audit finding text | Never in raw form | Critique text frequently quotes the submission verbatim |
| Chat refinement transcripts | Never as transcript content | Highest incidental-PII risk of any input surface |
Anonymization pipeline
+-
+
- Source-content exclusion. Raw text fields excluded at the schema and ETL level. +
- Structural extraction only. ETL reads scored and aggregated fields, never freeform text. +
- Identifier stripping. Review, user and org identifiers replaced with a one-way surrogate key. +
- Free-text quarantine. A stricter named-entity pass before any inclusion. +
- k-anonymity floor. Public content published only when the cohort exceeds a minimum threshold. +
+ Corpus contribution is off by default and requires explicit, separate opt-in. It is not bundled + into ToS acceptance and is revocable at any time from account settings. +
+ +10.6 Incident response
+Structured around the six functions of NIST CSF 2.0.
+| Function | VerdictTank action |
|---|---|
| Govern | Named incident commander. Breach classification criteria documented before any incident. |
| Identify | Asset inventory: transactional database, corpus database, credentials for all nine model vendors, white-label tenant segments. |
| Protect | Row-level-security multi-tenant isolation, sanitization gate, sub-processor training guard, per-vendor spend ceilings. |
| Detect | Alerting on anomalous data access, bulk export, and cross-org query attempts. Health-gate telemetry on every panel dispatch. |
| Respond | GDPR: 72-hour notification to the supervisory authority (Art. 33). CCPA: notification without unreasonable delay. |
| Recover | Post-incident review documented. White-label tenants notified per their individual DPA terms. |
Wrong-verdict liability
++ Scenario: a customer submits a proposal, receives a favorable panel verdict, + acts on it, and the verdict was wrong in a way that led to a bad decision. This is primarily a + reputational risk. The liability cap bounds legal exposure and does nothing for reputation. +
+-
+
- Legal layer: the AI disclaimer fails closed on every surface, a $100 or 12-months-of-fees liability cap applies, and the terms explicitly instruct users not to rely on AI output for investment decisions. +
- Confidence calibration: every verdict ships with panel spread, standard deviation, and outlier flags. A verdict with a 5.4-point spread is a materially different signal from a unanimous one, and the report says so on its face. +
- The published FAIL: section 3 of this document is itself part of the defense. We publish the case where our own thesis failed, which is a stronger honesty posture than any disclaimer. +
- Incident playbook: do not litigate merits publicly, point to the auditable disclaimer version shown to the user, offer a private re-review, and disclose plus correct any systematic flaw found. +
Model provider outage disclosure
+-
+
- Public status page distinguishing VerdictTank infrastructure incidents from upstream model provider incidents. +
- Degraded-mode behavior: a panel that ran short of its full eleven seats is flagged visibly with the seat count and which roles failed over. We never silently substitute a provider without disclosure. The 2026-08-12 run is reported at 8 of 11 seats throughout this document for exactly that reason. +
- SLA language: uptime commitments are qualified as dependent on upstream provider availability. +