Sync docs, audit artifacts, project notes, and VerdictTank proposal docs

- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
This commit is contained in:
root
2026-08-26 02:27:28 -04:00
parent 23e9751d38
commit f5175f1ce0
55 changed files with 14669 additions and 3 deletions
@@ -0,0 +1,411 @@
# VerdictTank Judge Pool Spec v2.0 - Primary Reviewer Scorecard
**Reviewer:** Claude Opus 5 (Primary Reviewer, Phase 2)
**Date:** 2026-08-12
**Spec:** https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md (17,641 bytes, v2.0)
**Method:** Document critique + live empirical validation against admin-ai (157-model catalog, real inference runs)
> **Note on version:** `web_extract` returned a cached **v1.0**. Verified against the live origin
> via `curl` - the deployed file is **v2.0**, byte-identical to
> `/root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md`. This review is of v2.0.
---
## Scorecard
| # | Dimension | Score | Verdict |
|---|---|---|---|
| 1 | Role Coverage | 7/10 | FIX |
| 2 | Model Assignments | 6/10 | FIX |
| 3 | Vendor Diversity | 8/10 | DEFEND |
| 4 | admin-ai Paths | 9/10 | DEFEND |
| 5 | Fallback Chains | 4/10 | FIX |
| 6 | Content-Based Rules | 5/10 | FIX |
| 7 | Latency Budgets | 2/10 | FIX |
| 8 | Feedback Loop | 5/10 | FIX |
| 9 | Pricing | 4/10 | FIX |
| 10 | Product Boundary | 7/10 | DEFER |
**Weighted mean: 5.7/10** - Architecturally literate, empirically unvalidated. Two dimensions (7, 9) are
launch-blocking. One unlisted finding (§Critical Finding) threatens the product thesis itself.
---
## CRITICAL FINDING (unlisted dimension - read first)
**The 12-judge panel does not measurably outperform one model run once.**
I tested this directly. Same proposal, same rubric, three dimensions, live models.
**A) Single-model noise floor** - `claude-opus-5` x6 at temp 1.0:
```
market: 5,5,5,4,4,4 stdev 0.50
team: 4,4,4,4,4,4 stdev 0.00
fin: 3,3,3,3,3,3 stdev 0.00
```
**B) Nine distinct judges across 8 vendors, one run each:**
```
Opus5 [5,4,3] Sonnet5 [4,4,3] Sol [6,4,3] GeminiPro [4,3,2] DeepSeek [5,3,2]
Terra [4,4,3] Qwen [4,3,2] MiniMax [6,3,4] Fable5 [6,4,4]
panel stdev: market 0.87, team 0.50, fin 0.74
```
**C) The result that matters:**
| | market | team | financials |
|---|---|---|---|
| Mean, 1 model | 4.50 | 4.00 | 3.00 |
| Mean, 9 judges / 8 vendors | 4.89 | 3.56 | 2.89 |
| **Delta** | **0.39** | **0.44** | **0.11** |
Nine frontier models from eight vendors, ~218s of wall clock and roughly 20x the token cost, move the
final number by **less than half a point on a 10-point scale**. On `market`, panel spread (0.87) is only
**1.7x** the single-model rerun noise (0.50) - statistically indistinguishable from rerunning one model
six times.
The individual judges *do* disagree (market scores span 4-6). But disagreement that averages to the same
answer is **variance, not signal**. The spec sells vendor diversity as the core accuracy mechanism and
never once tests whether diversity changes the output.
**This is the product thesis, and it is unvalidated.** Every downstream claim - premium pricing, the
"most accurate review tool available" positioning, the 12-vendor moat - rests on it.
**Required before build:** run 30-50 real proposals with known outcomes. Report panel-vs-single
correlation against ground truth. If the panel does not beat one good model by a margin that justifies
20x cost, the correct architecture is 3 judges, not 12 - and the pricing story needs rebuilding.
Better to learn this now than after a customer runs the same A/B.
---
## 1. Role Coverage - 7/10 - FIX
v2.0 deserves credit: it closed three of the four gaps flagged in the prior review (Financial Integrity,
Team/Founder, Legal/Regulatory now have dedicated phases). That is real progress.
**Technical Architecture remains uncovered** - the one gap explicitly identified in the v4.1 gap analysis
and silently dropped. Execution-Feasibility is operational (timeline, team, resources), not architectural
(stack, scalability, security posture, technical debt). For a product whose buyers are evaluating
*technical* startups, having no judge that reads the architecture is a conspicuous hole.
Two further gaps neither version names:
- **Traction/Evidence.** Nobody scores whether claims are *substantiated*. Research Agent gathers
citations but never scores; no downstream role is tasked with "the founder asserts 40% MoM growth -
is there evidence?" This is the single most common reason real proposals fail diligence.
- **Narrative/Communication quality.** For a *proposal* review tool, no judge assesses whether the
document actually persuades.
**Fix:** add Technical Architecture (merge into Execution-Feasibility if headcount is capped) and fold
an evidence-substantiation mandate into the Research Agent's brief so grounding produces a scored
claims-verification artifact, not just citations.
---
## 2. Model Assignments - 6/10 - FIX
Most seats are defensible. Grok 4.5 on Research is correct (only pool member with native live web
search). Qwen3.7 Plus on Market-Reality for a non-Western lens is genuinely thoughtful.
**Empirically-grounded objections:**
**Kimi K2.6 on Reasoning-Verification is misassigned.** In my scoring test it returned **empty content
after burning all 900 output tokens** - and on the 12-page critical-path run it took 15.1s. The role
requires emitting a structured contradiction report; a model that silently exhausts its budget in the
seat that validates every other judge's consistency is the worst possible placement. It is also the
*only* Moonshot seat, so there is no same-vendor fallback.
**Empty-response rate is a systemic risk the spec never models.** At `max_tokens=12`, four of six probed
models (MiniMax-M3, Kimi K2.6, Claude Fable 5, Gemini Pro) returned **empty content** - reasoning tokens
consumed the entire budget. §4.1 handles "malformed score" but not "well-formed empty response," which is
the actual failure mode I observed. Token budgets must be set per-model with reasoning-token headroom.
**Claude Sonnet 5 sits in two scoring seats** (Validation + Legal/Regulatory). The spec waves this off as
"read/analytical, not scoring-intense," but §1.1 marks Legal/Regulatory as **"Yes - scores."** The
justification contradicts the table two rows above it. Same weights, same biases, two votes.
**Gemini Pro Latest doubles as Cross-Check B and Audit Agent.** The Audit Agent's stated job is catching
*groupthink* - it cannot audit a panel it already voted in. This directly violates the spirit of §3.1's
own audit-independence rule.
---
## 3. Vendor Diversity - 8/10 - DEFEND
The strongest dimension. Eight vendors, max 25% - a genuine improvement over v1.0's 44% Anthropic
violation, and the rule was tightened from 40% to 33% rather than loosened to fit. That is the right
instinct and the spec should defend it.
**Two caveats worth documenting rather than fixing:**
*Vendor diversity is not architecture diversity.* Nearly every pool member is a transformer trained on
overlapping web corpora with similar RLHF conventions. My independence test showed all three "independent"
cross-checks scoring `market` at **exactly 6 - spread 0.00**. Eight logos, one prior. The correlation
data from the Critical Finding is the real story here.
*Infrastructure concentration.* All 12 models route through a single admin-ai/LiteLLM proxy. Eight-vendor
diversity buys nothing if the proxy is down - that is the actual SPOF, and the spec's §4.1 fallback table
implicitly assumes the proxy always answers.
---
## 4. admin-ai Paths - 9/10 - DEFEND
**Verified live, not taken on faith.** I queried the admin-ai catalog (157 models) and confirmed all 13
spec paths resolve exactly, then ran real inference against every one:
```
OK xai/grok-4.5 OK claude-opus-5 OK claude-sonnet-5
OK gpt-5.6-sol OK gemini/gemini-pro-latest OK deepseek-v4-pro
OK kimi-k2.6 OK gpt-5.6-terra OK qwen3.7-plus
OK MiniMax-M3 OK claude-fable-5 OK gemini/gemini-3.6-flash
OK gemini/gemini-3.5-flash-lite
```
All 11 judge models returned live responses. The spec's claim "all model paths confirmed against admin-ai"
is **true** - rare enough in a v2.0 draft to call out. This dimension should be defended as-is.
**The one point off:** `gemini/gemini-pro-latest` is a **floating alias**, not a pinned version. The
catalog carries pinned alternatives (`gemini/gemini-3.1-pro-preview`, `openrouter/google/gemini-3-pro-preview`).
Google can repoint that alias with no notice and silently change two seats - Cross-Check B *and* Audit.
My determinism probe on that alias returned `[6,3,3] / [6,4,3] / [6,4,2]` across three identical calls.
For a product whose entire value is score reproducibility, and which promises T+90/180/365 longitudinal
outcome tracking, an unpinned model destroys year-over-year comparability. **Pin every judge seat.**
---
## 5. Fallback Chains - 4/10 - FIX
§3.1 asserts "Fallback chains always cross vendor boundaries - Enforced at the assignment table (§2.3)."
Reading §2.3 against that claim, the rule is **violated in the first fallback hop of 4 of 11 rows**:
| Role | Primary | Fallback 1 | Violation |
|---|---|---|---|
| Primary Reviewer | Claude Opus 5 | **Claude Sonnet 5** | Anthropic → Anthropic |
| Cross-Check A | GPT-5.6 Sol | **GPT-5.6 Terra** | OpenAI → OpenAI |
| Cross-Check B | Gemini Pro Latest | **Gemini 3.6 Flash** | Google → Google |
| Team/Founder | Claude Fable 5 | **Claude Sonnet 5** | Anthropic → Anthropic |
This is the *identical* defect flagged in the v1.0 review ("§4.1 says different vendor but §2.3 says
Opus→Sonnet"). It was marked as a top-5 prioritized fix, and it was **not fixed** - the assertion text was
added to §3.1 without correcting the table it points at. A rule that is stated but not enforced is worse
than no rule: it will pass code review as "already handled."
The failure mode is precisely what fallbacks exist to prevent. Anthropic has a regional outage → Primary
Reviewer fails over to Sonnet 5 → also Anthropic → also down. Same for the OpenAI and Google rows.
**Additional defects:**
- **Convergence, not diversity.** DeepSeek V4 Pro appears as a fallback in **7 of 11 chains**. Under
broad degradation the 8-vendor panel collapses toward a single DeepSeek-dominated panel, blowing the
33% vendor cap at exactly the moment it matters most. No runtime re-check of the cap after failover.
- **Kimi K2.6 has no same-tier peer** - the sole Moonshot seat degrades straight to DeepSeek (Tier C).
- **§4.2 is unexecutable as written.** "Tier B for Primary/Validation only, Tier C for others" - Tier C
is `DeepSeek V4 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite`: two vendors, three models, for up to
five simultaneous seats. The cost-optimization path *cannot* satisfy the vendor rule. This is the same
unexecutability flagged in v1.0 ("Tier B has no OpenAI/Google model") reappearing in a new form.
---
## 6. Content-Based Rules - 5/10 - FIX
The vertical/stage triggers are directionally sensible and v2.0 adds a genuine improvement: the §3.3
missing-vertical handler with explicit user disclosure is honest product design.
**But the weighting system is undefined.** Every rule adjusts "weight ±N%" and **the spec never states
what the weights weight.** There is no aggregation formula anywhere in 302 lines - no statement of how 11
judges' 10-dimension scores combine into a final number. "Market-Reality weight +30%" is meaningless
without knowing the base weight, the combination function, and whether weights renormalize. This is the
single largest specification hole in the document: it is the core scoring algorithm, and it is absent.
**Composition rule is underspecified.** §3.9 caps stacking at +50% per role but doesn't say what happens
when the cap binds - are all triggers scaled proportionally, or does first-match win? Different answers
give different scores for the same proposal.
**Detection is hand-waved.** Every row says "Detected from proposal text" with no mechanism, no confidence
threshold, and no misclassification path. Vertical detection *changes the score* - a fintech misread as
SaaS loses its +30% Legal/Regulatory weight. That needs a confidence score and a human-review fallback,
and multi-vertical proposals (fintech + healthcare) have no representation at all.
**Page count is a poor complexity proxy.** A 10-page dense technical proposal triggers the *reduced*
panel; a 50-page deck with 40 pages of appendix triggers the full one. Use token count or content
density.
**Substantively questionable:** Seed-stage sets Financial Integrity **-20%**. Seed is where financial
models are *most* fictional and cap-table mistakes are most permanent. De-weighting financial scrutiny
where founders most need it inverts the diligence priority.
---
## 7. Latency Budgets - 2/10 - FIX (launch-blocking)
The spec revised Enterprise from 30s → 45s and annotated it *"v1.0 was unrealistic for 12-judge pipeline."*
The revision is still off by nearly 5x. I measured it.
**Sequential critical path, 12-page proposal (~5,500 tokens), following the spec's own dependency graph:**
```
P1 Research xai/grok-4.5 21.5s
P2 Primary claude-opus-5 20.0s
P3 Validation claude-sonnet-5 15.8s
P5 Reasoning kimi-k2.6 15.1s
P6 Execution gpt-5.6-terra 16.4s
P7 Market qwen3.7-plus 67.2s <-- single judge exceeds Enterprise MAX alone
P8 Financial MiniMax-M3 26.1s
P9 Team claude-fable-5 22.9s
P11 Audit gemini/gemini-pro-latest 13.0s
------------------------------------------------
TOTAL 218.0s
```
**218s against a 45s target and a 90s max - 4.8x over target, 2.4x over the stated maximum.**
This is not a tuning problem, it is structural:
- **The pipeline is inherently sequential.** Phases 2→3→5→6→7→8→9→11 each consume prior output by design.
Only 4a/4b/4c parallelize. You cannot fan out a dependency chain.
- **A single judge blows the entire budget.** Qwen3.7 Plus took **67.2s alone** - 1.5x the full Enterprise
target, before any other judge runs. It burned 1,450 reasoning tokens on a trivial scoring task.
- **Even perfect parallelism fails.** All 11 models fired simultaneously on a *short* prompt still took
**27.4s wall clock** - 61% of the Enterprise budget consumed by the slowest model, with zero
orchestration, retries, or assembly.
- **Fallbacks make it worse.** A 30s timeout + Fallback 1 retry adds 30s+ to an already-blown budget.
The §4.1 retry path and the §4.3 budget are mutually unsatisfiable.
- **Tier ordering is still inverted.** Enterprise runs the deepest pipeline (12 judges) on the tightest
budget (45s); Free runs 3 judges on 90s. v1.0 had this backwards and v2.0 preserved the inversion while
only adjusting magnitudes.
50+ page proposals (the spec's own "High complexity," which *adds* a secondary pass) will run
substantially past 218s.
**Fix:** these are not real-time interactions - they are deep-analysis jobs. Re-architect as an
**asynchronous job model**: submit → progress streaming → notify on completion. Budget **5-10 minutes**
for Enterprise and sell the depth. Then set per-phase timeouts against measured p95, not aspiration.
A 45s promise that reliably takes 218s is a support-ticket generator and a churn driver.
---
## 8. Feedback Loop - 5/10 - FIX
v2.0 genuinely improved here - §5.4 cold-start gate, min N=30/N=50 thresholds, and the immediate-recusal
rule for >2.0σ bias are all correct additions that address prior findings.
**The remaining problems are foundational:**
**The accuracy metric is not sound.** "Correlation between dimension score and T+90 outcome" -
- **Ninety days is far too short.** Seed rounds take 3-9 months; the T+90 signal is mostly noise about
fundraising *timing*, not proposal quality.
- **Survivorship and selection bias are unaddressed.** Response rates on outcome surveys skew heavily to
founders who succeeded. The spec plans to weight models on a systematically biased sample.
- **Confounding is total.** A proposal that scores 4/10, gets rewritten using the Fix-It plan, and then
raises successfully - did the judge score correctly or incorrectly? The product *intervenes* on the
outcome it measures. This is unfixable by more data; it needs a holdout design.
- **N is unreachable.** N≥50 per model per dimension per vertical, with 10 dimensions, 11 models and 5+
verticals, implies thousands of tracked reviews before a single threshold fires. At 100 reviews/mo
Enterprise capacity, that is **years**. The entire feedback loop is aspirational at realistic volume,
and the "data moat" narrative rests on it.
**Rotation rules:** "Remove from Tier A, demote to Tier B" as a demotion path is odd - a model with
<0.3 outcome correlation is not a *cheaper* model, it is an *inaccurate* one. Demoting it means budget
users get the judge known to be wrong. Also, the model-deprecation row says replace with "Fallback 1
from §2.3" - but §2.3 fallbacks are same-vendor in 4 rows, so vendor-caused deprecation cascades to a
sibling that may be deprecated in the same wave.
---
## 9. Pricing - 4/10 - FIX
I have no objection to premium positioning, and the "don't race to the bottom" instinct is right. The
objection is that **the price is not connected to demonstrated value**, and v2.0 raised it by
**3.3x-5.0x** (from $79/$299 to $249/$799/$1,499) on positioning reasoning alone, with zero customers,
zero LOIs, and - per the Critical Finding - no evidence the 12-judge panel beats one model.
**The benchmark is misapplied.** GC AI at $500/seat/mo is cited as the anchor comp, and I verified it
independently (gc.ai, corroborated by vaquill.ai's 2026 benchmark; note haqq.ai could not source it to
GC AI directly). But GC AI serves **in-house legal teams at 1,900+ companies** - daily-use workflow
software with seat-level lock-in. VerdictTank is **episodic**: a founder reviews a proposal during a
fundraise, then churns. Anchoring episodic tooling to daily-workflow pricing is a category error.
**The usage math undermines the tiers.** Enterprise at $799/mo = 100 reviews. Real founders raising a
round need **3-8 reviews over a 2-3 month window**. That is ~$100/review nominal at a utilization
almost no customer will reach - and Pro at $249 for 20 reviews has the same problem. Customers pay for
capacity they cannot consume, notice, and churn. **Per-review or credit-pack pricing fits actual
consumption far better than monthly seats**, and the spec never considers it.
**Margin honesty.** The v4.1 reference concedes "COGS 72-97% margin - pure positioning play." My cost
sampling supports that: a full panel run is roughly $0.30-0.80 in tokens. A 99.9% gross margin at
$799/mo is not premium positioning, it is an unanchored price waiting for a competitor to undercut it
with the same off-the-shelf models. The moat is claimed to be the outcome-tracking corpus - which
dimension 8 shows is years away at this volume.
**Free tier at 1/mo is too stingy** for a trust-first product. The entire pitch is "we tell you the
brutal truth" - that requires *experiencing* the depth. A 1-review Tier-C sample (3 judges, no Fix-It)
demonstrates the weakest possible version of the product to every prospective buyer.
**Fix:** validate willingness-to-pay with 10-20 design partners before locking. Offer per-review pricing
alongside subscriptions. Anchor to *outcome value* (a better raise) rather than to a competitor in an
adjacent category.
---
## 10. Product Boundary - 7/10 - DEFER
Conceptually clean and easy to communicate: VerdictTank = "is this good?", RFP Tank = "does this match
what they asked for?" The superset framing is right, and inheriting one engine is the correct build
decision.
**Unresolved, but not urgent:**
- **RFP Tank is the better business and it is the side project.** RFP responses are recurring, deadline-
driven, budgeted, and B2B - structurally superior to episodic founder fundraising. The spec treats it
as a discount add-on. Strategically inverted.
- **Cannibalization is unpriced.** RFP Tank is a strict superset at (presumably) a higher price. A
rational buyer needing both buys RFP Tank only. Enterprise VT + 20% off RFP Tank is then a discount
on a product that replaces the one just paid for.
- **10-dimension rubric is asserted, never enumerated.** The document references "10-dimension scoring"
in §7 and §8 but **never lists the ten dimensions.** For a build spec, the thing being scored should
be defined; I am scoring against dimensions the spec assumes I already know.
- No shared-account model, no cross-product SSO, no migration path.
**DEFER** because the boundary is directionally correct and none of this blocks the judge-pool build.
Revisit before RFP Tank pricing is set.
---
## Priority Actions
| # | Action | Dim | Severity |
|---|---|---|---|
| 1 | **Validate panel-vs-single-model accuracy on 30-50 real proposals.** Product thesis is unproven. | - | **BLOCKER** |
| 2 | **Re-architect to async jobs; budget 5-10 min.** Measured 218s vs 45s target. | 7 | **BLOCKER** |
| 3 | **Fix 4 same-vendor Fallback-1 hops.** Flagged in v1.0 review, still unfixed. | 5 | **BLOCKER** |
| 4 | **Define the score aggregation formula.** "Weight +30%" is meaningless; core algorithm absent. | 6 | **BLOCKER** |
| 5 | Pin all model versions - drop floating `gemini-pro-latest` alias. | 4 | High |
| 6 | Reassign Kimi K2.6 off Reasoning-Verification (empty output under budget). | 2 | High |
| 7 | Set per-model token budgets with reasoning-token headroom; handle empty-but-valid responses. | 2 | High |
| 8 | Split Gemini Pro's Cross-Check B / Audit double-seat; auditor cannot audit itself. | 2 | High |
| 9 | Add runtime vendor-cap re-check after failover (DeepSeek is fallback in 7 of 11 chains). | 5 | High |
| 10 | Validate pricing with design partners; add per-review option. | 9 | High |
| 11 | Add Technical Architecture + evidence-substantiation coverage. | 1 | Medium |
| 12 | Replace page-count complexity proxy with token count. | 6 | Medium |
| 13 | Enumerate the 10 scoring dimensions in the spec. | 10 | Medium |
| 14 | Reconsider Financial Integrity -20% at seed stage. | 6 | Medium |
---
## Bottom Line
v2.0 is a real improvement over v1.0 - the vendor rule was tightened rather than loosened to fit, the
cold-start gates are correct, the role gaps were mostly closed, and **every model path is genuinely live**,
which I verified rather than assumed. The document is architecturally literate.
But it is a **design document wearing a build spec's clothes**. Its two most load-bearing quantitative
claims fail on contact with the live system: the latency budget is off by 4.8x, and the fallback
cross-vendor guarantee is contradicted by its own assignment table in four rows - *the same defect flagged
in the v1.0 review and marked as a prioritized fix.* The assertion was added; the table was not corrected.
Most seriously: the panel-diversity premise that justifies the pricing, the vendor spread, and the entire
category claim has **never been tested**, and my measurement suggests it may not survive testing. Nine
judges across eight vendors moved the score by under half a point versus a single model.
**Recommendation: do not proceed to build.** Run the accuracy validation (Action 1) and re-baseline
latency (Action 2) first. If the panel advantage is real, this architecture is worth building and the
premium price defensible. If it is not, the correct product is 3 judges at a third of the price - and it
is far cheaper to discover that now than after the first enterprise customer runs the same A/B I just ran.