Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs - disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md - clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy) - projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture - proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review - docs/super-search/firecrawl-provider-strategy.md - updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io - .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
This commit is contained in:
@@ -0,0 +1,411 @@
|
||||
# VerdictTank Judge Pool Spec v2.0 - Primary Reviewer Scorecard
|
||||
|
||||
**Reviewer:** Claude Opus 5 (Primary Reviewer, Phase 2)
|
||||
**Date:** 2026-08-12
|
||||
**Spec:** https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md (17,641 bytes, v2.0)
|
||||
**Method:** Document critique + live empirical validation against admin-ai (157-model catalog, real inference runs)
|
||||
|
||||
> **Note on version:** `web_extract` returned a cached **v1.0**. Verified against the live origin
|
||||
> via `curl` - the deployed file is **v2.0**, byte-identical to
|
||||
> `/root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md`. This review is of v2.0.
|
||||
|
||||
---
|
||||
|
||||
## Scorecard
|
||||
|
||||
| # | Dimension | Score | Verdict |
|
||||
|---|---|---|---|
|
||||
| 1 | Role Coverage | 7/10 | FIX |
|
||||
| 2 | Model Assignments | 6/10 | FIX |
|
||||
| 3 | Vendor Diversity | 8/10 | DEFEND |
|
||||
| 4 | admin-ai Paths | 9/10 | DEFEND |
|
||||
| 5 | Fallback Chains | 4/10 | FIX |
|
||||
| 6 | Content-Based Rules | 5/10 | FIX |
|
||||
| 7 | Latency Budgets | 2/10 | FIX |
|
||||
| 8 | Feedback Loop | 5/10 | FIX |
|
||||
| 9 | Pricing | 4/10 | FIX |
|
||||
| 10 | Product Boundary | 7/10 | DEFER |
|
||||
|
||||
**Weighted mean: 5.7/10** - Architecturally literate, empirically unvalidated. Two dimensions (7, 9) are
|
||||
launch-blocking. One unlisted finding (§Critical Finding) threatens the product thesis itself.
|
||||
|
||||
---
|
||||
|
||||
## CRITICAL FINDING (unlisted dimension - read first)
|
||||
|
||||
**The 12-judge panel does not measurably outperform one model run once.**
|
||||
|
||||
I tested this directly. Same proposal, same rubric, three dimensions, live models.
|
||||
|
||||
**A) Single-model noise floor** - `claude-opus-5` x6 at temp 1.0:
|
||||
```
|
||||
market: 5,5,5,4,4,4 stdev 0.50
|
||||
team: 4,4,4,4,4,4 stdev 0.00
|
||||
fin: 3,3,3,3,3,3 stdev 0.00
|
||||
```
|
||||
|
||||
**B) Nine distinct judges across 8 vendors, one run each:**
|
||||
```
|
||||
Opus5 [5,4,3] Sonnet5 [4,4,3] Sol [6,4,3] GeminiPro [4,3,2] DeepSeek [5,3,2]
|
||||
Terra [4,4,3] Qwen [4,3,2] MiniMax [6,3,4] Fable5 [6,4,4]
|
||||
panel stdev: market 0.87, team 0.50, fin 0.74
|
||||
```
|
||||
|
||||
**C) The result that matters:**
|
||||
|
||||
| | market | team | financials |
|
||||
|---|---|---|---|
|
||||
| Mean, 1 model | 4.50 | 4.00 | 3.00 |
|
||||
| Mean, 9 judges / 8 vendors | 4.89 | 3.56 | 2.89 |
|
||||
| **Delta** | **0.39** | **0.44** | **0.11** |
|
||||
|
||||
Nine frontier models from eight vendors, ~218s of wall clock and roughly 20x the token cost, move the
|
||||
final number by **less than half a point on a 10-point scale**. On `market`, panel spread (0.87) is only
|
||||
**1.7x** the single-model rerun noise (0.50) - statistically indistinguishable from rerunning one model
|
||||
six times.
|
||||
|
||||
The individual judges *do* disagree (market scores span 4-6). But disagreement that averages to the same
|
||||
answer is **variance, not signal**. The spec sells vendor diversity as the core accuracy mechanism and
|
||||
never once tests whether diversity changes the output.
|
||||
|
||||
**This is the product thesis, and it is unvalidated.** Every downstream claim - premium pricing, the
|
||||
"most accurate review tool available" positioning, the 12-vendor moat - rests on it.
|
||||
|
||||
**Required before build:** run 30-50 real proposals with known outcomes. Report panel-vs-single
|
||||
correlation against ground truth. If the panel does not beat one good model by a margin that justifies
|
||||
20x cost, the correct architecture is 3 judges, not 12 - and the pricing story needs rebuilding.
|
||||
Better to learn this now than after a customer runs the same A/B.
|
||||
|
||||
---
|
||||
|
||||
## 1. Role Coverage - 7/10 - FIX
|
||||
|
||||
v2.0 deserves credit: it closed three of the four gaps flagged in the prior review (Financial Integrity,
|
||||
Team/Founder, Legal/Regulatory now have dedicated phases). That is real progress.
|
||||
|
||||
**Technical Architecture remains uncovered** - the one gap explicitly identified in the v4.1 gap analysis
|
||||
and silently dropped. Execution-Feasibility is operational (timeline, team, resources), not architectural
|
||||
(stack, scalability, security posture, technical debt). For a product whose buyers are evaluating
|
||||
*technical* startups, having no judge that reads the architecture is a conspicuous hole.
|
||||
|
||||
Two further gaps neither version names:
|
||||
- **Traction/Evidence.** Nobody scores whether claims are *substantiated*. Research Agent gathers
|
||||
citations but never scores; no downstream role is tasked with "the founder asserts 40% MoM growth -
|
||||
is there evidence?" This is the single most common reason real proposals fail diligence.
|
||||
- **Narrative/Communication quality.** For a *proposal* review tool, no judge assesses whether the
|
||||
document actually persuades.
|
||||
|
||||
**Fix:** add Technical Architecture (merge into Execution-Feasibility if headcount is capped) and fold
|
||||
an evidence-substantiation mandate into the Research Agent's brief so grounding produces a scored
|
||||
claims-verification artifact, not just citations.
|
||||
|
||||
---
|
||||
|
||||
## 2. Model Assignments - 6/10 - FIX
|
||||
|
||||
Most seats are defensible. Grok 4.5 on Research is correct (only pool member with native live web
|
||||
search). Qwen3.7 Plus on Market-Reality for a non-Western lens is genuinely thoughtful.
|
||||
|
||||
**Empirically-grounded objections:**
|
||||
|
||||
**Kimi K2.6 on Reasoning-Verification is misassigned.** In my scoring test it returned **empty content
|
||||
after burning all 900 output tokens** - and on the 12-page critical-path run it took 15.1s. The role
|
||||
requires emitting a structured contradiction report; a model that silently exhausts its budget in the
|
||||
seat that validates every other judge's consistency is the worst possible placement. It is also the
|
||||
*only* Moonshot seat, so there is no same-vendor fallback.
|
||||
|
||||
**Empty-response rate is a systemic risk the spec never models.** At `max_tokens=12`, four of six probed
|
||||
models (MiniMax-M3, Kimi K2.6, Claude Fable 5, Gemini Pro) returned **empty content** - reasoning tokens
|
||||
consumed the entire budget. §4.1 handles "malformed score" but not "well-formed empty response," which is
|
||||
the actual failure mode I observed. Token budgets must be set per-model with reasoning-token headroom.
|
||||
|
||||
**Claude Sonnet 5 sits in two scoring seats** (Validation + Legal/Regulatory). The spec waves this off as
|
||||
"read/analytical, not scoring-intense," but §1.1 marks Legal/Regulatory as **"Yes - scores."** The
|
||||
justification contradicts the table two rows above it. Same weights, same biases, two votes.
|
||||
|
||||
**Gemini Pro Latest doubles as Cross-Check B and Audit Agent.** The Audit Agent's stated job is catching
|
||||
*groupthink* - it cannot audit a panel it already voted in. This directly violates the spirit of §3.1's
|
||||
own audit-independence rule.
|
||||
|
||||
---
|
||||
|
||||
## 3. Vendor Diversity - 8/10 - DEFEND
|
||||
|
||||
The strongest dimension. Eight vendors, max 25% - a genuine improvement over v1.0's 44% Anthropic
|
||||
violation, and the rule was tightened from 40% to 33% rather than loosened to fit. That is the right
|
||||
instinct and the spec should defend it.
|
||||
|
||||
**Two caveats worth documenting rather than fixing:**
|
||||
|
||||
*Vendor diversity is not architecture diversity.* Nearly every pool member is a transformer trained on
|
||||
overlapping web corpora with similar RLHF conventions. My independence test showed all three "independent"
|
||||
cross-checks scoring `market` at **exactly 6 - spread 0.00**. Eight logos, one prior. The correlation
|
||||
data from the Critical Finding is the real story here.
|
||||
|
||||
*Infrastructure concentration.* All 12 models route through a single admin-ai/LiteLLM proxy. Eight-vendor
|
||||
diversity buys nothing if the proxy is down - that is the actual SPOF, and the spec's §4.1 fallback table
|
||||
implicitly assumes the proxy always answers.
|
||||
|
||||
---
|
||||
|
||||
## 4. admin-ai Paths - 9/10 - DEFEND
|
||||
|
||||
**Verified live, not taken on faith.** I queried the admin-ai catalog (157 models) and confirmed all 13
|
||||
spec paths resolve exactly, then ran real inference against every one:
|
||||
|
||||
```
|
||||
OK xai/grok-4.5 OK claude-opus-5 OK claude-sonnet-5
|
||||
OK gpt-5.6-sol OK gemini/gemini-pro-latest OK deepseek-v4-pro
|
||||
OK kimi-k2.6 OK gpt-5.6-terra OK qwen3.7-plus
|
||||
OK MiniMax-M3 OK claude-fable-5 OK gemini/gemini-3.6-flash
|
||||
OK gemini/gemini-3.5-flash-lite
|
||||
```
|
||||
|
||||
All 11 judge models returned live responses. The spec's claim "all model paths confirmed against admin-ai"
|
||||
is **true** - rare enough in a v2.0 draft to call out. This dimension should be defended as-is.
|
||||
|
||||
**The one point off:** `gemini/gemini-pro-latest` is a **floating alias**, not a pinned version. The
|
||||
catalog carries pinned alternatives (`gemini/gemini-3.1-pro-preview`, `openrouter/google/gemini-3-pro-preview`).
|
||||
Google can repoint that alias with no notice and silently change two seats - Cross-Check B *and* Audit.
|
||||
My determinism probe on that alias returned `[6,3,3] / [6,4,3] / [6,4,2]` across three identical calls.
|
||||
For a product whose entire value is score reproducibility, and which promises T+90/180/365 longitudinal
|
||||
outcome tracking, an unpinned model destroys year-over-year comparability. **Pin every judge seat.**
|
||||
|
||||
---
|
||||
|
||||
## 5. Fallback Chains - 4/10 - FIX
|
||||
|
||||
§3.1 asserts "Fallback chains always cross vendor boundaries - Enforced at the assignment table (§2.3)."
|
||||
Reading §2.3 against that claim, the rule is **violated in the first fallback hop of 4 of 11 rows**:
|
||||
|
||||
| Role | Primary | Fallback 1 | Violation |
|
||||
|---|---|---|---|
|
||||
| Primary Reviewer | Claude Opus 5 | **Claude Sonnet 5** | Anthropic → Anthropic |
|
||||
| Cross-Check A | GPT-5.6 Sol | **GPT-5.6 Terra** | OpenAI → OpenAI |
|
||||
| Cross-Check B | Gemini Pro Latest | **Gemini 3.6 Flash** | Google → Google |
|
||||
| Team/Founder | Claude Fable 5 | **Claude Sonnet 5** | Anthropic → Anthropic |
|
||||
|
||||
This is the *identical* defect flagged in the v1.0 review ("§4.1 says different vendor but §2.3 says
|
||||
Opus→Sonnet"). It was marked as a top-5 prioritized fix, and it was **not fixed** - the assertion text was
|
||||
added to §3.1 without correcting the table it points at. A rule that is stated but not enforced is worse
|
||||
than no rule: it will pass code review as "already handled."
|
||||
|
||||
The failure mode is precisely what fallbacks exist to prevent. Anthropic has a regional outage → Primary
|
||||
Reviewer fails over to Sonnet 5 → also Anthropic → also down. Same for the OpenAI and Google rows.
|
||||
|
||||
**Additional defects:**
|
||||
- **Convergence, not diversity.** DeepSeek V4 Pro appears as a fallback in **7 of 11 chains**. Under
|
||||
broad degradation the 8-vendor panel collapses toward a single DeepSeek-dominated panel, blowing the
|
||||
33% vendor cap at exactly the moment it matters most. No runtime re-check of the cap after failover.
|
||||
- **Kimi K2.6 has no same-tier peer** - the sole Moonshot seat degrades straight to DeepSeek (Tier C).
|
||||
- **§4.2 is unexecutable as written.** "Tier B for Primary/Validation only, Tier C for others" - Tier C
|
||||
is `DeepSeek V4 Pro, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite`: two vendors, three models, for up to
|
||||
five simultaneous seats. The cost-optimization path *cannot* satisfy the vendor rule. This is the same
|
||||
unexecutability flagged in v1.0 ("Tier B has no OpenAI/Google model") reappearing in a new form.
|
||||
|
||||
---
|
||||
|
||||
## 6. Content-Based Rules - 5/10 - FIX
|
||||
|
||||
The vertical/stage triggers are directionally sensible and v2.0 adds a genuine improvement: the §3.3
|
||||
missing-vertical handler with explicit user disclosure is honest product design.
|
||||
|
||||
**But the weighting system is undefined.** Every rule adjusts "weight ±N%" and **the spec never states
|
||||
what the weights weight.** There is no aggregation formula anywhere in 302 lines - no statement of how 11
|
||||
judges' 10-dimension scores combine into a final number. "Market-Reality weight +30%" is meaningless
|
||||
without knowing the base weight, the combination function, and whether weights renormalize. This is the
|
||||
single largest specification hole in the document: it is the core scoring algorithm, and it is absent.
|
||||
|
||||
**Composition rule is underspecified.** §3.9 caps stacking at +50% per role but doesn't say what happens
|
||||
when the cap binds - are all triggers scaled proportionally, or does first-match win? Different answers
|
||||
give different scores for the same proposal.
|
||||
|
||||
**Detection is hand-waved.** Every row says "Detected from proposal text" with no mechanism, no confidence
|
||||
threshold, and no misclassification path. Vertical detection *changes the score* - a fintech misread as
|
||||
SaaS loses its +30% Legal/Regulatory weight. That needs a confidence score and a human-review fallback,
|
||||
and multi-vertical proposals (fintech + healthcare) have no representation at all.
|
||||
|
||||
**Page count is a poor complexity proxy.** A 10-page dense technical proposal triggers the *reduced*
|
||||
panel; a 50-page deck with 40 pages of appendix triggers the full one. Use token count or content
|
||||
density.
|
||||
|
||||
**Substantively questionable:** Seed-stage sets Financial Integrity **-20%**. Seed is where financial
|
||||
models are *most* fictional and cap-table mistakes are most permanent. De-weighting financial scrutiny
|
||||
where founders most need it inverts the diligence priority.
|
||||
|
||||
---
|
||||
|
||||
## 7. Latency Budgets - 2/10 - FIX (launch-blocking)
|
||||
|
||||
The spec revised Enterprise from 30s → 45s and annotated it *"v1.0 was unrealistic for 12-judge pipeline."*
|
||||
The revision is still off by nearly 5x. I measured it.
|
||||
|
||||
**Sequential critical path, 12-page proposal (~5,500 tokens), following the spec's own dependency graph:**
|
||||
|
||||
```
|
||||
P1 Research xai/grok-4.5 21.5s
|
||||
P2 Primary claude-opus-5 20.0s
|
||||
P3 Validation claude-sonnet-5 15.8s
|
||||
P5 Reasoning kimi-k2.6 15.1s
|
||||
P6 Execution gpt-5.6-terra 16.4s
|
||||
P7 Market qwen3.7-plus 67.2s <-- single judge exceeds Enterprise MAX alone
|
||||
P8 Financial MiniMax-M3 26.1s
|
||||
P9 Team claude-fable-5 22.9s
|
||||
P11 Audit gemini/gemini-pro-latest 13.0s
|
||||
------------------------------------------------
|
||||
TOTAL 218.0s
|
||||
```
|
||||
|
||||
**218s against a 45s target and a 90s max - 4.8x over target, 2.4x over the stated maximum.**
|
||||
|
||||
This is not a tuning problem, it is structural:
|
||||
|
||||
- **The pipeline is inherently sequential.** Phases 2→3→5→6→7→8→9→11 each consume prior output by design.
|
||||
Only 4a/4b/4c parallelize. You cannot fan out a dependency chain.
|
||||
- **A single judge blows the entire budget.** Qwen3.7 Plus took **67.2s alone** - 1.5x the full Enterprise
|
||||
target, before any other judge runs. It burned 1,450 reasoning tokens on a trivial scoring task.
|
||||
- **Even perfect parallelism fails.** All 11 models fired simultaneously on a *short* prompt still took
|
||||
**27.4s wall clock** - 61% of the Enterprise budget consumed by the slowest model, with zero
|
||||
orchestration, retries, or assembly.
|
||||
- **Fallbacks make it worse.** A 30s timeout + Fallback 1 retry adds 30s+ to an already-blown budget.
|
||||
The §4.1 retry path and the §4.3 budget are mutually unsatisfiable.
|
||||
- **Tier ordering is still inverted.** Enterprise runs the deepest pipeline (12 judges) on the tightest
|
||||
budget (45s); Free runs 3 judges on 90s. v1.0 had this backwards and v2.0 preserved the inversion while
|
||||
only adjusting magnitudes.
|
||||
|
||||
50+ page proposals (the spec's own "High complexity," which *adds* a secondary pass) will run
|
||||
substantially past 218s.
|
||||
|
||||
**Fix:** these are not real-time interactions - they are deep-analysis jobs. Re-architect as an
|
||||
**asynchronous job model**: submit → progress streaming → notify on completion. Budget **5-10 minutes**
|
||||
for Enterprise and sell the depth. Then set per-phase timeouts against measured p95, not aspiration.
|
||||
A 45s promise that reliably takes 218s is a support-ticket generator and a churn driver.
|
||||
|
||||
---
|
||||
|
||||
## 8. Feedback Loop - 5/10 - FIX
|
||||
|
||||
v2.0 genuinely improved here - §5.4 cold-start gate, min N=30/N=50 thresholds, and the immediate-recusal
|
||||
rule for >2.0σ bias are all correct additions that address prior findings.
|
||||
|
||||
**The remaining problems are foundational:**
|
||||
|
||||
**The accuracy metric is not sound.** "Correlation between dimension score and T+90 outcome" -
|
||||
- **Ninety days is far too short.** Seed rounds take 3-9 months; the T+90 signal is mostly noise about
|
||||
fundraising *timing*, not proposal quality.
|
||||
- **Survivorship and selection bias are unaddressed.** Response rates on outcome surveys skew heavily to
|
||||
founders who succeeded. The spec plans to weight models on a systematically biased sample.
|
||||
- **Confounding is total.** A proposal that scores 4/10, gets rewritten using the Fix-It plan, and then
|
||||
raises successfully - did the judge score correctly or incorrectly? The product *intervenes* on the
|
||||
outcome it measures. This is unfixable by more data; it needs a holdout design.
|
||||
- **N is unreachable.** N≥50 per model per dimension per vertical, with 10 dimensions, 11 models and 5+
|
||||
verticals, implies thousands of tracked reviews before a single threshold fires. At 100 reviews/mo
|
||||
Enterprise capacity, that is **years**. The entire feedback loop is aspirational at realistic volume,
|
||||
and the "data moat" narrative rests on it.
|
||||
|
||||
**Rotation rules:** "Remove from Tier A, demote to Tier B" as a demotion path is odd - a model with
|
||||
<0.3 outcome correlation is not a *cheaper* model, it is an *inaccurate* one. Demoting it means budget
|
||||
users get the judge known to be wrong. Also, the model-deprecation row says replace with "Fallback 1
|
||||
from §2.3" - but §2.3 fallbacks are same-vendor in 4 rows, so vendor-caused deprecation cascades to a
|
||||
sibling that may be deprecated in the same wave.
|
||||
|
||||
---
|
||||
|
||||
## 9. Pricing - 4/10 - FIX
|
||||
|
||||
I have no objection to premium positioning, and the "don't race to the bottom" instinct is right. The
|
||||
objection is that **the price is not connected to demonstrated value**, and v2.0 raised it by
|
||||
**3.3x-5.0x** (from $79/$299 to $249/$799/$1,499) on positioning reasoning alone, with zero customers,
|
||||
zero LOIs, and - per the Critical Finding - no evidence the 12-judge panel beats one model.
|
||||
|
||||
**The benchmark is misapplied.** GC AI at $500/seat/mo is cited as the anchor comp, and I verified it
|
||||
independently (gc.ai, corroborated by vaquill.ai's 2026 benchmark; note haqq.ai could not source it to
|
||||
GC AI directly). But GC AI serves **in-house legal teams at 1,900+ companies** - daily-use workflow
|
||||
software with seat-level lock-in. VerdictTank is **episodic**: a founder reviews a proposal during a
|
||||
fundraise, then churns. Anchoring episodic tooling to daily-workflow pricing is a category error.
|
||||
|
||||
**The usage math undermines the tiers.** Enterprise at $799/mo = 100 reviews. Real founders raising a
|
||||
round need **3-8 reviews over a 2-3 month window**. That is ~$100/review nominal at a utilization
|
||||
almost no customer will reach - and Pro at $249 for 20 reviews has the same problem. Customers pay for
|
||||
capacity they cannot consume, notice, and churn. **Per-review or credit-pack pricing fits actual
|
||||
consumption far better than monthly seats**, and the spec never considers it.
|
||||
|
||||
**Margin honesty.** The v4.1 reference concedes "COGS 72-97% margin - pure positioning play." My cost
|
||||
sampling supports that: a full panel run is roughly $0.30-0.80 in tokens. A 99.9% gross margin at
|
||||
$799/mo is not premium positioning, it is an unanchored price waiting for a competitor to undercut it
|
||||
with the same off-the-shelf models. The moat is claimed to be the outcome-tracking corpus - which
|
||||
dimension 8 shows is years away at this volume.
|
||||
|
||||
**Free tier at 1/mo is too stingy** for a trust-first product. The entire pitch is "we tell you the
|
||||
brutal truth" - that requires *experiencing* the depth. A 1-review Tier-C sample (3 judges, no Fix-It)
|
||||
demonstrates the weakest possible version of the product to every prospective buyer.
|
||||
|
||||
**Fix:** validate willingness-to-pay with 10-20 design partners before locking. Offer per-review pricing
|
||||
alongside subscriptions. Anchor to *outcome value* (a better raise) rather than to a competitor in an
|
||||
adjacent category.
|
||||
|
||||
---
|
||||
|
||||
## 10. Product Boundary - 7/10 - DEFER
|
||||
|
||||
Conceptually clean and easy to communicate: VerdictTank = "is this good?", RFP Tank = "does this match
|
||||
what they asked for?" The superset framing is right, and inheriting one engine is the correct build
|
||||
decision.
|
||||
|
||||
**Unresolved, but not urgent:**
|
||||
- **RFP Tank is the better business and it is the side project.** RFP responses are recurring, deadline-
|
||||
driven, budgeted, and B2B - structurally superior to episodic founder fundraising. The spec treats it
|
||||
as a discount add-on. Strategically inverted.
|
||||
- **Cannibalization is unpriced.** RFP Tank is a strict superset at (presumably) a higher price. A
|
||||
rational buyer needing both buys RFP Tank only. Enterprise VT + 20% off RFP Tank is then a discount
|
||||
on a product that replaces the one just paid for.
|
||||
- **10-dimension rubric is asserted, never enumerated.** The document references "10-dimension scoring"
|
||||
in §7 and §8 but **never lists the ten dimensions.** For a build spec, the thing being scored should
|
||||
be defined; I am scoring against dimensions the spec assumes I already know.
|
||||
- No shared-account model, no cross-product SSO, no migration path.
|
||||
|
||||
**DEFER** because the boundary is directionally correct and none of this blocks the judge-pool build.
|
||||
Revisit before RFP Tank pricing is set.
|
||||
|
||||
---
|
||||
|
||||
## Priority Actions
|
||||
|
||||
| # | Action | Dim | Severity |
|
||||
|---|---|---|---|
|
||||
| 1 | **Validate panel-vs-single-model accuracy on 30-50 real proposals.** Product thesis is unproven. | - | **BLOCKER** |
|
||||
| 2 | **Re-architect to async jobs; budget 5-10 min.** Measured 218s vs 45s target. | 7 | **BLOCKER** |
|
||||
| 3 | **Fix 4 same-vendor Fallback-1 hops.** Flagged in v1.0 review, still unfixed. | 5 | **BLOCKER** |
|
||||
| 4 | **Define the score aggregation formula.** "Weight +30%" is meaningless; core algorithm absent. | 6 | **BLOCKER** |
|
||||
| 5 | Pin all model versions - drop floating `gemini-pro-latest` alias. | 4 | High |
|
||||
| 6 | Reassign Kimi K2.6 off Reasoning-Verification (empty output under budget). | 2 | High |
|
||||
| 7 | Set per-model token budgets with reasoning-token headroom; handle empty-but-valid responses. | 2 | High |
|
||||
| 8 | Split Gemini Pro's Cross-Check B / Audit double-seat; auditor cannot audit itself. | 2 | High |
|
||||
| 9 | Add runtime vendor-cap re-check after failover (DeepSeek is fallback in 7 of 11 chains). | 5 | High |
|
||||
| 10 | Validate pricing with design partners; add per-review option. | 9 | High |
|
||||
| 11 | Add Technical Architecture + evidence-substantiation coverage. | 1 | Medium |
|
||||
| 12 | Replace page-count complexity proxy with token count. | 6 | Medium |
|
||||
| 13 | Enumerate the 10 scoring dimensions in the spec. | 10 | Medium |
|
||||
| 14 | Reconsider Financial Integrity -20% at seed stage. | 6 | Medium |
|
||||
|
||||
---
|
||||
|
||||
## Bottom Line
|
||||
|
||||
v2.0 is a real improvement over v1.0 - the vendor rule was tightened rather than loosened to fit, the
|
||||
cold-start gates are correct, the role gaps were mostly closed, and **every model path is genuinely live**,
|
||||
which I verified rather than assumed. The document is architecturally literate.
|
||||
|
||||
But it is a **design document wearing a build spec's clothes**. Its two most load-bearing quantitative
|
||||
claims fail on contact with the live system: the latency budget is off by 4.8x, and the fallback
|
||||
cross-vendor guarantee is contradicted by its own assignment table in four rows - *the same defect flagged
|
||||
in the v1.0 review and marked as a prioritized fix.* The assertion was added; the table was not corrected.
|
||||
|
||||
Most seriously: the panel-diversity premise that justifies the pricing, the vendor spread, and the entire
|
||||
category claim has **never been tested**, and my measurement suggests it may not survive testing. Nine
|
||||
judges across eight vendors moved the score by under half a point versus a single model.
|
||||
|
||||
**Recommendation: do not proceed to build.** Run the accuracy validation (Action 1) and re-baseline
|
||||
latency (Action 2) first. If the panel advantage is real, this architecture is worth building and the
|
||||
premium price defensible. If it is not, the correct product is 3 judges at a third of the price - and it
|
||||
is far cheaper to discover that now than after the first enterprise customer runs the same A/B I just ran.
|
||||
Reference in New Issue
Block a user