Files
root f5175f1ce0 Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
2026-08-26 02:27:28 -04:00

21 KiB
Raw Permalink Blame History

Cross-Check C — DeepSeek V4 Pro Independent Re-Score

Spec: VerdictTank Judge Pool Specification v2.0
Source of truth used: /root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md (local v2.0)
Note: Deployed URL still serves v1.0 — review is against local v2.0 only.
Reviewer lens: Chinese-lab independence (DeepSeek) — different training data, RLHF, and reasoning patterns from Anthropic / OpenAI / Google
Date: 2026-08-12
Mode: Blind independent re-score. FIX / DEFEND / DEFER per dimension.


Executive Summary

Metric Value
Overall 6.4 / 10
Dimensions FIX 5
Dimensions DEFEND 3
Dimensions DEFER 2
Fatal structural issues 2 (fallback vendor-correlation; SaaS-default US-centrism)
Strongest advance vs v1.0 8-vendor spread, Financial/Team/Legal roles, cold-start gate

v2.0 is a real architectural upgrade from the Anthropic-heavy v1.0 panel. The Chinese-ecosystem seats (DeepSeek CC-C, Qwen Market-Reality, Kimi Reasoning, MiniMax Financial) are the right instinct. But the fallback table still hides same-vendor correlation, vertical rules are Western-startup shaped, and "SaaS default" silently erases superapps / WeChat ecosystems / B2B marketplaces that dominate non-US deal flow.


Scorecard (10 Dimensions)

# Dimension Score Verdict One-line
1 Role coverage 7 DEFEND 12 phases close the v1 gaps; still thin on tech architecture + geo-market fit
2 Model assignments 7 DEFEND Chinese seats well placed; Legal on Sonnet is weak; CC-C as Tier C undervalues independence
3 admin-ai paths 6 DEFER Paths look plausible but unproven in this review; Gemini "Pro Latest" alias risk
4 Vendor diversity 8 DEFEND 8 vendors / ≤25% is strong on paper; concentration reappears under fallback
5 Fallback chains 4 FIX Claims "always cross vendor" — 4/12 primaries share vendor with Fallback 1
6 Content-based rules 4 FIX Missing superapp / WeChat / B2B marketplace / cross-border / gov-tech verticals
7 Latency budgets 6 DEFER Better than v1 inverted budgets; 12-judge @ 45s still aspirational without parallelism proof
8 Feedback loop 7 DEFEND Cold-start + N gates good; outcome signal still US-funding-centric
9 Pricing 7 DEFEND Premium hold is coherent; APAC willingness-to-pay and seat economics unaddressed
10 Product boundary + SaaS default 5 FIX VT/RFP split clean; SaaS-default assumption is US-centric and quietly wrong for global deal flow

Mean: 6.1 → weighted overall 6.4 (fallback + verticals + SaaS-default weighted higher under Chinese-lab lens)


Dimension Detail

1. Role Coverage — 7/10 — DEFEND

What works

  • v1.0 left Financial Integrity, Team/Founder, Legal/Regulatory uncovered. v2.0 adds Phases 810. Correct fix.
  • Research never scores; Audit is read-only; three blind cross-checks is a genuine independence architecture.
  • Non-scoring Reasoning-Verification (Phase 5) as synthesis before execution/market is sound sequencing.

What is still thin

  • No dedicated Technical Architecture judge (stack, scalability, security posture). Execution-Feasibility is ops/timeline, not architecture diligence. For deep-tech / infra / AI-infra proposals this is a material hole.
  • No Geo-Market / Localization role. Market-Reality with Qwen helps, but China/SEA/MENA GTM (ICP, channel, regulatory market access) is not a first-class dimension.
  • "10-dimension scoring" in §7 product boundary is never mapped to the 12 phases. Which phases emit which of the 10 scores? Spec is silent → assembly ambiguity.

Verdict: DEFEND the 12-role expansion as directionally correct. Do not expand further before v4.1 ship — but log Technical Architecture + Geo-Market as v4.2 candidates.


2. Model Assignments — 7/10 — DEFEND

Chinese-lab view of seat quality

Role Model Assessment
Research Grok 4.5 Correct — native live web/X is load-bearing
Primary Opus 5 Correct — brutal critique seat
Validation Sonnet 5 Acceptable — cost/quality trade
CC-A GPT-5.6 Sol Correct — OpenAI lineage independence
CC-B Gemini Pro Latest Correct — Google lineage
CC-C DeepSeek V4 Pro Right vendor, wrong tier signal
Reasoning Kimi K2.6 Strong — Moonshot is a real 4th ecosystem
Execution GPT-5.6 Terra Acceptable ops grounding
Market Qwen3.7 Plus Best seat in the roster for non-Western lens
Financial MiniMax-M3 Interesting; unproven for cap-table math specifically
Team Fable 5 Overkill / expensive for founder-fit; also Anthropic stack concentration with Primary/Validation
Legal Sonnet 5 (shared) Mismatch — legal needs specialized calibration + citation discipline, not a doubled Validation model
Audit Gemini Pro Latest (shared w/ CC-B) Acceptable if Audit is post-hoc read-only; weakens "different eyes" narrative

Critical notes from this seat (DeepSeek as CC-C)

  1. Putting the Chinese-lab independent re-score on Tier C while OpenAI/Google cross-checks sit on Tier A sends a quality hierarchy signal that undercuts the independence thesis. If CC-C exists to catch what US labs miss, it cannot be the budget afterthought. Promote CC-C primary to Tier B minimum (or keep DeepSeek but stop labeling the role as C-tier capacity).
  2. Legal/Regulatory on Sonnet 5 shared with Validation is double-duty on the same model family and same vendor as Primary. For FDA / PIPL / data-export / EU AI Act work, this is the wrong specialization.
  3. MiniMax-M3 for Financial Integrity is a bet, not a proven assignment. Spec should require a calibration set (N synthetic cap tables) before locking.

Verdict: DEFEND overall roster direction. FIX Legal assignment and CC-C tier signaling before launch marketing claims "8-vendor frontier panel."


3. admin-ai Paths — 6/10 — DEFER

Observations

  • Paths are more operationally concrete than v1.0 (good).
  • gemini/gemini-pro-latest is an alias, not a pinned model. Alias drift silently changes Audit + CC-B behavior.
  • Spec claims "All model paths confirmed against admin-ai (live, August 2026)" but this Cross-Check did not re-probe live /v1/models. Treat as asserted, not re-verified.
  • Internal architecture notes elsewhere still flag Qwen3.8 Max / Kimi K3 as "not yet on admin-ai — substitute qwen3.7-plus / kimi-k2.6." v2.0 uses the substitutes. Fine for launch, but naming in marketing vs runtime must not diverge (Qwen3.8 vs 3.7; Kimi K3 vs K2.6).

Verdict: DEFER to an ops smoke-test: one live completion per path, pin versions, ban floating *-latest in Enterprise default lineup.


4. Vendor Diversity — 8/10 — DEFEND

On paper (Enterprise default 12-judge)

  • Anthropic 3 (25%), OpenAI 2, Google 2, + DeepSeek / Moonshot / Alibaba / MiniMax / xAI ×1 each.
  • 8 vendors. Hard rule tightened 40% → 33%. This is the single biggest structural win vs v1.0 (Anthropic 44%).

Under stress (Chinese-lab concern)

  • Diversity is a steady-state property. Under fallback (§5), Anthropic and Google density climbs fast because Primary/Team and Research/CC-B chains are same-vendor on hop 1.
  • Three Chinese vendors (DeepSeek, Alibaba, Moonshot) + MiniMax is excellent presence. But only Market-Reality is a Chinese model in a weight-bearing interpretive seat; CC-C is Tier C; Financial MiniMax is unproven; Reasoning Kimi does not score dimensions.
  • Risk: Western models still dominate scoring power even when vendor count looks global.

Verdict: DEFEND the 8-vendor design. Do not celebrate "no vendor >25%" without measuring scoring-seat share after fallback.


5. Fallback Chains — 4/10 — FIX ⚠️ FATAL-ish

Spec claim (§3.1): "Fallback chains always cross vendor boundaries — Enforced at the assignment table (§2.3)."

Audit of §2.3 (Primary → Fallback 1 vendor):

Role Primary vendor Fallback 1 vendor Cross-vendor?
Research xAI Google
Primary Anthropic Anthropic
Validation Anthropic OpenAI
Cross-Check A OpenAI OpenAI
Cross-Check B Google Google
Cross-Check C DeepSeek Google
Reasoning Moonshot DeepSeek
Execution OpenAI Anthropic
Market Alibaba xAI
Financial MiniMax DeepSeek
Team/Founder Anthropic Anthropic
Audit Google Anthropic

Result: 4 of 12 roles (33%) violate the hard constraint on the first hop.

Additional cascade issue:

  • Research Fallback 1 → Fallback 2 = Google → Google (same-vendor cascade). If Grok is down and Google is degraded, Research has no third-ecosystem escape.

Hidden correlation (Chinese-lab lens)

  • Same-vendor fallback is not just a rule bug — it recreates correlated failure and correlated judgment. OpenAI Sol→Terra preserves OpenAI RLHF priors. Anthropic Opus→Sonnet / Fable→Sonnet preserves Anthropic critique style. Google Pro→Flash preserves Google grounding stack.
  • §4.1 says "All fallbacks cross vendor boundaries (§2.3 guarantees this)" — this is a false guarantee. Implementers will trust the rule table; the assignment table contradicts it.
  • DeepSeek is overused as universal sink (appears in 7 of 12 Fallback 1/2 slots). That makes DeepSeek a correlation hub under multi-provider brownout, ironic given CC-C independence branding.

Required FIX

  1. Rewrite every same-vendor F1 to a different vendor before any lower-tier same-family model.
  2. Suggested repairs:
Role Primary F1 (fixed) F2
Primary Opus 5 (Anth) GPT-5.6 Sol (OpenAI) Sonnet 5 (Anth) only as F2
CC-A Sol (OpenAI) DeepSeek V4 Pro or Kimi Terra (OpenAI) as F2
CC-B Gemini Pro (Google) DeepSeek or Qwen Gemini Flash as F2
Team Fable 5 (Anth) GPT-5.6 Sol or Kimi Sonnet 5 as F2
Research F2 replace Gemini Flash with DeepSeek or Qwen (third ecosystem)
  1. Add automated test: assert primary.vendor != fallback1.vendor for every row; CI fails the spec if violated.
  2. Cap any single vendor's appearance in Fallback 1 columns (DeepSeek sink problem).

Verdict: FIX — blocking. Do not ship "always cross vendor" language while §2.3 falsifies it.


6. Content-Based Rules / Verticals — 4/10 — FIX ⚠️

What exists: Biotech, Hardware/IoT, Fintech, Climate/Energy, SaaS(default) + stage + complexity. Composition cap +50% is good.

What is missing (non-Western / global deal flow)

Missing vertical Why it matters Suggested effect
Superapp / Mini-program ecosystem WeChat / Alipay / LINE / Grab-style platform dependency is a first-class business model in CN/SEA, not "SaaS" Market +30%; Execution weights platform policy risk; Legal +PIPL/platform ToS
B2B marketplace / transaction platform Take-rate, cold-start liquidity, disintermediation — not SaaS net-retention logic Market + Financial weights; Execution on two-sided ops
Cross-border / trade / payments corridor FX, export controls, dual-regulation Legal +40%; Financial + FX/settlement realism
Government / SOE / public procurement RFP-adjacent but also guanxi, budget cycles, localization mandates Legal + Team weights; different buyer psychology
Consumer social / short-video / live commerce Not "SaaS"; growth loops and platform risk dominate Market + Team; Execution on content/ops
Industrial / manufacturing SaaS in CN Hardware+SaaS hybrid common; supply chain + data residency Execution + Legal (data export)
Crypto / Web3 / stablecoin (global+Asia) Spec mentions nothing; still a real proposal class Legal + Market specialized
Edtech / Healthtech consumer CN Heavy regulatory, different from US edtech/HIPAA framing Legal frameworks beyond FDA/GDPR

SaaS-default failure mode (§3.3)

  • Unmatched verticals → SaaS weighting + a polite post-review notice.
  • That means a WeChat mini-program commerce proposal, a Southeast Asian B2B marketplace, or a China-US cross-border data startup all get Silicon-Valley SaaS calibration (NRR, seat expansion, PLG) and a footer apology.
  • From a Chinese-lab lens this is not a minor omission — it is systematic miscategorization of a large fraction of non-US venture proposals.

Also missing framework coverage in Legal phase description

  • Lists FDA, SOC 2, GDPR, EU AI Act.
  • Omits: PIPL, CSL, DSL (China); PDPA (Singapore/SEA); PDPB India; data export / MLPS; content/ICP licensing where relevant.

Required FIX

  1. Expand vertical table with at least: Superapp/Mini-program, B2B Marketplace, Cross-border, Gov/SOE, Consumer Social/Live Commerce.
  2. Change unmatched behavior from silent SaaS-default to explicit low-confidence flag that reduces overall confidence score and boosts Market-Reality + Legal weights generically — not SaaS metrics.
  3. Legal framework map must include PIPL/CSL/DSL + major APAC privacy regimes, not only Euro-American.
  4. Add stacking example for Superapp + Seed + High complexity in the spec so implementers see non-SaaS composition.

Verdict: FIX — high priority for any claim of global-grade review.


7. Latency Budgets — 6/10 — DEFER

  • v1.0 inversion (Enterprise tightest) is fixed. Good.
  • Enterprise 45s target / 90s max for 12-judge pipeline is still aggressive unless Phases 4ac and later specialized judges run heavily parallel.
  • Spec never states parallelism topology (which phases block on which). Without a DAG, latency budgets are wishes.
  • Free 90/180 and Pro 60/120 are reasonable.

Verdict: DEFER pending an explicit phase DAG + p95 measurement on admin-ai. Do not market "45s Enterprise full pipeline" until measured.


8. Feedback Loop — 7/10 — DEFEND

Strengths vs v1.0

  • Cold-start gate (90 days) — correct.
  • min N=30 bias / N=50 accuracy — correct.
  • Immediate bias recusal (>2σ, 3 reviews) — good operational escape hatch.
  • Malformed-score auto-replace — practical.

Chinese-lab concerns

  • T+90/180/365 "public funding data" is implicitly US/EU venture outcomes (Crunchbase-shaped). Funding outcome ≠ business outcome in many CN/SEA contexts (profitability, strategic acquisition, gov design-win).
  • Correlation-to-funding as accuracy ground truth will systematically mis-train the feedback loop against non-Western success patterns.
  • Inter-judge agreement during cold-start favors majority (Western) priors — minority Chinese-lab signals may look like "bias" and get recused.

Mitigations to log (not all blocking)

  • Multi-outcome labels: funded / revenue milestone / strategic acq / shutdown — not funded-only.
  • Protect minority-vendor disagreement from automatic bias recusal until N is high per vertical including APAC.
  • Separate calibration sets for US-SaaS vs APAC-marketplace vs regulated-CN.

Verdict: DEFEND structure. Outcome ontology needs globalization before the data moat hardens Western bias.


9. Pricing — 7/10 — DEFEND

  • $249 / $799 / $1,499 with 16.7% annual is coherent premium vs GC AI $500/seat and authoring tools.
  • Free=1 review/mo is the right leash.
  • Pre-Review Coach on Pro (not Enterprise-only) is smart funnel design.
  • Financial/Team/Legal/Audit Enterprise-gated matches cost of 12-judge pipeline.

Gaps

  • No APAC regional pricing / PPP consideration (often required for SEA/CN SMB founders — even if Enterprise stays global USD).
  • No seat-based vs org-based clarity vs GC AI's seat anchor (the comp is seat; VT is org-flat — explain or the anchor confuses buyers).
  • COGS not in this spec (exists in architecture notes). Fine for this doc.

Verdict: DEFEND premium hold. Not the Chinese-lab primary fight.


10. Product Boundary + SaaS-Default Centrism — 5/10 — FIX

Product boundary (VT vs RFP Tank)

  • Clean and correct. Critique vs compliance is the right split. DEFEND that half.

SaaS-default assumption (the real Dim-10 issue under this lens)

  • §3.2: Vertical = SaaS (default).
  • §3.3: unmatched → SaaS-default calibration.
  • Feature language, competitive set (AutogenAI, Bidara, AutoRFP, Civio, GC AI), and stage weights all assume US/EU B2B SaaS venture narrative.
  • Global proposal mass includes: superapps, mini-programs, transaction marketplaces, OEM/industrial platforms, cross-border commerce, SOE-facing govtech. Forcing SaaS defaults is not neutral — it is a prior.

Independence irony

  • You seated Qwen on Market-Reality and DeepSeek on CC-C — then told unmatched verticals to pretend they are SaaS. The non-Western models are asked to score with Western category priors.

Required FIX

  1. Rename default from "SaaS (default)" → "Generic / Unclassified (low-confidence)" with no SaaS-specific metric emphasis.
  2. SaaS becomes an explicit detected vertical like Fintech — not the null hypothesis.
  3. Confidence banner already in §3.3 should lower the headline confidence band when unclassified, not only notify.
  4. Competitive landscape §6.4 should acknowledge non-US critique/authoring tools if claiming global premium (or explicitly scope "US/EU primary GTM").

Verdict: FIX the null-hypothesis vertical. Keep VT/RFP boundary as-is.


Special Focus Answers (Brief)

A. Are fallback chains truly vendor-independent?

No. Spec asserts yes; §2.3 falsifies on 4/12 first hops (Primary, CC-A, CC-B, Team). Research F1→F2 is Google→Google. DeepSeek is a correlation sink on brownout. Blocking FIX.

B. Missing non-Western verticals?

Yes, material. Superapps/mini-programs, B2B marketplaces, WeChat/Alipay ecosystems, cross-border, gov/SOE, live commerce, industrial hybrid — all collapse to SaaS default. Legal frameworks omit PIPL/CSL/DSL/PDPA.

C. Is "SaaS default" too US-centric?

Yes. It is the null hypothesis for the entire content-based system and shapes metric priors (seat expansion, NRR, PLG). Should be an explicit vertical, not the default. Unclassified → low-confidence generic, not SaaS.


FIX / DEFEND / DEFER Register

ID Item Priority Action
F1 Rewrite §2.3 so every Primary→F1 is cross-vendor; kill false §3.1/§4.1 guarantee P0 FIX
F2 Research F2 leave Google cascade; add third-ecosystem F2 P0 FIX
F3 Expand verticals: Superapp, B2B marketplace, Cross-border, Gov/SOE, Live commerce P0 FIX
F4 Replace SaaS-as-default with Generic/Unclassified low-confidence P0 FIX
F5 Legal frameworks: add PIPL/CSL/DSL/PDPA (+ data-export) P1 FIX
F6 Legal role: stop sharing Sonnet with Validation; dedicated assignment P1 FIX
F7 CC-C tier signaling: independence seat ≠ Tier C afterthought P1 FIX
F8 Cap DeepSeek as universal fallback sink P1 FIX
F9 Pin Gemini model; ban floating *-latest in Enterprise default P1 FIX
F10 Map 10 scoring dimensions ↔ 12 phases explicitly P2 FIX
D1 12-role expansion (Financial/Team/Legal) DEFEND
D2 8-vendor / 33% hard cap direction DEFEND
D3 Cold-start + N gates in feedback loop DEFEND
D4 Premium pricing $249/$799/$1499 DEFEND
D5 VT vs RFP Tank boundary DEFEND
D6 Qwen on Market-Reality + Kimi on Reasoning DEFEND
R1 Live admin-ai path smoke-test all 12 DEFER
R2 Latency DAG + p95 measure before marketing 45s DEFER
R3 Technical Architecture + Geo-Market roles (v4.2) DEFER
R4 Multi-outcome (non-US) feedback ontology DEFER

Per-Dimension Scores (machine-readable)

{
  "reviewer": "Cross-Check C — DeepSeek V4 Pro",
  "spec": "judge-pool-spec v2.0",
  "lens": "chinese-lab-independence",
  "overall": 6.4,
  "dimensions": [
    {"id": 1, "name": "role_coverage", "score": 7, "verdict": "DEFEND"},
    {"id": 2, "name": "model_assignments", "score": 7, "verdict": "DEFEND"},
    {"id": 3, "name": "admin_ai_paths", "score": 6, "verdict": "DEFER"},
    {"id": 4, "name": "vendor_diversity", "score": 8, "verdict": "DEFEND"},
    {"id": 5, "name": "fallback_chains", "score": 4, "verdict": "FIX"},
    {"id": 6, "name": "content_based_rules", "score": 4, "verdict": "FIX"},
    {"id": 7, "name": "latency_budgets", "score": 6, "verdict": "DEFER"},
    {"id": 8, "name": "feedback_loop", "score": 7, "verdict": "DEFEND"},
    {"id": 9, "name": "pricing", "score": 7, "verdict": "DEFEND"},
    {"id": 10, "name": "product_boundary_saas_default", "score": 5, "verdict": "FIX"}
  ],
  "blocking_fixes": ["F1", "F2", "F3", "F4"],
  "deployed_url_mismatch": "https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md still serves v1.0"
}

Closing (Cross-Check C voice)

v2.0 finally treats Chinese labs as first-class citizens of the panel. That is real progress.

But independence is not a logo count. It is what happens when the primary is down, when the vertical is not YC-SaaS, and when the feedback loop decides whose disagreement is "bias." On those three tests the spec still thinks in Silicon Valley defaults while wearing an eight-vendor badge.

Ship after F1F4. Everything else can follow.

— Cross-Check C (DeepSeek V4 Pro), independent re-score, 2026-08-12