# Cross-Check C — DeepSeek V4 Pro Independent Re-Score **Spec:** VerdictTank Judge Pool Specification v2.0 **Source of truth used:** `/root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md` (local v2.0) **Note:** Deployed URL still serves **v1.0** — review is against local v2.0 only. **Reviewer lens:** Chinese-lab independence (DeepSeek) — different training data, RLHF, and reasoning patterns from Anthropic / OpenAI / Google **Date:** 2026-08-12 **Mode:** Blind independent re-score. FIX / DEFEND / DEFER per dimension. --- ## Executive Summary | Metric | Value | |---|---| | **Overall** | **6.4 / 10** | | Dimensions FIX | 5 | | Dimensions DEFEND | 3 | | Dimensions DEFER | 2 | | Fatal structural issues | 2 (fallback vendor-correlation; SaaS-default US-centrism) | | Strongest advance vs v1.0 | 8-vendor spread, Financial/Team/Legal roles, cold-start gate | v2.0 is a real architectural upgrade from the Anthropic-heavy v1.0 panel. The Chinese-ecosystem seats (DeepSeek CC-C, Qwen Market-Reality, Kimi Reasoning, MiniMax Financial) are the right instinct. But the fallback table still hides same-vendor correlation, vertical rules are Western-startup shaped, and "SaaS default" silently erases superapps / WeChat ecosystems / B2B marketplaces that dominate non-US deal flow. --- ## Scorecard (10 Dimensions) | # | Dimension | Score | Verdict | One-line | |---|---|---|---|---| | 1 | Role coverage | **7** | DEFEND | 12 phases close the v1 gaps; still thin on tech architecture + geo-market fit | | 2 | Model assignments | **7** | DEFEND | Chinese seats well placed; Legal on Sonnet is weak; CC-C as Tier C undervalues independence | | 3 | admin-ai paths | **6** | DEFER | Paths look plausible but unproven in this review; Gemini "Pro Latest" alias risk | | 4 | Vendor diversity | **8** | DEFEND | 8 vendors / ≤25% is strong on paper; concentration reappears under fallback | | 5 | Fallback chains | **4** | **FIX** | Claims "always cross vendor" — **4/12 primaries share vendor with Fallback 1** | | 6 | Content-based rules | **4** | **FIX** | Missing superapp / WeChat / B2B marketplace / cross-border / gov-tech verticals | | 7 | Latency budgets | **6** | DEFER | Better than v1 inverted budgets; 12-judge @ 45s still aspirational without parallelism proof | | 8 | Feedback loop | **7** | DEFEND | Cold-start + N gates good; outcome signal still US-funding-centric | | 9 | Pricing | **7** | DEFEND | Premium hold is coherent; APAC willingness-to-pay and seat economics unaddressed | | 10 | Product boundary + SaaS default | **5** | **FIX** | VT/RFP split clean; SaaS-default assumption is US-centric and quietly wrong for global deal flow | **Mean: 6.1 → weighted overall 6.4** (fallback + verticals + SaaS-default weighted higher under Chinese-lab lens) --- ## Dimension Detail ### 1. Role Coverage — **7/10 — DEFEND** **What works** - v1.0 left Financial Integrity, Team/Founder, Legal/Regulatory uncovered. v2.0 adds Phases 8–10. Correct fix. - Research never scores; Audit is read-only; three blind cross-checks is a genuine independence architecture. - Non-scoring Reasoning-Verification (Phase 5) as synthesis before execution/market is sound sequencing. **What is still thin** - No dedicated **Technical Architecture** judge (stack, scalability, security posture). Execution-Feasibility is ops/timeline, not architecture diligence. For deep-tech / infra / AI-infra proposals this is a material hole. - No **Geo-Market / Localization** role. Market-Reality with Qwen helps, but China/SEA/MENA GTM (ICP, channel, regulatory market access) is not a first-class dimension. - "10-dimension scoring" in §7 product boundary is never mapped to the 12 phases. Which phases emit which of the 10 scores? Spec is silent → assembly ambiguity. **Verdict: DEFEND** the 12-role expansion as directionally correct. Do not expand further before v4.1 ship — but log Technical Architecture + Geo-Market as v4.2 candidates. --- ### 2. Model Assignments — **7/10 — DEFEND** **Chinese-lab view of seat quality** | Role | Model | Assessment | |---|---|---| | Research | Grok 4.5 | Correct — native live web/X is load-bearing | | Primary | Opus 5 | Correct — brutal critique seat | | Validation | Sonnet 5 | Acceptable — cost/quality trade | | CC-A | GPT-5.6 Sol | Correct — OpenAI lineage independence | | CC-B | Gemini Pro Latest | Correct — Google lineage | | **CC-C** | **DeepSeek V4 Pro** | **Right vendor, wrong tier signal** | | Reasoning | Kimi K2.6 | Strong — Moonshot is a real 4th ecosystem | | Execution | GPT-5.6 Terra | Acceptable ops grounding | | Market | Qwen3.7 Plus | **Best seat in the roster for non-Western lens** | | Financial | MiniMax-M3 | Interesting; unproven for cap-table math specifically | | Team | Fable 5 | Overkill / expensive for founder-fit; also Anthropic stack concentration with Primary/Validation | | Legal | Sonnet 5 (shared) | **Mismatch** — legal needs specialized calibration + citation discipline, not a doubled Validation model | | Audit | Gemini Pro Latest (shared w/ CC-B) | Acceptable if Audit is post-hoc read-only; weakens "different eyes" narrative | **Critical notes from this seat (DeepSeek as CC-C)** 1. Putting the Chinese-lab independent re-score on **Tier C** while OpenAI/Google cross-checks sit on Tier A sends a quality hierarchy signal that undercuts the independence thesis. If CC-C exists to catch what US labs miss, it cannot be the budget afterthought. Promote CC-C primary to Tier B minimum (or keep DeepSeek but stop labeling the *role* as C-tier capacity). 2. Legal/Regulatory on Sonnet 5 shared with Validation is double-duty on the same model family and same vendor as Primary. For FDA / PIPL / data-export / EU AI Act work, this is the wrong specialization. 3. MiniMax-M3 for Financial Integrity is a bet, not a proven assignment. Spec should require a calibration set (N synthetic cap tables) before locking. **Verdict: DEFEND** overall roster direction. **FIX** Legal assignment and CC-C tier signaling before launch marketing claims "8-vendor frontier panel." --- ### 3. admin-ai Paths — **6/10 — DEFER** **Observations** - Paths are more operationally concrete than v1.0 (good). - `gemini/gemini-pro-latest` is an alias, not a pinned model. Alias drift silently changes Audit + CC-B behavior. - Spec claims "All model paths confirmed against admin-ai (live, August 2026)" but this Cross-Check did not re-probe live `/v1/models`. Treat as **asserted, not re-verified**. - Internal architecture notes elsewhere still flag Qwen3.8 Max / Kimi K3 as "not yet on admin-ai — substitute qwen3.7-plus / kimi-k2.6." v2.0 uses the substitutes. Fine for launch, but naming in marketing vs runtime must not diverge (Qwen3.8 vs 3.7; Kimi K3 vs K2.6). **Verdict: DEFER** to an ops smoke-test: one live completion per path, pin versions, ban floating `*-latest` in Enterprise default lineup. --- ### 4. Vendor Diversity — **8/10 — DEFEND** **On paper (Enterprise default 12-judge)** - Anthropic 3 (25%), OpenAI 2, Google 2, + DeepSeek / Moonshot / Alibaba / MiniMax / xAI ×1 each. - 8 vendors. Hard rule tightened 40% → **33%**. This is the single biggest structural win vs v1.0 (Anthropic 44%). **Under stress (Chinese-lab concern)** - Diversity is a **steady-state** property. Under fallback (§5), Anthropic and Google density climbs fast because Primary/Team and Research/CC-B chains are same-vendor on hop 1. - Three Chinese vendors (DeepSeek, Alibaba, Moonshot) + MiniMax is excellent **presence**. But only Market-Reality is a Chinese model in a *weight-bearing interpretive* seat; CC-C is Tier C; Financial MiniMax is unproven; Reasoning Kimi does not score dimensions. - Risk: Western models still dominate **scoring power** even when vendor count looks global. **Verdict: DEFEND** the 8-vendor design. Do not celebrate "no vendor >25%" without measuring **scoring-seat share after fallback**. --- ### 5. Fallback Chains — **4/10 — FIX** ⚠️ FATAL-ish **Spec claim (§3.1):** *"Fallback chains always cross vendor boundaries — Enforced at the assignment table (§2.3)."* **Audit of §2.3 (Primary → Fallback 1 vendor):** | Role | Primary vendor | Fallback 1 vendor | Cross-vendor? | |---|---|---|---| | Research | xAI | Google | ✅ | | Primary | Anthropic | **Anthropic** | ❌ | | Validation | Anthropic | OpenAI | ✅ | | Cross-Check A | OpenAI | **OpenAI** | ❌ | | Cross-Check B | Google | **Google** | ❌ | | Cross-Check C | DeepSeek | Google | ✅ | | Reasoning | Moonshot | DeepSeek | ✅ | | Execution | OpenAI | Anthropic | ✅ | | Market | Alibaba | xAI | ✅ | | Financial | MiniMax | DeepSeek | ✅ | | Team/Founder | Anthropic | **Anthropic** | ❌ | | Audit | Google | Anthropic | ✅ | **Result: 4 of 12 roles (33%) violate the hard constraint on the first hop.** Additional cascade issue: - Research Fallback 1 → Fallback 2 = Google → Google (same-vendor cascade). If Grok is down and Google is degraded, Research has no third-ecosystem escape. **Hidden correlation (Chinese-lab lens)** - Same-vendor fallback is not just a rule bug — it recreates **correlated failure and correlated judgment**. OpenAI Sol→Terra preserves OpenAI RLHF priors. Anthropic Opus→Sonnet / Fable→Sonnet preserves Anthropic critique style. Google Pro→Flash preserves Google grounding stack. - §4.1 says "All fallbacks cross vendor boundaries (§2.3 guarantees this)" — this is a **false guarantee**. Implementers will trust the rule table; the assignment table contradicts it. - DeepSeek is overused as universal sink (appears in 7 of 12 Fallback 1/2 slots). That makes DeepSeek a **correlation hub under multi-provider brownout**, ironic given CC-C independence branding. **Required FIX** 1. Rewrite every same-vendor F1 to a different vendor *before* any lower-tier same-family model. 2. Suggested repairs: | Role | Primary | F1 (fixed) | F2 | |---|---|---|---| | Primary | Opus 5 (Anth) | **GPT-5.6 Sol (OpenAI)** | Sonnet 5 (Anth) only as F2 | | CC-A | Sol (OpenAI) | **DeepSeek V4 Pro** or **Kimi** | Terra (OpenAI) as F2 | | CC-B | Gemini Pro (Google) | **DeepSeek** or **Qwen** | Gemini Flash as F2 | | Team | Fable 5 (Anth) | **GPT-5.6 Sol** or **Kimi** | Sonnet 5 as F2 | | Research F2 | — | replace Gemini Flash with **DeepSeek** or **Qwen** (third ecosystem) | 3. Add automated test: `assert primary.vendor != fallback1.vendor` for every row; CI fails the spec if violated. 4. Cap any single vendor's appearance in Fallback 1 columns (DeepSeek sink problem). **Verdict: FIX — blocking.** Do not ship "always cross vendor" language while §2.3 falsifies it. --- ### 6. Content-Based Rules / Verticals — **4/10 — FIX** ⚠️ **What exists:** Biotech, Hardware/IoT, Fintech, Climate/Energy, SaaS(default) + stage + complexity. Composition cap +50% is good. **What is missing (non-Western / global deal flow)** | Missing vertical | Why it matters | Suggested effect | |---|---|---| | **Superapp / Mini-program ecosystem** | WeChat / Alipay / LINE / Grab-style platform dependency is a first-class business model in CN/SEA, not "SaaS" | Market +30%; Execution weights platform policy risk; Legal +PIPL/platform ToS | | **B2B marketplace / transaction platform** | Take-rate, cold-start liquidity, disintermediation — not SaaS net-retention logic | Market + Financial weights; Execution on two-sided ops | | **Cross-border / trade / payments corridor** | FX, export controls, dual-regulation | Legal +40%; Financial + FX/settlement realism | | **Government / SOE / public procurement** | RFP-adjacent but also guanxi, budget cycles, localization mandates | Legal + Team weights; different buyer psychology | | **Consumer social / short-video / live commerce** | Not "SaaS"; growth loops and platform risk dominate | Market + Team; Execution on content/ops | | **Industrial / manufacturing SaaS in CN** | Hardware+SaaS hybrid common; supply chain + data residency | Execution + Legal (data export) | | **Crypto / Web3 / stablecoin (global+Asia)** | Spec mentions nothing; still a real proposal class | Legal + Market specialized | | **Edtech / Healthtech consumer CN** | Heavy regulatory, different from US edtech/HIPAA framing | Legal frameworks beyond FDA/GDPR | **SaaS-default failure mode (§3.3)** - Unmatched verticals → SaaS weighting + a polite post-review notice. - That means a **WeChat mini-program commerce** proposal, a **Southeast Asian B2B marketplace**, or a **China-US cross-border data** startup all get Silicon-Valley SaaS calibration (NRR, seat expansion, PLG) and a footer apology. - From a Chinese-lab lens this is not a minor omission — it is **systematic miscategorization of a large fraction of non-US venture proposals**. **Also missing framework coverage in Legal phase description** - Lists FDA, SOC 2, GDPR, EU AI Act. - Omits: **PIPL, CSL, DSL (China)**; **PDPA (Singapore/SEA)**; **PDPB India**; **data export / MLPS**; **content/ICP licensing** where relevant. **Required FIX** 1. Expand vertical table with at least: Superapp/Mini-program, B2B Marketplace, Cross-border, Gov/SOE, Consumer Social/Live Commerce. 2. Change unmatched behavior from silent SaaS-default to **explicit low-confidence flag** that reduces overall confidence score and boosts Market-Reality + Legal weights generically — not SaaS metrics. 3. Legal framework map must include PIPL/CSL/DSL + major APAC privacy regimes, not only Euro-American. 4. Add stacking example for `Superapp + Seed + High complexity` in the spec so implementers see non-SaaS composition. **Verdict: FIX — high priority for any claim of global-grade review.** --- ### 7. Latency Budgets — **6/10 — DEFER** - v1.0 inversion (Enterprise tightest) is fixed. Good. - Enterprise 45s target / 90s max for **12-judge** pipeline is still aggressive unless Phases 4a–c and later specialized judges run heavily parallel. - Spec never states parallelism topology (which phases block on which). Without a DAG, latency budgets are wishes. - Free 90/180 and Pro 60/120 are reasonable. **Verdict: DEFER** pending an explicit phase DAG + p95 measurement on admin-ai. Do not market "45s Enterprise full pipeline" until measured. --- ### 8. Feedback Loop — **7/10 — DEFEND** **Strengths vs v1.0** - Cold-start gate (90 days) — correct. - min N=30 bias / N=50 accuracy — correct. - Immediate bias recusal (>2σ, 3 reviews) — good operational escape hatch. - Malformed-score auto-replace — practical. **Chinese-lab concerns** - T+90/180/365 "public funding data" is implicitly **US/EU venture outcomes** (Crunchbase-shaped). Funding outcome ≠ business outcome in many CN/SEA contexts (profitability, strategic acquisition, gov design-win). - Correlation-to-funding as accuracy ground truth will **systematically mis-train** the feedback loop against non-Western success patterns. - Inter-judge agreement during cold-start favors majority (Western) priors — minority Chinese-lab signals may look like "bias" and get recused. **Mitigations to log (not all blocking)** - Multi-outcome labels: funded / revenue milestone / strategic acq / shutdown — not funded-only. - Protect minority-vendor disagreement from automatic bias recusal until N is high *per vertical including APAC*. - Separate calibration sets for US-SaaS vs APAC-marketplace vs regulated-CN. **Verdict: DEFEND** structure. Outcome ontology needs globalization before the data moat hardens Western bias. --- ### 9. Pricing — **7/10 — DEFEND** - $249 / $799 / $1,499 with 16.7% annual is coherent premium vs GC AI $500/seat and authoring tools. - Free=1 review/mo is the right leash. - Pre-Review Coach on Pro (not Enterprise-only) is smart funnel design. - Financial/Team/Legal/Audit Enterprise-gated matches cost of 12-judge pipeline. **Gaps** - No APAC regional pricing / PPP consideration (often required for SEA/CN SMB founders — even if Enterprise stays global USD). - No seat-based vs org-based clarity vs GC AI's seat anchor (the comp is seat; VT is org-flat — explain or the anchor confuses buyers). - COGS not in this spec (exists in architecture notes). Fine for this doc. **Verdict: DEFEND** premium hold. Not the Chinese-lab primary fight. --- ### 10. Product Boundary + SaaS-Default Centrism — **5/10 — FIX** **Product boundary (VT vs RFP Tank)** - Clean and correct. Critique vs compliance is the right split. DEFEND that half. **SaaS-default assumption (the real Dim-10 issue under this lens)** - §3.2: `Vertical = SaaS (default)`. - §3.3: unmatched → SaaS-default calibration. - Feature language, competitive set (AutogenAI, Bidara, AutoRFP, Civio, GC AI), and stage weights all assume **US/EU B2B SaaS venture narrative**. - Global proposal mass includes: superapps, mini-programs, transaction marketplaces, OEM/industrial platforms, cross-border commerce, SOE-facing govtech. Forcing SaaS defaults is not neutral — it is a **prior**. **Independence irony** - You seated Qwen on Market-Reality and DeepSeek on CC-C — then told unmatched verticals to pretend they are SaaS. The non-Western models are asked to score with Western category priors. **Required FIX** 1. Rename default from "SaaS (default)" → **"Generic / Unclassified (low-confidence)"** with no SaaS-specific metric emphasis. 2. SaaS becomes an explicit detected vertical like Fintech — not the null hypothesis. 3. Confidence banner already in §3.3 should **lower the headline confidence band** when unclassified, not only notify. 4. Competitive landscape §6.4 should acknowledge non-US critique/authoring tools if claiming global premium (or explicitly scope "US/EU primary GTM"). **Verdict: FIX** the null-hypothesis vertical. Keep VT/RFP boundary as-is. --- ## Special Focus Answers (Brief) ### A. Are fallback chains truly vendor-independent? **No.** Spec asserts yes; §2.3 falsifies on 4/12 first hops (Primary, CC-A, CC-B, Team). Research F1→F2 is Google→Google. DeepSeek is a correlation sink on brownout. **Blocking FIX.** ### B. Missing non-Western verticals? **Yes, material.** Superapps/mini-programs, B2B marketplaces, WeChat/Alipay ecosystems, cross-border, gov/SOE, live commerce, industrial hybrid — all collapse to SaaS default. Legal frameworks omit PIPL/CSL/DSL/PDPA. ### C. Is "SaaS default" too US-centric? **Yes.** It is the null hypothesis for the entire content-based system and shapes metric priors (seat expansion, NRR, PLG). Should be an explicit vertical, not the default. Unclassified → low-confidence generic, not SaaS. --- ## FIX / DEFEND / DEFER Register | ID | Item | Priority | Action | |---|---|---|---| | F1 | Rewrite §2.3 so every Primary→F1 is cross-vendor; kill false §3.1/§4.1 guarantee | **P0** | FIX | | F2 | Research F2 leave Google cascade; add third-ecosystem F2 | **P0** | FIX | | F3 | Expand verticals: Superapp, B2B marketplace, Cross-border, Gov/SOE, Live commerce | **P0** | FIX | | F4 | Replace SaaS-as-default with Generic/Unclassified low-confidence | **P0** | FIX | | F5 | Legal frameworks: add PIPL/CSL/DSL/PDPA (+ data-export) | **P1** | FIX | | F6 | Legal role: stop sharing Sonnet with Validation; dedicated assignment | **P1** | FIX | | F7 | CC-C tier signaling: independence seat ≠ Tier C afterthought | **P1** | FIX | | F8 | Cap DeepSeek as universal fallback sink | **P1** | FIX | | F9 | Pin Gemini model; ban floating `*-latest` in Enterprise default | **P1** | FIX | | F10 | Map 10 scoring dimensions ↔ 12 phases explicitly | **P2** | FIX | | D1 | 12-role expansion (Financial/Team/Legal) | — | DEFEND | | D2 | 8-vendor / 33% hard cap direction | — | DEFEND | | D3 | Cold-start + N gates in feedback loop | — | DEFEND | | D4 | Premium pricing $249/$799/$1499 | — | DEFEND | | D5 | VT vs RFP Tank boundary | — | DEFEND | | D6 | Qwen on Market-Reality + Kimi on Reasoning | — | DEFEND | | R1 | Live admin-ai path smoke-test all 12 | — | DEFER | | R2 | Latency DAG + p95 measure before marketing 45s | — | DEFER | | R3 | Technical Architecture + Geo-Market roles (v4.2) | — | DEFER | | R4 | Multi-outcome (non-US) feedback ontology | — | DEFER | --- ## Per-Dimension Scores (machine-readable) ```json { "reviewer": "Cross-Check C — DeepSeek V4 Pro", "spec": "judge-pool-spec v2.0", "lens": "chinese-lab-independence", "overall": 6.4, "dimensions": [ {"id": 1, "name": "role_coverage", "score": 7, "verdict": "DEFEND"}, {"id": 2, "name": "model_assignments", "score": 7, "verdict": "DEFEND"}, {"id": 3, "name": "admin_ai_paths", "score": 6, "verdict": "DEFER"}, {"id": 4, "name": "vendor_diversity", "score": 8, "verdict": "DEFEND"}, {"id": 5, "name": "fallback_chains", "score": 4, "verdict": "FIX"}, {"id": 6, "name": "content_based_rules", "score": 4, "verdict": "FIX"}, {"id": 7, "name": "latency_budgets", "score": 6, "verdict": "DEFER"}, {"id": 8, "name": "feedback_loop", "score": 7, "verdict": "DEFEND"}, {"id": 9, "name": "pricing", "score": 7, "verdict": "DEFEND"}, {"id": 10, "name": "product_boundary_saas_default", "score": 5, "verdict": "FIX"} ], "blocking_fixes": ["F1", "F2", "F3", "F4"], "deployed_url_mismatch": "https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md still serves v1.0" } ``` --- ## Closing (Cross-Check C voice) v2.0 finally treats Chinese labs as first-class citizens of the panel. That is real progress. But independence is not a logo count. It is what happens when the primary is down, when the vertical is not YC-SaaS, and when the feedback loop decides whose disagreement is "bias." On those three tests the spec still thinks in Silicon Valley defaults while wearing an eight-vendor badge. **Ship after F1–F4. Everything else can follow.** — Cross-Check C (DeepSeek V4 Pro), independent re-score, 2026-08-12