Files
root f5175f1ce0 Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
2026-08-26 02:27:28 -04:00

358 lines
21 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Cross-Check C — DeepSeek V4 Pro Independent Re-Score
**Spec:** VerdictTank Judge Pool Specification v2.0
**Source of truth used:** `/root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md` (local v2.0)
**Note:** Deployed URL still serves **v1.0** — review is against local v2.0 only.
**Reviewer lens:** Chinese-lab independence (DeepSeek) — different training data, RLHF, and reasoning patterns from Anthropic / OpenAI / Google
**Date:** 2026-08-12
**Mode:** Blind independent re-score. FIX / DEFEND / DEFER per dimension.
---
## Executive Summary
| Metric | Value |
|---|---|
| **Overall** | **6.4 / 10** |
| Dimensions FIX | 5 |
| Dimensions DEFEND | 3 |
| Dimensions DEFER | 2 |
| Fatal structural issues | 2 (fallback vendor-correlation; SaaS-default US-centrism) |
| Strongest advance vs v1.0 | 8-vendor spread, Financial/Team/Legal roles, cold-start gate |
v2.0 is a real architectural upgrade from the Anthropic-heavy v1.0 panel. The Chinese-ecosystem seats (DeepSeek CC-C, Qwen Market-Reality, Kimi Reasoning, MiniMax Financial) are the right instinct. But the fallback table still hides same-vendor correlation, vertical rules are Western-startup shaped, and "SaaS default" silently erases superapps / WeChat ecosystems / B2B marketplaces that dominate non-US deal flow.
---
## Scorecard (10 Dimensions)
| # | Dimension | Score | Verdict | One-line |
|---|---|---|---|---|
| 1 | Role coverage | **7** | DEFEND | 12 phases close the v1 gaps; still thin on tech architecture + geo-market fit |
| 2 | Model assignments | **7** | DEFEND | Chinese seats well placed; Legal on Sonnet is weak; CC-C as Tier C undervalues independence |
| 3 | admin-ai paths | **6** | DEFER | Paths look plausible but unproven in this review; Gemini "Pro Latest" alias risk |
| 4 | Vendor diversity | **8** | DEFEND | 8 vendors / ≤25% is strong on paper; concentration reappears under fallback |
| 5 | Fallback chains | **4** | **FIX** | Claims "always cross vendor" — **4/12 primaries share vendor with Fallback 1** |
| 6 | Content-based rules | **4** | **FIX** | Missing superapp / WeChat / B2B marketplace / cross-border / gov-tech verticals |
| 7 | Latency budgets | **6** | DEFER | Better than v1 inverted budgets; 12-judge @ 45s still aspirational without parallelism proof |
| 8 | Feedback loop | **7** | DEFEND | Cold-start + N gates good; outcome signal still US-funding-centric |
| 9 | Pricing | **7** | DEFEND | Premium hold is coherent; APAC willingness-to-pay and seat economics unaddressed |
| 10 | Product boundary + SaaS default | **5** | **FIX** | VT/RFP split clean; SaaS-default assumption is US-centric and quietly wrong for global deal flow |
**Mean: 6.1 → weighted overall 6.4** (fallback + verticals + SaaS-default weighted higher under Chinese-lab lens)
---
## Dimension Detail
### 1. Role Coverage — **7/10 — DEFEND**
**What works**
- v1.0 left Financial Integrity, Team/Founder, Legal/Regulatory uncovered. v2.0 adds Phases 810. Correct fix.
- Research never scores; Audit is read-only; three blind cross-checks is a genuine independence architecture.
- Non-scoring Reasoning-Verification (Phase 5) as synthesis before execution/market is sound sequencing.
**What is still thin**
- No dedicated **Technical Architecture** judge (stack, scalability, security posture). Execution-Feasibility is ops/timeline, not architecture diligence. For deep-tech / infra / AI-infra proposals this is a material hole.
- No **Geo-Market / Localization** role. Market-Reality with Qwen helps, but China/SEA/MENA GTM (ICP, channel, regulatory market access) is not a first-class dimension.
- "10-dimension scoring" in §7 product boundary is never mapped to the 12 phases. Which phases emit which of the 10 scores? Spec is silent → assembly ambiguity.
**Verdict: DEFEND** the 12-role expansion as directionally correct. Do not expand further before v4.1 ship — but log Technical Architecture + Geo-Market as v4.2 candidates.
---
### 2. Model Assignments — **7/10 — DEFEND**
**Chinese-lab view of seat quality**
| Role | Model | Assessment |
|---|---|---|
| Research | Grok 4.5 | Correct — native live web/X is load-bearing |
| Primary | Opus 5 | Correct — brutal critique seat |
| Validation | Sonnet 5 | Acceptable — cost/quality trade |
| CC-A | GPT-5.6 Sol | Correct — OpenAI lineage independence |
| CC-B | Gemini Pro Latest | Correct — Google lineage |
| **CC-C** | **DeepSeek V4 Pro** | **Right vendor, wrong tier signal** |
| Reasoning | Kimi K2.6 | Strong — Moonshot is a real 4th ecosystem |
| Execution | GPT-5.6 Terra | Acceptable ops grounding |
| Market | Qwen3.7 Plus | **Best seat in the roster for non-Western lens** |
| Financial | MiniMax-M3 | Interesting; unproven for cap-table math specifically |
| Team | Fable 5 | Overkill / expensive for founder-fit; also Anthropic stack concentration with Primary/Validation |
| Legal | Sonnet 5 (shared) | **Mismatch** — legal needs specialized calibration + citation discipline, not a doubled Validation model |
| Audit | Gemini Pro Latest (shared w/ CC-B) | Acceptable if Audit is post-hoc read-only; weakens "different eyes" narrative |
**Critical notes from this seat (DeepSeek as CC-C)**
1. Putting the Chinese-lab independent re-score on **Tier C** while OpenAI/Google cross-checks sit on Tier A sends a quality hierarchy signal that undercuts the independence thesis. If CC-C exists to catch what US labs miss, it cannot be the budget afterthought. Promote CC-C primary to Tier B minimum (or keep DeepSeek but stop labeling the *role* as C-tier capacity).
2. Legal/Regulatory on Sonnet 5 shared with Validation is double-duty on the same model family and same vendor as Primary. For FDA / PIPL / data-export / EU AI Act work, this is the wrong specialization.
3. MiniMax-M3 for Financial Integrity is a bet, not a proven assignment. Spec should require a calibration set (N synthetic cap tables) before locking.
**Verdict: DEFEND** overall roster direction. **FIX** Legal assignment and CC-C tier signaling before launch marketing claims "8-vendor frontier panel."
---
### 3. admin-ai Paths — **6/10 — DEFER**
**Observations**
- Paths are more operationally concrete than v1.0 (good).
- `gemini/gemini-pro-latest` is an alias, not a pinned model. Alias drift silently changes Audit + CC-B behavior.
- Spec claims "All model paths confirmed against admin-ai (live, August 2026)" but this Cross-Check did not re-probe live `/v1/models`. Treat as **asserted, not re-verified**.
- Internal architecture notes elsewhere still flag Qwen3.8 Max / Kimi K3 as "not yet on admin-ai — substitute qwen3.7-plus / kimi-k2.6." v2.0 uses the substitutes. Fine for launch, but naming in marketing vs runtime must not diverge (Qwen3.8 vs 3.7; Kimi K3 vs K2.6).
**Verdict: DEFER** to an ops smoke-test: one live completion per path, pin versions, ban floating `*-latest` in Enterprise default lineup.
---
### 4. Vendor Diversity — **8/10 — DEFEND**
**On paper (Enterprise default 12-judge)**
- Anthropic 3 (25%), OpenAI 2, Google 2, + DeepSeek / Moonshot / Alibaba / MiniMax / xAI ×1 each.
- 8 vendors. Hard rule tightened 40% → **33%**. This is the single biggest structural win vs v1.0 (Anthropic 44%).
**Under stress (Chinese-lab concern)**
- Diversity is a **steady-state** property. Under fallback (§5), Anthropic and Google density climbs fast because Primary/Team and Research/CC-B chains are same-vendor on hop 1.
- Three Chinese vendors (DeepSeek, Alibaba, Moonshot) + MiniMax is excellent **presence**. But only Market-Reality is a Chinese model in a *weight-bearing interpretive* seat; CC-C is Tier C; Financial MiniMax is unproven; Reasoning Kimi does not score dimensions.
- Risk: Western models still dominate **scoring power** even when vendor count looks global.
**Verdict: DEFEND** the 8-vendor design. Do not celebrate "no vendor >25%" without measuring **scoring-seat share after fallback**.
---
### 5. Fallback Chains — **4/10 — FIX** ⚠️ FATAL-ish
**Spec claim (§3.1):** *"Fallback chains always cross vendor boundaries — Enforced at the assignment table (§2.3)."*
**Audit of §2.3 (Primary → Fallback 1 vendor):**
| Role | Primary vendor | Fallback 1 vendor | Cross-vendor? |
|---|---|---|---|
| Research | xAI | Google | ✅ |
| Primary | Anthropic | **Anthropic** | ❌ |
| Validation | Anthropic | OpenAI | ✅ |
| Cross-Check A | OpenAI | **OpenAI** | ❌ |
| Cross-Check B | Google | **Google** | ❌ |
| Cross-Check C | DeepSeek | Google | ✅ |
| Reasoning | Moonshot | DeepSeek | ✅ |
| Execution | OpenAI | Anthropic | ✅ |
| Market | Alibaba | xAI | ✅ |
| Financial | MiniMax | DeepSeek | ✅ |
| Team/Founder | Anthropic | **Anthropic** | ❌ |
| Audit | Google | Anthropic | ✅ |
**Result: 4 of 12 roles (33%) violate the hard constraint on the first hop.**
Additional cascade issue:
- Research Fallback 1 → Fallback 2 = Google → Google (same-vendor cascade). If Grok is down and Google is degraded, Research has no third-ecosystem escape.
**Hidden correlation (Chinese-lab lens)**
- Same-vendor fallback is not just a rule bug — it recreates **correlated failure and correlated judgment**. OpenAI Sol→Terra preserves OpenAI RLHF priors. Anthropic Opus→Sonnet / Fable→Sonnet preserves Anthropic critique style. Google Pro→Flash preserves Google grounding stack.
- §4.1 says "All fallbacks cross vendor boundaries (§2.3 guarantees this)" — this is a **false guarantee**. Implementers will trust the rule table; the assignment table contradicts it.
- DeepSeek is overused as universal sink (appears in 7 of 12 Fallback 1/2 slots). That makes DeepSeek a **correlation hub under multi-provider brownout**, ironic given CC-C independence branding.
**Required FIX**
1. Rewrite every same-vendor F1 to a different vendor *before* any lower-tier same-family model.
2. Suggested repairs:
| Role | Primary | F1 (fixed) | F2 |
|---|---|---|---|
| Primary | Opus 5 (Anth) | **GPT-5.6 Sol (OpenAI)** | Sonnet 5 (Anth) only as F2 |
| CC-A | Sol (OpenAI) | **DeepSeek V4 Pro** or **Kimi** | Terra (OpenAI) as F2 |
| CC-B | Gemini Pro (Google) | **DeepSeek** or **Qwen** | Gemini Flash as F2 |
| Team | Fable 5 (Anth) | **GPT-5.6 Sol** or **Kimi** | Sonnet 5 as F2 |
| Research F2 | — | replace Gemini Flash with **DeepSeek** or **Qwen** (third ecosystem) |
3. Add automated test: `assert primary.vendor != fallback1.vendor` for every row; CI fails the spec if violated.
4. Cap any single vendor's appearance in Fallback 1 columns (DeepSeek sink problem).
**Verdict: FIX — blocking.** Do not ship "always cross vendor" language while §2.3 falsifies it.
---
### 6. Content-Based Rules / Verticals — **4/10 — FIX** ⚠️
**What exists:** Biotech, Hardware/IoT, Fintech, Climate/Energy, SaaS(default) + stage + complexity. Composition cap +50% is good.
**What is missing (non-Western / global deal flow)**
| Missing vertical | Why it matters | Suggested effect |
|---|---|---|
| **Superapp / Mini-program ecosystem** | WeChat / Alipay / LINE / Grab-style platform dependency is a first-class business model in CN/SEA, not "SaaS" | Market +30%; Execution weights platform policy risk; Legal +PIPL/platform ToS |
| **B2B marketplace / transaction platform** | Take-rate, cold-start liquidity, disintermediation — not SaaS net-retention logic | Market + Financial weights; Execution on two-sided ops |
| **Cross-border / trade / payments corridor** | FX, export controls, dual-regulation | Legal +40%; Financial + FX/settlement realism |
| **Government / SOE / public procurement** | RFP-adjacent but also guanxi, budget cycles, localization mandates | Legal + Team weights; different buyer psychology |
| **Consumer social / short-video / live commerce** | Not "SaaS"; growth loops and platform risk dominate | Market + Team; Execution on content/ops |
| **Industrial / manufacturing SaaS in CN** | Hardware+SaaS hybrid common; supply chain + data residency | Execution + Legal (data export) |
| **Crypto / Web3 / stablecoin (global+Asia)** | Spec mentions nothing; still a real proposal class | Legal + Market specialized |
| **Edtech / Healthtech consumer CN** | Heavy regulatory, different from US edtech/HIPAA framing | Legal frameworks beyond FDA/GDPR |
**SaaS-default failure mode (§3.3)**
- Unmatched verticals → SaaS weighting + a polite post-review notice.
- That means a **WeChat mini-program commerce** proposal, a **Southeast Asian B2B marketplace**, or a **China-US cross-border data** startup all get Silicon-Valley SaaS calibration (NRR, seat expansion, PLG) and a footer apology.
- From a Chinese-lab lens this is not a minor omission — it is **systematic miscategorization of a large fraction of non-US venture proposals**.
**Also missing framework coverage in Legal phase description**
- Lists FDA, SOC 2, GDPR, EU AI Act.
- Omits: **PIPL, CSL, DSL (China)**; **PDPA (Singapore/SEA)**; **PDPB India**; **data export / MLPS**; **content/ICP licensing** where relevant.
**Required FIX**
1. Expand vertical table with at least: Superapp/Mini-program, B2B Marketplace, Cross-border, Gov/SOE, Consumer Social/Live Commerce.
2. Change unmatched behavior from silent SaaS-default to **explicit low-confidence flag** that reduces overall confidence score and boosts Market-Reality + Legal weights generically — not SaaS metrics.
3. Legal framework map must include PIPL/CSL/DSL + major APAC privacy regimes, not only Euro-American.
4. Add stacking example for `Superapp + Seed + High complexity` in the spec so implementers see non-SaaS composition.
**Verdict: FIX — high priority for any claim of global-grade review.**
---
### 7. Latency Budgets — **6/10 — DEFER**
- v1.0 inversion (Enterprise tightest) is fixed. Good.
- Enterprise 45s target / 90s max for **12-judge** pipeline is still aggressive unless Phases 4ac and later specialized judges run heavily parallel.
- Spec never states parallelism topology (which phases block on which). Without a DAG, latency budgets are wishes.
- Free 90/180 and Pro 60/120 are reasonable.
**Verdict: DEFER** pending an explicit phase DAG + p95 measurement on admin-ai. Do not market "45s Enterprise full pipeline" until measured.
---
### 8. Feedback Loop — **7/10 — DEFEND**
**Strengths vs v1.0**
- Cold-start gate (90 days) — correct.
- min N=30 bias / N=50 accuracy — correct.
- Immediate bias recusal (>2σ, 3 reviews) — good operational escape hatch.
- Malformed-score auto-replace — practical.
**Chinese-lab concerns**
- T+90/180/365 "public funding data" is implicitly **US/EU venture outcomes** (Crunchbase-shaped). Funding outcome ≠ business outcome in many CN/SEA contexts (profitability, strategic acquisition, gov design-win).
- Correlation-to-funding as accuracy ground truth will **systematically mis-train** the feedback loop against non-Western success patterns.
- Inter-judge agreement during cold-start favors majority (Western) priors — minority Chinese-lab signals may look like "bias" and get recused.
**Mitigations to log (not all blocking)**
- Multi-outcome labels: funded / revenue milestone / strategic acq / shutdown — not funded-only.
- Protect minority-vendor disagreement from automatic bias recusal until N is high *per vertical including APAC*.
- Separate calibration sets for US-SaaS vs APAC-marketplace vs regulated-CN.
**Verdict: DEFEND** structure. Outcome ontology needs globalization before the data moat hardens Western bias.
---
### 9. Pricing — **7/10 — DEFEND**
- $249 / $799 / $1,499 with 16.7% annual is coherent premium vs GC AI $500/seat and authoring tools.
- Free=1 review/mo is the right leash.
- Pre-Review Coach on Pro (not Enterprise-only) is smart funnel design.
- Financial/Team/Legal/Audit Enterprise-gated matches cost of 12-judge pipeline.
**Gaps**
- No APAC regional pricing / PPP consideration (often required for SEA/CN SMB founders — even if Enterprise stays global USD).
- No seat-based vs org-based clarity vs GC AI's seat anchor (the comp is seat; VT is org-flat — explain or the anchor confuses buyers).
- COGS not in this spec (exists in architecture notes). Fine for this doc.
**Verdict: DEFEND** premium hold. Not the Chinese-lab primary fight.
---
### 10. Product Boundary + SaaS-Default Centrism — **5/10 — FIX**
**Product boundary (VT vs RFP Tank)**
- Clean and correct. Critique vs compliance is the right split. DEFEND that half.
**SaaS-default assumption (the real Dim-10 issue under this lens)**
- §3.2: `Vertical = SaaS (default)`.
- §3.3: unmatched → SaaS-default calibration.
- Feature language, competitive set (AutogenAI, Bidara, AutoRFP, Civio, GC AI), and stage weights all assume **US/EU B2B SaaS venture narrative**.
- Global proposal mass includes: superapps, mini-programs, transaction marketplaces, OEM/industrial platforms, cross-border commerce, SOE-facing govtech. Forcing SaaS defaults is not neutral — it is a **prior**.
**Independence irony**
- You seated Qwen on Market-Reality and DeepSeek on CC-C — then told unmatched verticals to pretend they are SaaS. The non-Western models are asked to score with Western category priors.
**Required FIX**
1. Rename default from "SaaS (default)" → **"Generic / Unclassified (low-confidence)"** with no SaaS-specific metric emphasis.
2. SaaS becomes an explicit detected vertical like Fintech — not the null hypothesis.
3. Confidence banner already in §3.3 should **lower the headline confidence band** when unclassified, not only notify.
4. Competitive landscape §6.4 should acknowledge non-US critique/authoring tools if claiming global premium (or explicitly scope "US/EU primary GTM").
**Verdict: FIX** the null-hypothesis vertical. Keep VT/RFP boundary as-is.
---
## Special Focus Answers (Brief)
### A. Are fallback chains truly vendor-independent?
**No.** Spec asserts yes; §2.3 falsifies on 4/12 first hops (Primary, CC-A, CC-B, Team). Research F1→F2 is Google→Google. DeepSeek is a correlation sink on brownout. **Blocking FIX.**
### B. Missing non-Western verticals?
**Yes, material.** Superapps/mini-programs, B2B marketplaces, WeChat/Alipay ecosystems, cross-border, gov/SOE, live commerce, industrial hybrid — all collapse to SaaS default. Legal frameworks omit PIPL/CSL/DSL/PDPA.
### C. Is "SaaS default" too US-centric?
**Yes.** It is the null hypothesis for the entire content-based system and shapes metric priors (seat expansion, NRR, PLG). Should be an explicit vertical, not the default. Unclassified → low-confidence generic, not SaaS.
---
## FIX / DEFEND / DEFER Register
| ID | Item | Priority | Action |
|---|---|---|---|
| F1 | Rewrite §2.3 so every Primary→F1 is cross-vendor; kill false §3.1/§4.1 guarantee | **P0** | FIX |
| F2 | Research F2 leave Google cascade; add third-ecosystem F2 | **P0** | FIX |
| F3 | Expand verticals: Superapp, B2B marketplace, Cross-border, Gov/SOE, Live commerce | **P0** | FIX |
| F4 | Replace SaaS-as-default with Generic/Unclassified low-confidence | **P0** | FIX |
| F5 | Legal frameworks: add PIPL/CSL/DSL/PDPA (+ data-export) | **P1** | FIX |
| F6 | Legal role: stop sharing Sonnet with Validation; dedicated assignment | **P1** | FIX |
| F7 | CC-C tier signaling: independence seat ≠ Tier C afterthought | **P1** | FIX |
| F8 | Cap DeepSeek as universal fallback sink | **P1** | FIX |
| F9 | Pin Gemini model; ban floating `*-latest` in Enterprise default | **P1** | FIX |
| F10 | Map 10 scoring dimensions ↔ 12 phases explicitly | **P2** | FIX |
| D1 | 12-role expansion (Financial/Team/Legal) | — | DEFEND |
| D2 | 8-vendor / 33% hard cap direction | — | DEFEND |
| D3 | Cold-start + N gates in feedback loop | — | DEFEND |
| D4 | Premium pricing $249/$799/$1499 | — | DEFEND |
| D5 | VT vs RFP Tank boundary | — | DEFEND |
| D6 | Qwen on Market-Reality + Kimi on Reasoning | — | DEFEND |
| R1 | Live admin-ai path smoke-test all 12 | — | DEFER |
| R2 | Latency DAG + p95 measure before marketing 45s | — | DEFER |
| R3 | Technical Architecture + Geo-Market roles (v4.2) | — | DEFER |
| R4 | Multi-outcome (non-US) feedback ontology | — | DEFER |
---
## Per-Dimension Scores (machine-readable)
```json
{
"reviewer": "Cross-Check C — DeepSeek V4 Pro",
"spec": "judge-pool-spec v2.0",
"lens": "chinese-lab-independence",
"overall": 6.4,
"dimensions": [
{"id": 1, "name": "role_coverage", "score": 7, "verdict": "DEFEND"},
{"id": 2, "name": "model_assignments", "score": 7, "verdict": "DEFEND"},
{"id": 3, "name": "admin_ai_paths", "score": 6, "verdict": "DEFER"},
{"id": 4, "name": "vendor_diversity", "score": 8, "verdict": "DEFEND"},
{"id": 5, "name": "fallback_chains", "score": 4, "verdict": "FIX"},
{"id": 6, "name": "content_based_rules", "score": 4, "verdict": "FIX"},
{"id": 7, "name": "latency_budgets", "score": 6, "verdict": "DEFER"},
{"id": 8, "name": "feedback_loop", "score": 7, "verdict": "DEFEND"},
{"id": 9, "name": "pricing", "score": 7, "verdict": "DEFEND"},
{"id": 10, "name": "product_boundary_saas_default", "score": 5, "verdict": "FIX"}
],
"blocking_fixes": ["F1", "F2", "F3", "F4"],
"deployed_url_mismatch": "https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md still serves v1.0"
}
```
---
## Closing (Cross-Check C voice)
v2.0 finally treats Chinese labs as first-class citizens of the panel. That is real progress.
But independence is not a logo count. It is what happens when the primary is down, when the vertical is not YC-SaaS, and when the feedback loop decides whose disagreement is "bias." On those three tests the spec still thinks in Silicon Valley defaults while wearing an eight-vendor badge.
**Ship after F1F4. Everything else can follow.**
— Cross-Check C (DeepSeek V4 Pro), independent re-score, 2026-08-12