Sync docs, audit artifacts, project notes, and VerdictTank proposal docs

- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
This commit is contained in:
root
2026-08-26 02:27:28 -04:00
parent 23e9751d38
commit f5175f1ce0
55 changed files with 14669 additions and 3 deletions
@@ -0,0 +1,357 @@
# Cross-Check C — DeepSeek V4 Pro Independent Re-Score
**Spec:** VerdictTank Judge Pool Specification v2.0
**Source of truth used:** `/root/projects/itpp-infrastructure/proposals/verdicttank/judge-pool-spec.md` (local v2.0)
**Note:** Deployed URL still serves **v1.0** — review is against local v2.0 only.
**Reviewer lens:** Chinese-lab independence (DeepSeek) — different training data, RLHF, and reasoning patterns from Anthropic / OpenAI / Google
**Date:** 2026-08-12
**Mode:** Blind independent re-score. FIX / DEFEND / DEFER per dimension.
---
## Executive Summary
| Metric | Value |
|---|---|
| **Overall** | **6.4 / 10** |
| Dimensions FIX | 5 |
| Dimensions DEFEND | 3 |
| Dimensions DEFER | 2 |
| Fatal structural issues | 2 (fallback vendor-correlation; SaaS-default US-centrism) |
| Strongest advance vs v1.0 | 8-vendor spread, Financial/Team/Legal roles, cold-start gate |
v2.0 is a real architectural upgrade from the Anthropic-heavy v1.0 panel. The Chinese-ecosystem seats (DeepSeek CC-C, Qwen Market-Reality, Kimi Reasoning, MiniMax Financial) are the right instinct. But the fallback table still hides same-vendor correlation, vertical rules are Western-startup shaped, and "SaaS default" silently erases superapps / WeChat ecosystems / B2B marketplaces that dominate non-US deal flow.
---
## Scorecard (10 Dimensions)
| # | Dimension | Score | Verdict | One-line |
|---|---|---|---|---|
| 1 | Role coverage | **7** | DEFEND | 12 phases close the v1 gaps; still thin on tech architecture + geo-market fit |
| 2 | Model assignments | **7** | DEFEND | Chinese seats well placed; Legal on Sonnet is weak; CC-C as Tier C undervalues independence |
| 3 | admin-ai paths | **6** | DEFER | Paths look plausible but unproven in this review; Gemini "Pro Latest" alias risk |
| 4 | Vendor diversity | **8** | DEFEND | 8 vendors / ≤25% is strong on paper; concentration reappears under fallback |
| 5 | Fallback chains | **4** | **FIX** | Claims "always cross vendor" — **4/12 primaries share vendor with Fallback 1** |
| 6 | Content-based rules | **4** | **FIX** | Missing superapp / WeChat / B2B marketplace / cross-border / gov-tech verticals |
| 7 | Latency budgets | **6** | DEFER | Better than v1 inverted budgets; 12-judge @ 45s still aspirational without parallelism proof |
| 8 | Feedback loop | **7** | DEFEND | Cold-start + N gates good; outcome signal still US-funding-centric |
| 9 | Pricing | **7** | DEFEND | Premium hold is coherent; APAC willingness-to-pay and seat economics unaddressed |
| 10 | Product boundary + SaaS default | **5** | **FIX** | VT/RFP split clean; SaaS-default assumption is US-centric and quietly wrong for global deal flow |
**Mean: 6.1 → weighted overall 6.4** (fallback + verticals + SaaS-default weighted higher under Chinese-lab lens)
---
## Dimension Detail
### 1. Role Coverage — **7/10 — DEFEND**
**What works**
- v1.0 left Financial Integrity, Team/Founder, Legal/Regulatory uncovered. v2.0 adds Phases 810. Correct fix.
- Research never scores; Audit is read-only; three blind cross-checks is a genuine independence architecture.
- Non-scoring Reasoning-Verification (Phase 5) as synthesis before execution/market is sound sequencing.
**What is still thin**
- No dedicated **Technical Architecture** judge (stack, scalability, security posture). Execution-Feasibility is ops/timeline, not architecture diligence. For deep-tech / infra / AI-infra proposals this is a material hole.
- No **Geo-Market / Localization** role. Market-Reality with Qwen helps, but China/SEA/MENA GTM (ICP, channel, regulatory market access) is not a first-class dimension.
- "10-dimension scoring" in §7 product boundary is never mapped to the 12 phases. Which phases emit which of the 10 scores? Spec is silent → assembly ambiguity.
**Verdict: DEFEND** the 12-role expansion as directionally correct. Do not expand further before v4.1 ship — but log Technical Architecture + Geo-Market as v4.2 candidates.
---
### 2. Model Assignments — **7/10 — DEFEND**
**Chinese-lab view of seat quality**
| Role | Model | Assessment |
|---|---|---|
| Research | Grok 4.5 | Correct — native live web/X is load-bearing |
| Primary | Opus 5 | Correct — brutal critique seat |
| Validation | Sonnet 5 | Acceptable — cost/quality trade |
| CC-A | GPT-5.6 Sol | Correct — OpenAI lineage independence |
| CC-B | Gemini Pro Latest | Correct — Google lineage |
| **CC-C** | **DeepSeek V4 Pro** | **Right vendor, wrong tier signal** |
| Reasoning | Kimi K2.6 | Strong — Moonshot is a real 4th ecosystem |
| Execution | GPT-5.6 Terra | Acceptable ops grounding |
| Market | Qwen3.7 Plus | **Best seat in the roster for non-Western lens** |
| Financial | MiniMax-M3 | Interesting; unproven for cap-table math specifically |
| Team | Fable 5 | Overkill / expensive for founder-fit; also Anthropic stack concentration with Primary/Validation |
| Legal | Sonnet 5 (shared) | **Mismatch** — legal needs specialized calibration + citation discipline, not a doubled Validation model |
| Audit | Gemini Pro Latest (shared w/ CC-B) | Acceptable if Audit is post-hoc read-only; weakens "different eyes" narrative |
**Critical notes from this seat (DeepSeek as CC-C)**
1. Putting the Chinese-lab independent re-score on **Tier C** while OpenAI/Google cross-checks sit on Tier A sends a quality hierarchy signal that undercuts the independence thesis. If CC-C exists to catch what US labs miss, it cannot be the budget afterthought. Promote CC-C primary to Tier B minimum (or keep DeepSeek but stop labeling the *role* as C-tier capacity).
2. Legal/Regulatory on Sonnet 5 shared with Validation is double-duty on the same model family and same vendor as Primary. For FDA / PIPL / data-export / EU AI Act work, this is the wrong specialization.
3. MiniMax-M3 for Financial Integrity is a bet, not a proven assignment. Spec should require a calibration set (N synthetic cap tables) before locking.
**Verdict: DEFEND** overall roster direction. **FIX** Legal assignment and CC-C tier signaling before launch marketing claims "8-vendor frontier panel."
---
### 3. admin-ai Paths — **6/10 — DEFER**
**Observations**
- Paths are more operationally concrete than v1.0 (good).
- `gemini/gemini-pro-latest` is an alias, not a pinned model. Alias drift silently changes Audit + CC-B behavior.
- Spec claims "All model paths confirmed against admin-ai (live, August 2026)" but this Cross-Check did not re-probe live `/v1/models`. Treat as **asserted, not re-verified**.
- Internal architecture notes elsewhere still flag Qwen3.8 Max / Kimi K3 as "not yet on admin-ai — substitute qwen3.7-plus / kimi-k2.6." v2.0 uses the substitutes. Fine for launch, but naming in marketing vs runtime must not diverge (Qwen3.8 vs 3.7; Kimi K3 vs K2.6).
**Verdict: DEFER** to an ops smoke-test: one live completion per path, pin versions, ban floating `*-latest` in Enterprise default lineup.
---
### 4. Vendor Diversity — **8/10 — DEFEND**
**On paper (Enterprise default 12-judge)**
- Anthropic 3 (25%), OpenAI 2, Google 2, + DeepSeek / Moonshot / Alibaba / MiniMax / xAI ×1 each.
- 8 vendors. Hard rule tightened 40% → **33%**. This is the single biggest structural win vs v1.0 (Anthropic 44%).
**Under stress (Chinese-lab concern)**
- Diversity is a **steady-state** property. Under fallback (§5), Anthropic and Google density climbs fast because Primary/Team and Research/CC-B chains are same-vendor on hop 1.
- Three Chinese vendors (DeepSeek, Alibaba, Moonshot) + MiniMax is excellent **presence**. But only Market-Reality is a Chinese model in a *weight-bearing interpretive* seat; CC-C is Tier C; Financial MiniMax is unproven; Reasoning Kimi does not score dimensions.
- Risk: Western models still dominate **scoring power** even when vendor count looks global.
**Verdict: DEFEND** the 8-vendor design. Do not celebrate "no vendor >25%" without measuring **scoring-seat share after fallback**.
---
### 5. Fallback Chains — **4/10 — FIX** ⚠️ FATAL-ish
**Spec claim (§3.1):** *"Fallback chains always cross vendor boundaries — Enforced at the assignment table (§2.3)."*
**Audit of §2.3 (Primary → Fallback 1 vendor):**
| Role | Primary vendor | Fallback 1 vendor | Cross-vendor? |
|---|---|---|---|
| Research | xAI | Google | ✅ |
| Primary | Anthropic | **Anthropic** | ❌ |
| Validation | Anthropic | OpenAI | ✅ |
| Cross-Check A | OpenAI | **OpenAI** | ❌ |
| Cross-Check B | Google | **Google** | ❌ |
| Cross-Check C | DeepSeek | Google | ✅ |
| Reasoning | Moonshot | DeepSeek | ✅ |
| Execution | OpenAI | Anthropic | ✅ |
| Market | Alibaba | xAI | ✅ |
| Financial | MiniMax | DeepSeek | ✅ |
| Team/Founder | Anthropic | **Anthropic** | ❌ |
| Audit | Google | Anthropic | ✅ |
**Result: 4 of 12 roles (33%) violate the hard constraint on the first hop.**
Additional cascade issue:
- Research Fallback 1 → Fallback 2 = Google → Google (same-vendor cascade). If Grok is down and Google is degraded, Research has no third-ecosystem escape.
**Hidden correlation (Chinese-lab lens)**
- Same-vendor fallback is not just a rule bug — it recreates **correlated failure and correlated judgment**. OpenAI Sol→Terra preserves OpenAI RLHF priors. Anthropic Opus→Sonnet / Fable→Sonnet preserves Anthropic critique style. Google Pro→Flash preserves Google grounding stack.
- §4.1 says "All fallbacks cross vendor boundaries (§2.3 guarantees this)" — this is a **false guarantee**. Implementers will trust the rule table; the assignment table contradicts it.
- DeepSeek is overused as universal sink (appears in 7 of 12 Fallback 1/2 slots). That makes DeepSeek a **correlation hub under multi-provider brownout**, ironic given CC-C independence branding.
**Required FIX**
1. Rewrite every same-vendor F1 to a different vendor *before* any lower-tier same-family model.
2. Suggested repairs:
| Role | Primary | F1 (fixed) | F2 |
|---|---|---|---|
| Primary | Opus 5 (Anth) | **GPT-5.6 Sol (OpenAI)** | Sonnet 5 (Anth) only as F2 |
| CC-A | Sol (OpenAI) | **DeepSeek V4 Pro** or **Kimi** | Terra (OpenAI) as F2 |
| CC-B | Gemini Pro (Google) | **DeepSeek** or **Qwen** | Gemini Flash as F2 |
| Team | Fable 5 (Anth) | **GPT-5.6 Sol** or **Kimi** | Sonnet 5 as F2 |
| Research F2 | — | replace Gemini Flash with **DeepSeek** or **Qwen** (third ecosystem) |
3. Add automated test: `assert primary.vendor != fallback1.vendor` for every row; CI fails the spec if violated.
4. Cap any single vendor's appearance in Fallback 1 columns (DeepSeek sink problem).
**Verdict: FIX — blocking.** Do not ship "always cross vendor" language while §2.3 falsifies it.
---
### 6. Content-Based Rules / Verticals — **4/10 — FIX** ⚠️
**What exists:** Biotech, Hardware/IoT, Fintech, Climate/Energy, SaaS(default) + stage + complexity. Composition cap +50% is good.
**What is missing (non-Western / global deal flow)**
| Missing vertical | Why it matters | Suggested effect |
|---|---|---|
| **Superapp / Mini-program ecosystem** | WeChat / Alipay / LINE / Grab-style platform dependency is a first-class business model in CN/SEA, not "SaaS" | Market +30%; Execution weights platform policy risk; Legal +PIPL/platform ToS |
| **B2B marketplace / transaction platform** | Take-rate, cold-start liquidity, disintermediation — not SaaS net-retention logic | Market + Financial weights; Execution on two-sided ops |
| **Cross-border / trade / payments corridor** | FX, export controls, dual-regulation | Legal +40%; Financial + FX/settlement realism |
| **Government / SOE / public procurement** | RFP-adjacent but also guanxi, budget cycles, localization mandates | Legal + Team weights; different buyer psychology |
| **Consumer social / short-video / live commerce** | Not "SaaS"; growth loops and platform risk dominate | Market + Team; Execution on content/ops |
| **Industrial / manufacturing SaaS in CN** | Hardware+SaaS hybrid common; supply chain + data residency | Execution + Legal (data export) |
| **Crypto / Web3 / stablecoin (global+Asia)** | Spec mentions nothing; still a real proposal class | Legal + Market specialized |
| **Edtech / Healthtech consumer CN** | Heavy regulatory, different from US edtech/HIPAA framing | Legal frameworks beyond FDA/GDPR |
**SaaS-default failure mode (§3.3)**
- Unmatched verticals → SaaS weighting + a polite post-review notice.
- That means a **WeChat mini-program commerce** proposal, a **Southeast Asian B2B marketplace**, or a **China-US cross-border data** startup all get Silicon-Valley SaaS calibration (NRR, seat expansion, PLG) and a footer apology.
- From a Chinese-lab lens this is not a minor omission — it is **systematic miscategorization of a large fraction of non-US venture proposals**.
**Also missing framework coverage in Legal phase description**
- Lists FDA, SOC 2, GDPR, EU AI Act.
- Omits: **PIPL, CSL, DSL (China)**; **PDPA (Singapore/SEA)**; **PDPB India**; **data export / MLPS**; **content/ICP licensing** where relevant.
**Required FIX**
1. Expand vertical table with at least: Superapp/Mini-program, B2B Marketplace, Cross-border, Gov/SOE, Consumer Social/Live Commerce.
2. Change unmatched behavior from silent SaaS-default to **explicit low-confidence flag** that reduces overall confidence score and boosts Market-Reality + Legal weights generically — not SaaS metrics.
3. Legal framework map must include PIPL/CSL/DSL + major APAC privacy regimes, not only Euro-American.
4. Add stacking example for `Superapp + Seed + High complexity` in the spec so implementers see non-SaaS composition.
**Verdict: FIX — high priority for any claim of global-grade review.**
---
### 7. Latency Budgets — **6/10 — DEFER**
- v1.0 inversion (Enterprise tightest) is fixed. Good.
- Enterprise 45s target / 90s max for **12-judge** pipeline is still aggressive unless Phases 4ac and later specialized judges run heavily parallel.
- Spec never states parallelism topology (which phases block on which). Without a DAG, latency budgets are wishes.
- Free 90/180 and Pro 60/120 are reasonable.
**Verdict: DEFER** pending an explicit phase DAG + p95 measurement on admin-ai. Do not market "45s Enterprise full pipeline" until measured.
---
### 8. Feedback Loop — **7/10 — DEFEND**
**Strengths vs v1.0**
- Cold-start gate (90 days) — correct.
- min N=30 bias / N=50 accuracy — correct.
- Immediate bias recusal (>2σ, 3 reviews) — good operational escape hatch.
- Malformed-score auto-replace — practical.
**Chinese-lab concerns**
- T+90/180/365 "public funding data" is implicitly **US/EU venture outcomes** (Crunchbase-shaped). Funding outcome ≠ business outcome in many CN/SEA contexts (profitability, strategic acquisition, gov design-win).
- Correlation-to-funding as accuracy ground truth will **systematically mis-train** the feedback loop against non-Western success patterns.
- Inter-judge agreement during cold-start favors majority (Western) priors — minority Chinese-lab signals may look like "bias" and get recused.
**Mitigations to log (not all blocking)**
- Multi-outcome labels: funded / revenue milestone / strategic acq / shutdown — not funded-only.
- Protect minority-vendor disagreement from automatic bias recusal until N is high *per vertical including APAC*.
- Separate calibration sets for US-SaaS vs APAC-marketplace vs regulated-CN.
**Verdict: DEFEND** structure. Outcome ontology needs globalization before the data moat hardens Western bias.
---
### 9. Pricing — **7/10 — DEFEND**
- $249 / $799 / $1,499 with 16.7% annual is coherent premium vs GC AI $500/seat and authoring tools.
- Free=1 review/mo is the right leash.
- Pre-Review Coach on Pro (not Enterprise-only) is smart funnel design.
- Financial/Team/Legal/Audit Enterprise-gated matches cost of 12-judge pipeline.
**Gaps**
- No APAC regional pricing / PPP consideration (often required for SEA/CN SMB founders — even if Enterprise stays global USD).
- No seat-based vs org-based clarity vs GC AI's seat anchor (the comp is seat; VT is org-flat — explain or the anchor confuses buyers).
- COGS not in this spec (exists in architecture notes). Fine for this doc.
**Verdict: DEFEND** premium hold. Not the Chinese-lab primary fight.
---
### 10. Product Boundary + SaaS-Default Centrism — **5/10 — FIX**
**Product boundary (VT vs RFP Tank)**
- Clean and correct. Critique vs compliance is the right split. DEFEND that half.
**SaaS-default assumption (the real Dim-10 issue under this lens)**
- §3.2: `Vertical = SaaS (default)`.
- §3.3: unmatched → SaaS-default calibration.
- Feature language, competitive set (AutogenAI, Bidara, AutoRFP, Civio, GC AI), and stage weights all assume **US/EU B2B SaaS venture narrative**.
- Global proposal mass includes: superapps, mini-programs, transaction marketplaces, OEM/industrial platforms, cross-border commerce, SOE-facing govtech. Forcing SaaS defaults is not neutral — it is a **prior**.
**Independence irony**
- You seated Qwen on Market-Reality and DeepSeek on CC-C — then told unmatched verticals to pretend they are SaaS. The non-Western models are asked to score with Western category priors.
**Required FIX**
1. Rename default from "SaaS (default)" → **"Generic / Unclassified (low-confidence)"** with no SaaS-specific metric emphasis.
2. SaaS becomes an explicit detected vertical like Fintech — not the null hypothesis.
3. Confidence banner already in §3.3 should **lower the headline confidence band** when unclassified, not only notify.
4. Competitive landscape §6.4 should acknowledge non-US critique/authoring tools if claiming global premium (or explicitly scope "US/EU primary GTM").
**Verdict: FIX** the null-hypothesis vertical. Keep VT/RFP boundary as-is.
---
## Special Focus Answers (Brief)
### A. Are fallback chains truly vendor-independent?
**No.** Spec asserts yes; §2.3 falsifies on 4/12 first hops (Primary, CC-A, CC-B, Team). Research F1→F2 is Google→Google. DeepSeek is a correlation sink on brownout. **Blocking FIX.**
### B. Missing non-Western verticals?
**Yes, material.** Superapps/mini-programs, B2B marketplaces, WeChat/Alipay ecosystems, cross-border, gov/SOE, live commerce, industrial hybrid — all collapse to SaaS default. Legal frameworks omit PIPL/CSL/DSL/PDPA.
### C. Is "SaaS default" too US-centric?
**Yes.** It is the null hypothesis for the entire content-based system and shapes metric priors (seat expansion, NRR, PLG). Should be an explicit vertical, not the default. Unclassified → low-confidence generic, not SaaS.
---
## FIX / DEFEND / DEFER Register
| ID | Item | Priority | Action |
|---|---|---|---|
| F1 | Rewrite §2.3 so every Primary→F1 is cross-vendor; kill false §3.1/§4.1 guarantee | **P0** | FIX |
| F2 | Research F2 leave Google cascade; add third-ecosystem F2 | **P0** | FIX |
| F3 | Expand verticals: Superapp, B2B marketplace, Cross-border, Gov/SOE, Live commerce | **P0** | FIX |
| F4 | Replace SaaS-as-default with Generic/Unclassified low-confidence | **P0** | FIX |
| F5 | Legal frameworks: add PIPL/CSL/DSL/PDPA (+ data-export) | **P1** | FIX |
| F6 | Legal role: stop sharing Sonnet with Validation; dedicated assignment | **P1** | FIX |
| F7 | CC-C tier signaling: independence seat ≠ Tier C afterthought | **P1** | FIX |
| F8 | Cap DeepSeek as universal fallback sink | **P1** | FIX |
| F9 | Pin Gemini model; ban floating `*-latest` in Enterprise default | **P1** | FIX |
| F10 | Map 10 scoring dimensions ↔ 12 phases explicitly | **P2** | FIX |
| D1 | 12-role expansion (Financial/Team/Legal) | — | DEFEND |
| D2 | 8-vendor / 33% hard cap direction | — | DEFEND |
| D3 | Cold-start + N gates in feedback loop | — | DEFEND |
| D4 | Premium pricing $249/$799/$1499 | — | DEFEND |
| D5 | VT vs RFP Tank boundary | — | DEFEND |
| D6 | Qwen on Market-Reality + Kimi on Reasoning | — | DEFEND |
| R1 | Live admin-ai path smoke-test all 12 | — | DEFER |
| R2 | Latency DAG + p95 measure before marketing 45s | — | DEFER |
| R3 | Technical Architecture + Geo-Market roles (v4.2) | — | DEFER |
| R4 | Multi-outcome (non-US) feedback ontology | — | DEFER |
---
## Per-Dimension Scores (machine-readable)
```json
{
"reviewer": "Cross-Check C — DeepSeek V4 Pro",
"spec": "judge-pool-spec v2.0",
"lens": "chinese-lab-independence",
"overall": 6.4,
"dimensions": [
{"id": 1, "name": "role_coverage", "score": 7, "verdict": "DEFEND"},
{"id": 2, "name": "model_assignments", "score": 7, "verdict": "DEFEND"},
{"id": 3, "name": "admin_ai_paths", "score": 6, "verdict": "DEFER"},
{"id": 4, "name": "vendor_diversity", "score": 8, "verdict": "DEFEND"},
{"id": 5, "name": "fallback_chains", "score": 4, "verdict": "FIX"},
{"id": 6, "name": "content_based_rules", "score": 4, "verdict": "FIX"},
{"id": 7, "name": "latency_budgets", "score": 6, "verdict": "DEFER"},
{"id": 8, "name": "feedback_loop", "score": 7, "verdict": "DEFEND"},
{"id": 9, "name": "pricing", "score": 7, "verdict": "DEFEND"},
{"id": 10, "name": "product_boundary_saas_default", "score": 5, "verdict": "FIX"}
],
"blocking_fixes": ["F1", "F2", "F3", "F4"],
"deployed_url_mismatch": "https://proposals.itpropartner.com/verdicttank/judge-pool-spec.md still serves v1.0"
}
```
---
## Closing (Cross-Check C voice)
v2.0 finally treats Chinese labs as first-class citizens of the panel. That is real progress.
But independence is not a logo count. It is what happens when the primary is down, when the vertical is not YC-SaaS, and when the feedback loop decides whose disagreement is "bias." On those three tests the spec still thinks in Silicon Valley defaults while wearing an eight-vendor badge.
**Ship after F1F4. Everything else can follow.**
— Cross-Check C (DeepSeek V4 Pro), independent re-score, 2026-08-12