From 65b976027dc951ebbf751641bf89faddc4891973 Mon Sep 17 00:00:00 2001 From: root Date: Sun, 16 Aug 2026 09:18:10 -0400 Subject: [PATCH] rename Wall-O to Scirium: move project docs, add vision/positioning spec, changelog entry --- .../{wall-o => scirium}/01-product-model.md | 0 .../{wall-o => scirium}/02-m365-connector.md | 0 .../03-deployment-white-label.md | 0 projects/scirium/04-business-proposal.md | 447 ++++++++++++++++++ 4 files changed, 447 insertions(+) rename projects/{wall-o => scirium}/01-product-model.md (100%) rename projects/{wall-o => scirium}/02-m365-connector.md (100%) rename projects/{wall-o => scirium}/03-deployment-white-label.md (100%) create mode 100644 projects/scirium/04-business-proposal.md diff --git a/projects/wall-o/01-product-model.md b/projects/scirium/01-product-model.md similarity index 100% rename from projects/wall-o/01-product-model.md rename to projects/scirium/01-product-model.md diff --git a/projects/wall-o/02-m365-connector.md b/projects/scirium/02-m365-connector.md similarity index 100% rename from projects/wall-o/02-m365-connector.md rename to projects/scirium/02-m365-connector.md diff --git a/projects/wall-o/03-deployment-white-label.md b/projects/scirium/03-deployment-white-label.md similarity index 100% rename from projects/wall-o/03-deployment-white-label.md rename to projects/scirium/03-deployment-white-label.md diff --git a/projects/scirium/04-business-proposal.md b/projects/scirium/04-business-proposal.md new file mode 100644 index 0000000..cfa26ac --- /dev/null +++ b/projects/scirium/04-business-proposal.md @@ -0,0 +1,447 @@ +# Wall-O Business Proposal + +Status: OPEN (v1 draft, pre-critical-review) +Date: 2026-08-16 +Author: Sho'Nuff (Conductor) + Wall-O marketing team +First client: Wall Orthodontics (pilot beachhead) +Product scope: business-agnostic internal staff knowledge-base chat + +--- + +## Table of Contents + +1. Executive Summary +2. Elevator Pitch +3. Problem Statement (business-agnostic) +4. Market Analysis (TAM/SAM/SOM, competitive, real-world examples) +5. Product Overview (v1 definition, architecture, v2 roadmap) +6. Revenue Model (pricing tiers) +7. Competitive Advantages (moat) +8. Go-to-Market Strategy +9. Risk Analysis (pre-mortem) +10. Financial Projections (is it worth building) +11. The Ask (infrastructure, decisions, success) + +--- + +## 1. Executive Summary + +Wall-O is a multi-tenant internal staff knowledge-base chat platform. Every knowledge domain inside a business (employee handbook, billing, IT help, policies, procedures) becomes its own scoped AI chat channel, backed by that business's real documents. Staff ask questions in a chat interface and get grounded, cited answers drawn from the company's own files - never hallucinated, never leaking across domains. + +The core problem is universal and business-agnostic: employees waste a documented, measurable fraction of their week hunting for information that already exists somewhere in the organization. McKinsey and Gartner studies put that fraction in the double digits of weekly hours. Enterprise solutions (Glean, Moveworks, Coveo) solve this for large companies at enterprise prices and enterprise procurement complexity. The SMB and vertical-niche middle - dental practices, law firms, restaurants, retail groups, logistics firms - is priced out and under-served. + +Wall-O's answer is a hard invariant: one channel equals one knowledge domain, one attached AI agent, and one scoped knowledge source. This makes answers provably grounded (every reply carries citations to the source document) and provably isolated (a billing question can never surface a patient file or a neighboring tenant's data). The product is self-hosted, white-label, and business-agnostic by construction - the same core serves a 12-person dental practice and a 200-person law firm. + +Financially, Wall-O is a high-margin software product running on infrastructure IT Pro Partner already owns: ~$68K build cost, ~$0.50 per 1,000 queries marginal cost, ~95-98% gross margin, ~$1,000 CAC against ~$14K LTV, and full build-cost payback in ~10-13 months (realistic). The v1 is a grounded Q&A assistant. The v2 roadmap adds risk-tiered "do" capabilities (reporting, content generation, integrations) that deepen the moat without ever touching patient/PHI scope. + +This proposal defines v1 and v2 concretely, prices it against the market, and makes the fact-based case for building it. It is submitted for critical review. + +--- + +## 2. Elevator Pitch + +Wall-O is the AI coworker that reads your business's actual documents and answers your team's questions with citations - for the price of a streaming subscription, on infrastructure you control, scoped so nothing leaks. It is "a private, cited ChatGPT for your employee handbook, billing rules, and policies" - one tightly-scoped assistant per knowledge domain, white-labeled to your brand. + +--- + +## 3. Problem Statement (business-agnostic) + +### 3.1 The universal problem + +Every business with more than a handful of employees has the same problem, regardless of industry: + +- New hires take weeks to onboard because institutional knowledge lives in scattered SharePoint libraries, PDFs, and managers' heads. +- Repetitive questions ("how do I submit a PTO request", "what's our return policy", "who do I escalate a billing dispute to") interrupt managers and senior staff constantly. +- Policy and procedure changes never propagate - staff operate on outdated rules. +- The people who know the answer are busy; the answer itself is already written down somewhere, just not findable. + +This is not an industry-specific problem. A dental practice has the same shape of problem as a law firm, a restaurant group, a retail chain, a logistics company, or a real estate office. The documents differ; the pain is identical. + +### 3.2 How people solve it today + +| Current method | Time cost | Pain level | Why it fails | +|---|---|---|---| +| Ask a manager / senior coworker | High (interrupts them every time) | High | Doesn't scale; same questions repeat | +| Search SharePoint / Google Drive manually | Medium (10-20 min per hunt) | Medium | Fragmented, no ranking, poor recall | +| Printed binders / shared docs | Medium | Medium | Goes stale immediately | +| Generic LLM (public ChatGPT) | Low | High (risk) | Hallucinates, no access to internal docs, leaks data | +| Enterprise knowledge AI (Glean, Moveworks, Coveo) | Low | Low (but expensive) | $30+/seat/mo, enterprise procurement, overkill for SMB | + +The gap: there is no affordable, self-hosted, white-label answer for the SMB and vertical-niche middle. Enterprise tools solve the problem but are priced and packaged for enterprises. Public LLMs are cheap but ungrounded and unsafe for internal documents. Wall-O sits in that gap. + +### 3.3 The precise target user + +The buyer is the owner or office manager of an SMB (10 to 200 employees) who: +- Already uses Microsoft 365 (so documents live in SharePoint/OneDrive), +- Is frustrated by repetitive questions and slow onboarding, +- Will not pay enterprise pricing or survive enterprise procurement, +- Values data control (keeps documents on their own infrastructure), +- Wants a white-label experience that feels like their brand, not a third-party tool. + +The beachhead vertical is orthodontics/dental (Wall Orthodontics), but the product model is explicitly business-agnostic - the same channel primitive serves any document-driven business. + +### 3.4 The gap Wall-O fills + +| Need | Public LLM | Enterprise AI | Wall-O | +|---|---|---|---| +| Grounded, cited answers from my docs | No | Yes | Yes | +| Affordable for SMB | Yes (but unsafe) | No | Yes | +| Self-hosted / data control | No | Partial | Yes | +| White-label to my brand | No | Partial | Yes | +| Channel-scoped (no cross-domain leak) | No | Partial | Yes (hard invariant) | +| No PHI / patient scope creep | No guarantee | No guarantee | Enforced at ingestion | + +--- + +## 4. Market Analysis + +### 4.1 Market size (bottom-up, business-agnostic) + +Methodology: per-seat SaaS anchored on knowledge-worker count ("every employee could use an internal Q&A assistant"), blended at $10/user/month - well below the enterprise incumbents that charge $30-75 with minimums. + +``` +TAM = 1,000,000,000 global knowledge workers x $10/mo x 12 = $120B/year +SAM = 1,000,000,000 x 45% (SMB share of workers, SBA 45.9%) = 450M workers + x $10/mo x 12 = $54B/year +SOM = $54B x 1% (3-5 yr category-winner) = $540M/year + $54B x 0.1% (year 1-2 realistic) = $54M/year +``` + +US-only cross-check (tighter, verifiable): 100M US knowledge workers x 45.9% = ~46M SMB workers x $120/yr = **$5.5B US SMB SAM**. 1% capture = $55M/yr; 5% = $275M/yr. Even a single vertical slice (e.g. ~1M US small professional-services firms) is a meaningful SOM. + +Sources: ~1B global knowledge workers (Schroders); ~100M US knowledge workers (Upwork/BLS/Eurostat); SMBs employ 45.9% of US private-sector workers (SBA 2024). + +### 4.2 Growth rate and trends + +- AI in Knowledge Management: $6.7B (2023) to $62.4B (2033), **25% CAGR** (Market.us). +- Knowledge Management Software: $14.56B (2025) to $70.01B (2035), **16.9% CAGR**, fastest segment "intelligent chatbots and virtual agents" (MRFR). +- Enterprises deploying internal AI chatbots report **45-55% reduction in query resolution time** (MRFR). +- SMB AI adoption rose from 5.2% (Jan 2023) to 17.7% (end 2025), with entry cost falling to $20-30/mo (JPMorgan Chase Institute). The demand is rising and the price point is coming to Wall-O's band. + +### 4.3 Real-world examples (the fact-based case to build) + +These are the proof points that the category is real, funded, and acquirers pay billions for it: + +| Company | Signal | Number | +|---|---|---| +| Glean | Raised $765M, valued **$7.2B** (Series F, Jun 2025) | Passed $100M ARR 2025; claims up to **110 hours saved/user/year** | +| Moveworks | Acquired by ServiceNow for **$2.85B** (Mar 2025) | Customers HP, Unilever, Toyota, Marriott; **70,000 hours reclaimed** at one automaker; 75,000 hours at a biopharma | +| Microsoft 365 Copilot | **20M+ paid seats** (of 450M M365 seats) | 75% of knowledge workers already use gen AI at work (MSFT Work Trend Index) | +| Notion AI | **$500M annualized revenue** (Sept 2025) | AI add-on attach rate grew 10-20% to 50%+ in one year | + +On the "every business has this problem" claim, the independent studies converge regardless of industry or size: + +- McKinsey Global Institute: knowledge workers spend **~19% of time** searching/tracking down information (plus 14% communicating). +- IDC: ~2.5 hours/day, ~30% of the workday, spent searching; 60% of executives say staff could not find what they needed. +- Deloitte: **>25% of time** spent searching for information. + +These cluster around 20-30% of the workday lost to search - a universal, business-agnostic cost, not an enterprise-only one. That is the budget Wall-O attacks. + +### 4.4 Why the SMB/vertical segment is underserved + +1. Enterprise tools price SMBs out on purpose. Glean is ~$50-75/seat with a 100-seat minimum (~$60K/yr floor); M365 Copilot is $30/seat on top of a required M365 license; Moveworks is $100K+/yr; Coveo is opaque enterprise QPM. None serve a 10-30 person practice. +2. SMB AI budgets are tiny - median ~$28-50/month total (JPMorgan Chase Institute). A $60K/yr enterprise contract is 100x an SMB's entire AI budget. +3. Only 17.7% of SMBs had adopted AI by end-2025, versus near-universal experimentation in large enterprises - while 91% of AI-using SMBs report revenue impact (Salesforce). The wedge is a cheap, self-hosted, vertical-aware tool - exactly Wall-O. + +### 4.5 Competitive landscape (condensed) + +Full 10-competitor analysis is in the research appendix. The key fact: **the $300-800/month self-hosted band for 10-30 seat SMBs is structurally empty.** + +| Competitor | Pricing | Target | Wall-O's gap | +|---|---|---|---| +| Glean | $50-75/seat, 100-seat min (~$60K/yr) | Enterprise | Self-host + white-label + vertical templates + SMB price | +| Moveworks/ServiceNow | $100K+/yr | Enterprise IT/HR | Not ticket-automation; general internal knowledge | +| Guru | $10-25/seat, 10-seat min | Mid-market | No self-host, no white-label, no channel scoping | +| Notion AI | ~$10/seat add-on | General workspace | Q&A only over Notion content; no M365 connector | +| M365 Copilot | $21-30/seat + M365 base | M365 tenants | Microsoft-cloud lock-in; no self-host/white-label/scoping | +| Coveo | Enterprise QPM | Fortune 1000 | No SMB tier | +| Dust | $30-150/seat credits | AI-operator teams | DIY agent-builder, not turnkey | +| Open WebUI / Onyx / AnythingLLM | Free self-host | Developers | Raw RAG kits - no multi-tenant isolation, no M365 connector, no vertical layer | + +Two strategic notes from the research: + +1. The two tools an SMB owner already has are **M365 Copilot and Notion AI** - so Wall-O must lead with self-hosted data control + vertical specificity, not generic "AI chat over docs", which those already do. +2. **No competitor is white-label.** That is the cleanest MSP/reseller angle - ITPP can sell Wall-O under a partner's brand where Glean/Guru/Notion/Copilot cannot. + +--- + +## 5. Product Overview + +### 5.1 The core primitive (what makes Wall-O different) + +Every Wall-O channel is exactly three things bound together, and this is a hard invariant: + +1. One knowledge domain (e.g. "Employee Resources", "Billing", "IT Help"). +2. One attached AI agent (a domain-tuned persona). +3. One scoped knowledge source (one vector namespace over an allowlisted set of documents). + +A channel has exactly one agent and exactly one scope. A channel cannot span two domains, and an agent cannot serve two channels. If a business needs a second domain, it creates a second channel, a second agent, and a second scope. This invariant is what makes answers grounded (only one scope's documents feed the answer) and isolated (no cross-domain or cross-tenant leak), and it is enforced in the schema with unique constraints running in both directions. + +### 5.2 Architecture (canonical split) + +| Layer | Responsibility | Owns | +|---|---|---| +| Rocket.Chat | Chat transport only | One workspace per tenant (MIT core, EE stripped). Rooms, users, messages. Zero intelligence. | +| Orchestrator | All intelligence | Multi-tenant FastAPI service. Tenancy, agents, kb_scope, M365 connector, retrieval, LLM, posting. | +| Postgres + pgvector | State and vectors | Tenants, channels, agents, scopes, documents, chunks, messages. | +| admin-ai (LiteLLM) | LLM | DeepSeek V4 Pro primary, configured fallback chain. | +| Wasabi S3 | Object storage | M365 sync staging, backups, agent assets, audit exports. | + +The intelligence never lives in the chat layer. Rocket.Chat is a dumb transport; the orchestrator owns everything that thinks. This split is what makes multi-tenancy clean and white-labeling a server-side concern rather than a per-tenant code fork. + +### 5.3 v1 definition (grounded Q&A) + +v1 is the answer engine, nothing more. A staff member @mentions the channel's agent (or DMs it), and gets back a grounded, cited answer drawn from that channel's scoped documents. + +**The v1 message flow (6 steps):** + +1. Staff @mentions the agent in a channel (or DMs it). +2. Rocket.Chat fires a signed webhook to the orchestrator (HMAC-authenticated, timestamp-skew-checked). +3. Orchestrator resolves tenant -> channel -> agent -> scope in one request context, dedupes on message id. +4. Retrieval: pgvector semantic search over the channel's namespace, always filtered by `tenant_id` AND `kb_scope_id`. Falls back to Microsoft Graph search if the top score is below threshold. +5. Prompt build: system persona + retrieved chunks labeled `[1]`, `[2]`... + citation instruction + question. The agent is told to answer only from context and to say "I could not find an answer" rather than guess. +6. Answer posts back as the agent bot, with inline `[n]` citations and a Sources footer linking each cited document. + +**v1 hard behaviors (SETTLED):** +- Never hallucinate: empty/below-threshold retrieval returns "I could not find an answer", with no sources. +- Every answer carries citations to source documents. +- Strictly internal staff knowledge. No patient records, no PHI. Ingestion is allowlisted per scope and a content filter flags PHI markers (SSN, MRN, DOB+name) before indexing. +- Full audit log: every inbound question and outbound answer is a `messages` row with tokens, model, latency, and citations. + +**What is already built vs. to build (v1):** + +| Component | Status | Effort | +|---|---|---| +| Product model + data architecture | SETTLED (3 design docs committed) | Done | +| Data model (tenants, channels, agents, kb_scopes, documents, chunks, messages, users) | Spec'd, not built | Build | +| Orchestrator (FastAPI, tenancy, retrieval, LLM client) | Not built | Build | +| Postgres + pgvector schema | Spec'd, not built | Build | +| Rocket.Chat workspace + bot integration | Phase 0 checklist spec'd, not built | Build | +| M365 connector (Sites.Selected read) | Spec'd, not built | Build | +| White-label server rebrand (fossify FOSS build) | Phase 1 gate, not started | Build | +| Auth (Hexclave/Stack Auth, Entra OIDC for O365) | Partially existing (Hexclave on app3) | Integrate | + +**Phase 0 proof slice:** stand up one Rocket.Chat workspace for Wall Orthodontics, register a bot, create a channel, attach an agent, index one document, and close the loop on a single grounded answer. The 21-step Phase 0 checklist is spec'd and ready to execute. + +### 5.4 v2 roadmap (planned upgrades) + +v2 adds risk-tiered "do" capabilities on top of the v1 answer engine. These are deliberately ordered by write-risk, and none of them ever unlock patient/PHI or autonomous destructive action. + +| Capability | Tier | Write risk | Auto-approve? | Value | +|---|---|---|---|---| +| Reporting (structured extraction + SQL aggregation + chart render) | Reporting | Zero new write risk | Yes | Highest value, zero risk - recommended first v2 ship | +| Content generation (drafts, summaries, boilerplate) | Content generation | Drafts only | Yes | Saves drafting time | +| Housekeeping (move/rename/delete-to-recycle) | Housekeeping | Destructive | No - propose/approve/audit | Keeps knowledge fresh | +| Integrations (Graph delegated sendMail/calendar/webhooks) | Integrations | External side effects | No - approval | Connects Wall-O to workflows | +| Design/brand (template render + image gen) | Design/brand | Drafts only | Yes (template-driven only) | Flyers, branded assets | + +**Write rails (SETTLED):** destructive actions require propose -> approve -> narrow audited write with before/after and undo path. Harmless writes auto-approve by risk tier. Per-agent tool scoping preserves one-agent-one-domain. Template-driven rendering is required for text-accurate flyers; raw image generation is unreliable for text and is not shipped for that use. + +**Permanently out of scope (never unlocks):** patient records, PHI, HIPAA scope, clinical/x-ray image analysis (separate FDA-regulated product), and autonomous destructive actions. These are hard rails, not feature gaps. + +### 5.5 Business-agnostic by construction + +The v1 and v2 product model references no industry. "Tenant" is any business; "channel" is any knowledge domain; "documents" are any document type. The M365 connector reads SharePoint/OneDrive libraries generically. The only orthodontics-specific artifacts are the pilot client's name and branding. A law firm, restaurant group, or logistics company is onboarded by the same provisioning path with different documents, different channel names, and different branding - zero code change. + +--- + +## 6. Revenue Model + +### 6.1 Pricing philosophy + +Premium positioning, not race-to-bottom. Wall-O is "the most accurate, best-scoped internal knowledge assistant", not the cheapest. Price for value delivered (time saved, errors avoided), not competitor undercutting. The product is self-hosted software with high gross margin, so price is set by willingness-to-pay in the segment, not by our marginal cost. + +### 6.2 Tiers (from the financial model) + +| Tier | Price | Includes | +|---|---|---| +| Starter | $199/tenant/mo (up to 10 users, +$15/user beyond) | 1 M365 connector (SharePoint + OneDrive), Rocket.Chat, 50K docs, 1 channel scope set | +| Business | $499/tenant/mo (up to 25 users, +$25/user beyond) | + Teams connectors, unlimited channel scoping, SSO (OIDC/SAML), 250K docs, SLA support | +| Enterprise | $2,000/tenant/mo (annual contract) | Dedicated instance, unlimited seats, custom connectors, SCIM, DPA, custom retention, white-glove onboarding, 99.9% SLA | + +Blended ARPU at a 60/30/10 tier mix = $469/mo (modeled at $450 for safety). Every tier includes the core grounded Q&A engine, the channel-scoped isolation invariant, citations, audit logs, and the no-PHI guardrail - higher tiers unlock more v2 "do" capabilities and white-glove onboarding, not better core answer quality. + +### 6.3 Revenue streams + +1. Recurring SaaS subscription (primary). +2. One-time onboarding/setup fee (document library audit, channel configuration). +3. White-label native app (Phase 2) as a premium add-on. +4. Optional dedicated-infrastructure deployment (single-tenant netcup/Hetzner instance) for compliance-sensitive buyers. + +--- + +## 7. Competitive Advantages (moat) + +### 7.1 The hard invariant is the moat + +The one-channel-one-agent-one-scope invariant is not a feature a competitor can bolt on after the fact; it is the schema. It makes every answer provably grounded and provably isolated, and it is what lets Wall-O serve regulated-adjacent verticals (dental, legal) where data leakage is a dealbreaker. Enterprise tools do broad search over everything; Wall-O does scoped answers over exactly one domain. + +### 7.2 SMB price point with enterprise-grade isolation + +Wall-O pairs SMB affordability with the isolation discipline usually found only in enterprise platforms. That combination is the open gap in the market. + +### 7.3 Self-hosted and white-label + +Data never leaves the customer's estate (or ITPP's controlled netcup estate). Chat, documents, and vectors stay self-hosted; only the LLM call routes through the internal LiteLLM proxy, and by scope it never carries PHI. White-labeling is a server-side rebrand (fossify FOSS build), so every tenant feels like their own branded product, not a resold third-party tool. + +### 7.4 Distribution via existing ITPP MSP + +IT Pro Partner already runs the infrastructure and has the MSP relationship model. Wall-O is not a cold-start consumer product; it slots into an existing managed-services motion where the first customer (Wall Orthodontics) is already warm. + +### 7.5 The moat compounds in v2 + +The risk-tiered "do" capabilities (reporting first, then integrations) deepen switching cost once a tenant's staff builds habits and workflows on the assistant. Reporting is the recommended first v2 ship because it is highest value with zero new write risk. + +--- + +## 8. Go-to-Market Strategy + +### 8.1 Beachhead: Wall Orthodontics + +Prove the loop with one real client in one vertical. Phase 0 closes a single grounded answer; Phase 1 delivers a white-labeled workspace with 3-5 channels (Employee Resources, Billing, IT Help, and a marketing/design channel as a tool-belt extension). Collect concrete ROI (questions answered, time saved, onboarding speed) for the case study. + +### 8.2 Vertical expansion from the beachhead + +The orthodontics win becomes a repeatable template for adjacent verticals: general dental, then any document-driven SMB (law firms, accounting, real estate, restaurants, logistics). Each vertical gets a tailored channel template (e.g. "billing" channel for a dental practice vs. "matter intake" for a law firm) but zero core code change - the business-agnostic model pays off here. + +### 8.3 Channel: MSP-led, not direct consumer + +Sell through IT Pro Partner's managed-services relationships and peer referral (practice-to-practice, firm-to-firm). The pitch is "the AI that reads your actual documents", demonstrated with a live tenant, not a slide deck. White-label means each MSP or franchise group can offer it under their own brand. + +### 8.4 Motion + +1. Live demo on a real (anonymized) tenant - show a cited answer, not a mockup. +2. 14-day pilot on the prospect's own SharePoint library. +3. Convert pilot to subscription with onboarding fee. + +--- + +## 9. Risk Analysis (pre-mortem) + +| Risk | Likelihood | Impact | Mitigation | +|---|---|---|---| +| LLM answer quality (hallucination) | Medium | High | Never-fabricate rail, citation requirement, below-threshold "I could not find an answer", model fallback chain | +| Data leakage across tenants/channels | Low | Critical | Hard invariant, tenant_id on every row, RLS, vector namespace isolation, HMAC webhook auth | +| PHI accidentally ingested | Low | Critical | Allowlisted libraries only, PHI-marker pre-ingest filter, block-and-log | +| M365 Graph API rate limits / connector fragility | Medium | Medium | Delta sync, exponential backoff, Graph search fallback | +| Rocket.Chat EE licensing in resale | Low | High | fossify FOSS-only build, never ship stock EE image | +| Native app store cost (Phase 2) | Medium | Low | PWA first, native only after PWA validated | +| Slow build (scope creep into "do" features) | Medium | Medium | v1 is answer-only; v2 capabilities are risk-tiered and gated | +| Competition (enterprise tools move downmarket) | Medium | Medium | SMB price point + isolation + white-label + MSP distribution | +| Churn if ROI not demonstrated | Medium | High | Measure time-saved from day one, report it to the buyer monthly | + +### 9.1 The one risk that kills the product + +Hallucination. If Wall-O ever gives staff a wrong, confidently-stated answer about policy or billing, trust collapses and the product is done. The never-fabricate rail, mandatory citations, and below-threshold "I could not find an answer" behavior are not features - they are the product. This is why the answer engine ships before any "do" capability, and why the v2 write features are risk-tiered with approval gates. + +--- + +## 10. Financial Projections (is it worth building) + +### 10.1 Build cost + +Assumption: $100/hr blended engineering rate (loaded senior full-stack or mid contractor), no new hardware (runs on existing netcup estate). + +| Component | Hours | Note | +|---|---|---| +| M365 Graph connector (ACL mapping, delta sync, webhooks) | 145 | Hardest single line | +| Ingestion pipeline (parse, chunk, embed, pgvector) | 70 | | +| RAG retrieval + grounded generation + citations | 90 | | +| Multi-tenancy (isolation, config, vector namespaces) | 50 | | +| Rocket.Chat integration | 70 | | +| Auth + SSO (Stack Auth / Hexclave) | 70 | | +| Admin UI + billing + onboarding | 90 | | +| Testing, security, observability, deploy, docs | 90 | | +| **Total** | **675** | | + +**Build cost = 675 hrs x $100 = ~$68K** (range $45K-$100K). Ongoing engineering ~$3,500/mo post-launch. + +### 10.2 Marginal cost (near zero) + +DeepSeek V4 Flash via LiteLLM ($0.14/M in, $0.28/M out). Per answer: ~2,000 input + ~500 output tokens. + +``` +LLM = (0.002M x $0.14) + (0.0005M x $0.28) = $0.00042/query +Embedding + storage + retries (amortized) = ~$0.00008/query +TOTAL = ~$0.0005/query + = ~$0.50 per 1,000 queries +``` + +Sensitivity: even at V4 Pro rates ($0.435/$0.87) it is $1.31/1K queries; a 5x price rise is ~$2.10/1K queries. All negligible. Per-tenant COGS is ~$3-30/mo depending on tier. **Gross margin ~95-98%.** + +### 10.3 Unit economics + +| Metric | Value | +|---|---| +| Blended ARPU | ~$450/mo (tier-mix math gives $469) | +| Blended COGS/tenant | ~$7/mo | +| Gross margin | ~95-98% | +| Churn | 3%/mo base (5% conservative) | +| LTV | ~$14K (range $8.5K-$21K) | +| CAC | ~$1,000 (owned MSP distribution is the moat) | +| LTV:CAC | ~14:1 (healthy SaaS is >3:1) | +| Payback per tenant | ~2.3 months | + +### 10.4 12-month ramp (3 scenarios, churn excluded) + +| Scenario | Tenants (month 12) | MRR (month 12) | 12-month revenue | +|---|---|---|---| +| Conservative | 22 | $9,900 | ~$52.7K | +| Realistic | 48 | $21,600 | ~$107.6K | +| Aggressive | 90 | $40,500 | ~$198.9K | + +### 10.5 Breakeven + +- **Monthly operating breakeven:** ~9 tenants (MRR covers $4,000/mo fixed opex + COGS). Every tenant past 9 is ~$443/mo pure contribution. +- **Full build-cost payback:** ~13 months realistic, ~10 aggressive, ~23 conservative. + +### 10.6 Verdict: is it worth building? + +**Yes - build it.** The math is unambiguous because three things compound: + +1. Near-zero marginal cost (~$0.50 per 1,000 queries) on self-hosted infra and open components - no license fees, no per-seat third-party cost. +2. ~95-98% gross margin with premium pricing in a market that already has real floors (Glean ~$60K/yr minimum, Moveworks $100K+). Wall-O at $199-$2,000/mo is dramatically cheaper to the customer yet ~98% margin to us. +3. Owned distribution: ITPP already has the MSP client base and trust, so CAC is ~$1,000/tenant instead of a paid-acquisition crawl. + +Full payback in ~10-13 months realistic, profitable even in the conservative case within ~2 years. The honest caveats: the M365 connector is the riskiest engineering line (fund and pilot it first), DeepSeek may raise prices (LiteLLM abstraction absorbs it), and churn/adoption are unproven in a new category (activation, queries-per-active-user, is the KPI to watch). + +--- + +## 11. The Ask (infrastructure, decisions, success) + +### 11.1 Necessary infrastructure (already spec'd) + +The build requires, and the design docs already specify: + +| Infrastructure | Status | Detail | +|---|---|---| +| netcup host for Wall-O | To provision | Docker host, Caddy TLS edge, one Rocket.Chat workspace + MongoDB per tenant | +| Orchestrator (FastAPI) | To build | Multi-tenant, owns tenancy/agents/scopes/retrieval/LLM | +| Postgres + pgvector | To build | Tenancy tables, chunks + HNSW vector index | +| admin-ai (LiteLLM) | Existing | DeepSeek V4 Pro primary + fallback chain | +| Wasabi S3 | Existing | Sync staging, backups, assets, audit exports | +| Vaultwarden | Existing | All secrets by reference, never inline | +| Hexclave/Stack Auth (app3) | Existing | Customer-facing SaaS auth; Entra OIDC for O365 pilot | +| White-label FOSS build | To build (Phase 1 gate) | fossify script + CI image build, never ship stock EE image | + +### 11.2 Deployment options (standard ITPP framing) + +**Option A - ITPP-INFRA Shared:** runs on the existing netcup estate alongside ITPP operations, backed by the same Wasabi S3 backup pipeline. Lowest cost, fastest start, ITPP manages everything below the app layer. + +**Option B - Dedicated:** dedicated netcup/Hetzner instances with a dedicated S3 bucket, managed by ITPP, for tenants that demand physical separation or for the multi-tenant production fleet at scale. + +Shared responsibility is explicit: ITPP manages everything below the app layer (servers, Docker, TLS, Postgres, backups, monitoring); the customer owns their documents and how they use the assistant. + +### 11.3 Decisions needed + +1. Approve the v1 answer-engine build (Phase 0 proof slice through Phase 1 FOSS white-label). +2. Confirm the pricing tiers after the financial model returns. +3. Confirm beachhead scope for Wall Orthodontics (channel count, white-label depth). +4. Schedule the critical review of this proposal (in progress). + +### 11.4 Success definition + +- v1: a Wall Orthodontics staff member asks a policy question in a channel and gets a cited, correct answer - with zero hallucinations across a week of real use. +- Business: 3 paying tenants with demonstrated ROI (measured time-saved) within 6 months of v1 launch. +- Technical: FOSS white-label build in CI before the first non-pilot customer; no PHI ever ingested; isolation verified by an adversarial test. +