Files
itpp-infrastructure/projects/scirium/04-business-proposal.md
T

29 KiB

Scirium Business Proposal

Status: OPEN (v1 draft, pre-critical-review) Date: 2026-08-16 Author: Sho'Nuff (Conductor) + Scirium marketing team First client: Wall Orthodontics (pilot beachhead) Product scope: business-agnostic internal staff knowledge-base chat


Table of Contents

  1. Executive Summary
  2. Elevator Pitch
  3. Problem Statement (business-agnostic)
  4. Market Analysis (TAM/SAM/SOM, competitive, real-world examples)
  5. Product Overview (v1 definition, architecture, v2 roadmap)
  6. Revenue Model (pricing tiers)
  7. Competitive Advantages (moat)
  8. Go-to-Market Strategy
  9. Risk Analysis (pre-mortem)
  10. Financial Projections (is it worth building)
  11. The Ask (infrastructure, decisions, success)

1. Executive Summary

Scirium is a multi-tenant internal staff knowledge-base chat platform. Every knowledge domain inside a business (employee handbook, billing, IT help, policies, procedures) becomes its own scoped AI chat channel, backed by that business's real documents. Staff ask questions in a chat interface and get grounded, cited answers drawn from the company's own files - never hallucinated, never leaking across domains.

The core problem is universal and business-agnostic: employees waste a documented, measurable fraction of their week hunting for information that already exists somewhere in the organization. McKinsey and Gartner studies put that fraction in the double digits of weekly hours. Enterprise solutions (Glean, Moveworks, Coveo) solve this for large companies at enterprise prices and enterprise procurement complexity. The SMB and vertical-niche middle - dental practices, law firms, restaurants, retail groups, logistics firms - is priced out and under-served.

Scirium's answer is a hard invariant: one channel equals one knowledge domain, one attached AI agent, and one scoped knowledge source. This makes answers provably grounded (every reply carries citations to the source document) and provably isolated (a billing question can never surface a patient file or a neighboring tenant's data). The product is self-hosted, white-label, and business-agnostic by construction - the same core serves a 12-person dental practice and a 200-person law firm.

Financially, Scirium is a high-margin software product running on infrastructure IT Pro Partner already owns: ~$68K build cost, ~$0.50 per 1,000 queries marginal cost, ~95-98% gross margin, ~$1,000 CAC against ~$14K LTV, and full build-cost payback in ~10-13 months (realistic). The v1 is a grounded Q&A assistant. The v2 roadmap adds risk-tiered "do" capabilities (reporting, content generation, integrations) that deepen the moat without ever touching patient/PHI scope.

This proposal defines v1 and v2 concretely, prices it against the market, and makes the fact-based case for building it. It is submitted for critical review.


2. Elevator Pitch

Scirium is the AI coworker that reads your business's actual documents and answers your team's questions with citations - for the price of a streaming subscription, on infrastructure you control, scoped so nothing leaks. It is "a private, cited ChatGPT for your employee handbook, billing rules, and policies" - one tightly-scoped assistant per knowledge domain, white-labeled to your brand.


3. Problem Statement (business-agnostic)

3.1 The universal problem

Every business with more than a handful of employees has the same problem, regardless of industry:

  • New hires take weeks to onboard because institutional knowledge lives in scattered SharePoint libraries, PDFs, and managers' heads.
  • Repetitive questions ("how do I submit a PTO request", "what's our return policy", "who do I escalate a billing dispute to") interrupt managers and senior staff constantly.
  • Policy and procedure changes never propagate - staff operate on outdated rules.
  • The people who know the answer are busy; the answer itself is already written down somewhere, just not findable.

This is not an industry-specific problem. A dental practice has the same shape of problem as a law firm, a restaurant group, a retail chain, a logistics company, or a real estate office. The documents differ; the pain is identical.

3.2 How people solve it today

Current method Time cost Pain level Why it fails
Ask a manager / senior coworker High (interrupts them every time) High Doesn't scale; same questions repeat
Search SharePoint / Google Drive manually Medium (10-20 min per hunt) Medium Fragmented, no ranking, poor recall
Printed binders / shared docs Medium Medium Goes stale immediately
Generic LLM (public ChatGPT) Low High (risk) Hallucinates, no access to internal docs, leaks data
Enterprise knowledge AI (Glean, Moveworks, Coveo) Low Low (but expensive) $30+/seat/mo, enterprise procurement, overkill for SMB

The gap: there is no affordable, self-hosted, white-label answer for the SMB and vertical-niche middle. Enterprise tools solve the problem but are priced and packaged for enterprises. Public LLMs are cheap but ungrounded and unsafe for internal documents. Scirium sits in that gap.

3.3 The precise target user

The buyer is the owner or office manager of an SMB (10 to 200 employees) who:

  • Already uses Microsoft 365 (so documents live in SharePoint/OneDrive),
  • Is frustrated by repetitive questions and slow onboarding,
  • Will not pay enterprise pricing or survive enterprise procurement,
  • Values data control (keeps documents on their own infrastructure),
  • Wants a white-label experience that feels like their brand, not a third-party tool.

The beachhead vertical is orthodontics/dental (Wall Orthodontics), but the product model is explicitly business-agnostic - the same channel primitive serves any document-driven business.

3.4 The gap Scirium fills

Need Public LLM Enterprise AI Scirium
Grounded, cited answers from my docs No Yes Yes
Affordable for SMB Yes (but unsafe) No Yes
Self-hosted / data control No Partial Yes
White-label to my brand No Partial Yes
Channel-scoped (no cross-domain leak) No Partial Yes (hard invariant)
No PHI / patient scope creep No guarantee No guarantee Enforced at ingestion

4. Market Analysis

4.1 Market size (bottom-up, business-agnostic)

Methodology: per-seat SaaS anchored on knowledge-worker count ("every employee could use an internal Q&A assistant"), blended at $10/user/month - well below the enterprise incumbents that charge $30-75 with minimums.

TAM  = 1,000,000,000 global knowledge workers x $10/mo x 12 = $120B/year
SAM  = 1,000,000,000 x 45% (SMB share of workers, SBA 45.9%) = 450M workers
       x $10/mo x 12 = $54B/year
SOM  = $54B x 1% (3-5 yr category-winner)   = $540M/year
       $54B x 0.1% (year 1-2 realistic)     = $54M/year

US-only cross-check (tighter, verifiable): 100M US knowledge workers x 45.9% = ~46M SMB workers x $120/yr = $5.5B US SMB SAM. 1% capture = $55M/yr; 5% = $275M/yr. Even a single vertical slice (e.g. ~1M US small professional-services firms) is a meaningful SOM.

Sources: ~1B global knowledge workers (Schroders); ~100M US knowledge workers (Upwork/BLS/Eurostat); SMBs employ 45.9% of US private-sector workers (SBA 2024).

  • AI in Knowledge Management: $6.7B (2023) to $62.4B (2033), 25% CAGR (Market.us).
  • Knowledge Management Software: $14.56B (2025) to $70.01B (2035), 16.9% CAGR, fastest segment "intelligent chatbots and virtual agents" (MRFR).
  • Enterprises deploying internal AI chatbots report 45-55% reduction in query resolution time (MRFR).
  • SMB AI adoption rose from 5.2% (Jan 2023) to 17.7% (end 2025), with entry cost falling to $20-30/mo (JPMorgan Chase Institute). The demand is rising and the price point is coming to Scirium's band.

4.3 Real-world examples (the fact-based case to build)

These are the proof points that the category is real, funded, and acquirers pay billions for it:

Company Signal Number
Glean Raised $765M, valued $7.2B (Series F, Jun 2025) Passed $100M ARR 2025; claims up to 110 hours saved/user/year
Moveworks Acquired by ServiceNow for $2.85B (Mar 2025) Customers HP, Unilever, Toyota, Marriott; 70,000 hours reclaimed at one automaker; 75,000 hours at a biopharma
Microsoft 365 Copilot 20M+ paid seats (of 450M M365 seats) 75% of knowledge workers already use gen AI at work (MSFT Work Trend Index)
Notion AI $500M annualized revenue (Sept 2025) AI add-on attach rate grew 10-20% to 50%+ in one year

On the "every business has this problem" claim, the independent studies converge regardless of industry or size:

  • McKinsey Global Institute: knowledge workers spend ~19% of time searching/tracking down information (plus 14% communicating).
  • IDC: ~2.5 hours/day, ~30% of the workday, spent searching; 60% of executives say staff could not find what they needed.
  • Deloitte: >25% of time spent searching for information.

These cluster around 20-30% of the workday lost to search - a universal, business-agnostic cost, not an enterprise-only one. That is the budget Scirium attacks.

4.4 Why the SMB/vertical segment is underserved

  1. Enterprise tools price SMBs out on purpose. Glean is $50-75/seat with a 100-seat minimum ($60K/yr floor); M365 Copilot is $30/seat on top of a required M365 license; Moveworks is $100K+/yr; Coveo is opaque enterprise QPM. None serve a 10-30 person practice.
  2. SMB AI budgets are tiny - median ~$28-50/month total (JPMorgan Chase Institute). A $60K/yr enterprise contract is 100x an SMB's entire AI budget.
  3. Only 17.7% of SMBs had adopted AI by end-2025, versus near-universal experimentation in large enterprises - while 91% of AI-using SMBs report revenue impact (Salesforce). The wedge is a cheap, self-hosted, vertical-aware tool - exactly Scirium.

4.5 Competitive landscape (condensed)

Full 10-competitor analysis is in the research appendix. The key fact: the $300-800/month self-hosted band for 10-30 seat SMBs is structurally empty.

Competitor Pricing Target Scirium's gap
Glean $50-75/seat, 100-seat min (~$60K/yr) Enterprise Self-host + white-label + vertical templates + SMB price
Moveworks/ServiceNow $100K+/yr Enterprise IT/HR Not ticket-automation; general internal knowledge
Guru $10-25/seat, 10-seat min Mid-market No self-host, no white-label, no channel scoping
Notion AI ~$10/seat add-on General workspace Q&A only over Notion content; no M365 connector
M365 Copilot $21-30/seat + M365 base M365 tenants Microsoft-cloud lock-in; no self-host/white-label/scoping
Coveo Enterprise QPM Fortune 1000 No SMB tier
Dust $30-150/seat credits AI-operator teams DIY agent-builder, not turnkey
Open WebUI / Onyx / AnythingLLM Free self-host Developers Raw RAG kits - no multi-tenant isolation, no M365 connector, no vertical layer

Two strategic notes from the research:

  1. The two tools an SMB owner already has are M365 Copilot and Notion AI - so Scirium must lead with self-hosted data control + vertical specificity, not generic "AI chat over docs", which those already do.
  2. No competitor is white-label. That is the cleanest MSP/reseller angle - ITPP can sell Scirium under a partner's brand where Glean/Guru/Notion/Copilot cannot.

5. Product Overview

5.1 The core primitive (what makes Scirium different)

Every Scirium channel is exactly three things bound together, and this is a hard invariant:

  1. One knowledge domain (e.g. "Employee Resources", "Billing", "IT Help").
  2. One attached AI agent (a domain-tuned persona).
  3. One scoped knowledge source (one vector namespace over an allowlisted set of documents).

A channel has exactly one agent and exactly one scope. A channel cannot span two domains, and an agent cannot serve two channels. If a business needs a second domain, it creates a second channel, a second agent, and a second scope. This invariant is what makes answers grounded (only one scope's documents feed the answer) and isolated (no cross-domain or cross-tenant leak), and it is enforced in the schema with unique constraints running in both directions.

5.2 Architecture (canonical split)

Layer Responsibility Owns
Rocket.Chat Chat transport only One workspace per tenant (MIT core, EE stripped). Rooms, users, messages. Zero intelligence.
Orchestrator All intelligence Multi-tenant FastAPI service. Tenancy, agents, kb_scope, M365 connector, retrieval, LLM, posting.
Postgres + pgvector State and vectors Tenants, channels, agents, scopes, documents, chunks, messages.
admin-ai (LiteLLM) LLM DeepSeek V4 Pro primary, configured fallback chain.
Wasabi S3 Object storage M365 sync staging, backups, agent assets, audit exports.

The intelligence never lives in the chat layer. Rocket.Chat is a dumb transport; the orchestrator owns everything that thinks. This split is what makes multi-tenancy clean and white-labeling a server-side concern rather than a per-tenant code fork.

5.3 v1 definition (grounded Q&A)

v1 is the answer engine, nothing more. A staff member @mentions the channel's agent (or DMs it), and gets back a grounded, cited answer drawn from that channel's scoped documents.

The v1 message flow (6 steps):

  1. Staff @mentions the agent in a channel (or DMs it).
  2. Rocket.Chat fires a signed webhook to the orchestrator (HMAC-authenticated, timestamp-skew-checked).
  3. Orchestrator resolves tenant -> channel -> agent -> scope in one request context, dedupes on message id.
  4. Retrieval: pgvector semantic search over the channel's namespace, always filtered by tenant_id AND kb_scope_id. Falls back to Microsoft Graph search if the top score is below threshold.
  5. Prompt build: system persona + retrieved chunks labeled [1], [2]... + citation instruction + question. The agent is told to answer only from context and to say "I could not find an answer" rather than guess.
  6. Answer posts back as the agent bot, with inline [n] citations and a Sources footer linking each cited document.

v1 hard behaviors (SETTLED):

  • Never hallucinate: empty/below-threshold retrieval returns "I could not find an answer", with no sources.
  • Every answer carries citations to source documents.
  • Strictly internal staff knowledge. No patient records, no PHI. Ingestion is allowlisted per scope and a content filter flags PHI markers (SSN, MRN, DOB+name) before indexing.
  • Full audit log: every inbound question and outbound answer is a messages row with tokens, model, latency, and citations.

What is already built vs. to build (v1):

Component Status Effort
Product model + data architecture SETTLED (3 design docs committed) Done
Data model (tenants, channels, agents, kb_scopes, documents, chunks, messages, users) Spec'd, not built Build
Orchestrator (FastAPI, tenancy, retrieval, LLM client) Not built Build
Postgres + pgvector schema Spec'd, not built Build
Rocket.Chat workspace + bot integration Phase 0 checklist spec'd, not built Build
M365 connector (Sites.Selected read) Spec'd, not built Build
White-label server rebrand (fossify FOSS build) Phase 1 gate, not started Build
Auth (Hexclave/Stack Auth, Entra OIDC for O365) Partially existing (Hexclave on app3) Integrate

Phase 0 proof slice: stand up one Rocket.Chat workspace for Wall Orthodontics, register a bot, create a channel, attach an agent, index one document, and close the loop on a single grounded answer. The 21-step Phase 0 checklist is spec'd and ready to execute.

5.4 v2 roadmap (planned upgrades)

v2 adds risk-tiered "do" capabilities on top of the v1 answer engine. These are deliberately ordered by write-risk, and none of them ever unlock patient/PHI or autonomous destructive action.

Capability Tier Write risk Auto-approve? Value
Reporting (structured extraction + SQL aggregation + chart render) Reporting Zero new write risk Yes Highest value, zero risk - recommended first v2 ship
Content generation (drafts, summaries, boilerplate) Content generation Drafts only Yes Saves drafting time
Housekeeping (move/rename/delete-to-recycle) Housekeeping Destructive No - propose/approve/audit Keeps knowledge fresh
Integrations (Graph delegated sendMail/calendar/webhooks) Integrations External side effects No - approval Connects Scirium to workflows
Design/brand (template render + image gen) Design/brand Drafts only Yes (template-driven only) Flyers, branded assets

Write rails (SETTLED): destructive actions require propose -> approve -> narrow audited write with before/after and undo path. Harmless writes auto-approve by risk tier. Per-agent tool scoping preserves one-agent-one-domain. Template-driven rendering is required for text-accurate flyers; raw image generation is unreliable for text and is not shipped for that use.

Permanently out of scope (never unlocks): patient records, PHI, HIPAA scope, clinical/x-ray image analysis (separate FDA-regulated product), and autonomous destructive actions. These are hard rails, not feature gaps.

5.5 Business-agnostic by construction

The v1 and v2 product model references no industry. "Tenant" is any business; "channel" is any knowledge domain; "documents" are any document type. The M365 connector reads SharePoint/OneDrive libraries generically. The only orthodontics-specific artifacts are the pilot client's name and branding. A law firm, restaurant group, or logistics company is onboarded by the same provisioning path with different documents, different channel names, and different branding - zero code change.


6. Revenue Model

6.1 Pricing philosophy

Premium positioning, not race-to-bottom. Scirium is "the most accurate, best-scoped internal knowledge assistant", not the cheapest. Price for value delivered (time saved, errors avoided), not competitor undercutting. The product is self-hosted software with high gross margin, so price is set by willingness-to-pay in the segment, not by our marginal cost.

6.2 Tiers (from the financial model)

Tier Price Includes
Starter $199/tenant/mo (up to 10 users, +$15/user beyond) 1 M365 connector (SharePoint + OneDrive), Rocket.Chat, 50K docs, 1 channel scope set
Business $499/tenant/mo (up to 25 users, +$25/user beyond) + Teams connectors, unlimited channel scoping, SSO (OIDC/SAML), 250K docs, SLA support
Enterprise $2,000/tenant/mo (annual contract) Dedicated instance, unlimited seats, custom connectors, SCIM, DPA, custom retention, white-glove onboarding, 99.9% SLA

Blended ARPU at a 60/30/10 tier mix = $469/mo (modeled at $450 for safety). Every tier includes the core grounded Q&A engine, the channel-scoped isolation invariant, citations, audit logs, and the no-PHI guardrail - higher tiers unlock more v2 "do" capabilities and white-glove onboarding, not better core answer quality.

6.3 Revenue streams

  1. Recurring SaaS subscription (primary).
  2. One-time onboarding/setup fee (document library audit, channel configuration).
  3. White-label native app (Phase 2) as a premium add-on.
  4. Optional dedicated-infrastructure deployment (single-tenant netcup/Hetzner instance) for compliance-sensitive buyers.

7. Competitive Advantages (moat)

7.1 The hard invariant is the moat

The one-channel-one-agent-one-scope invariant is not a feature a competitor can bolt on after the fact; it is the schema. It makes every answer provably grounded and provably isolated, and it is what lets Scirium serve regulated-adjacent verticals (dental, legal) where data leakage is a dealbreaker. Enterprise tools do broad search over everything; Scirium does scoped answers over exactly one domain.

7.2 SMB price point with enterprise-grade isolation

Scirium pairs SMB affordability with the isolation discipline usually found only in enterprise platforms. That combination is the open gap in the market.

7.3 Self-hosted and white-label

Data never leaves the customer's estate (or ITPP's controlled netcup estate). Chat, documents, and vectors stay self-hosted; only the LLM call routes through the internal LiteLLM proxy, and by scope it never carries PHI. White-labeling is a server-side rebrand (fossify FOSS build), so every tenant feels like their own branded product, not a resold third-party tool.

7.4 Distribution via existing ITPP MSP

IT Pro Partner already runs the infrastructure and has the MSP relationship model. Scirium is not a cold-start consumer product; it slots into an existing managed-services motion where the first customer (Wall Orthodontics) is already warm.

7.5 The moat compounds in v2

The risk-tiered "do" capabilities (reporting first, then integrations) deepen switching cost once a tenant's staff builds habits and workflows on the assistant. Reporting is the recommended first v2 ship because it is highest value with zero new write risk.


8. Go-to-Market Strategy

8.1 Beachhead: Wall Orthodontics

Prove the loop with one real client in one vertical. Phase 0 closes a single grounded answer; Phase 1 delivers a white-labeled workspace with 3-5 channels (Employee Resources, Billing, IT Help, and a marketing/design channel as a tool-belt extension). Collect concrete ROI (questions answered, time saved, onboarding speed) for the case study.

8.2 Vertical expansion from the beachhead

The orthodontics win becomes a repeatable template for adjacent verticals: general dental, then any document-driven SMB (law firms, accounting, real estate, restaurants, logistics). Each vertical gets a tailored channel template (e.g. "billing" channel for a dental practice vs. "matter intake" for a law firm) but zero core code change - the business-agnostic model pays off here.

8.3 Channel: MSP-led, not direct consumer

Sell through IT Pro Partner's managed-services relationships and peer referral (practice-to-practice, firm-to-firm). The pitch is "the AI that reads your actual documents", demonstrated with a live tenant, not a slide deck. White-label means each MSP or franchise group can offer it under their own brand.

8.4 Motion

  1. Live demo on a real (anonymized) tenant - show a cited answer, not a mockup.
  2. 14-day pilot on the prospect's own SharePoint library.
  3. Convert pilot to subscription with onboarding fee.

9. Risk Analysis (pre-mortem)

Risk Likelihood Impact Mitigation
LLM answer quality (hallucination) Medium High Never-fabricate rail, citation requirement, below-threshold "I could not find an answer", model fallback chain
Data leakage across tenants/channels Low Critical Hard invariant, tenant_id on every row, RLS, vector namespace isolation, HMAC webhook auth
PHI accidentally ingested Low Critical Allowlisted libraries only, PHI-marker pre-ingest filter, block-and-log
M365 Graph API rate limits / connector fragility Medium Medium Delta sync, exponential backoff, Graph search fallback
Rocket.Chat EE licensing in resale Low High fossify FOSS-only build, never ship stock EE image
Native app store cost (Phase 2) Medium Low PWA first, native only after PWA validated
Slow build (scope creep into "do" features) Medium Medium v1 is answer-only; v2 capabilities are risk-tiered and gated
Competition (enterprise tools move downmarket) Medium Medium SMB price point + isolation + white-label + MSP distribution
Churn if ROI not demonstrated Medium High Measure time-saved from day one, report it to the buyer monthly

9.1 The one risk that kills the product

Hallucination. If Scirium ever gives staff a wrong, confidently-stated answer about policy or billing, trust collapses and the product is done. The never-fabricate rail, mandatory citations, and below-threshold "I could not find an answer" behavior are not features - they are the product. This is why the answer engine ships before any "do" capability, and why the v2 write features are risk-tiered with approval gates.


10. Financial Projections (is it worth building)

10.1 Build cost

Assumption: $100/hr blended engineering rate (loaded senior full-stack or mid contractor), no new hardware (runs on existing netcup estate).

Component Hours Note
M365 Graph connector (ACL mapping, delta sync, webhooks) 145 Hardest single line
Ingestion pipeline (parse, chunk, embed, pgvector) 70
RAG retrieval + grounded generation + citations 90
Multi-tenancy (isolation, config, vector namespaces) 50
Rocket.Chat integration 70
Auth + SSO (Stack Auth / Hexclave) 70
Admin UI + billing + onboarding 90
Testing, security, observability, deploy, docs 90
Total 675

Build cost = 675 hrs x $100 = ~$68K (range $45K-$100K). Ongoing engineering ~$3,500/mo post-launch.

10.2 Marginal cost (near zero)

DeepSeek V4 Flash via LiteLLM ($0.14/M in, $0.28/M out). Per answer: ~2,000 input + ~500 output tokens.

LLM       = (0.002M x $0.14) + (0.0005M x $0.28) = $0.00042/query
Embedding + storage + retries (amortized)        = ~$0.00008/query
TOTAL                                            = ~$0.0005/query
                                                = ~$0.50 per 1,000 queries

Sensitivity: even at V4 Pro rates ($0.435/$0.87) it is $1.31/1K queries; a 5x price rise is ~$2.10/1K queries. All negligible. Per-tenant COGS is ~$3-30/mo depending on tier. Gross margin ~95-98%.

10.3 Unit economics

Metric Value
Blended ARPU ~$450/mo (tier-mix math gives $469)
Blended COGS/tenant ~$7/mo
Gross margin ~95-98%
Churn 3%/mo base (5% conservative)
LTV ~$14K (range $8.5K-$21K)
CAC ~$1,000 (owned MSP distribution is the moat)
LTV:CAC ~14:1 (healthy SaaS is >3:1)
Payback per tenant ~2.3 months

10.4 12-month ramp (3 scenarios, churn excluded)

Scenario Tenants (month 12) MRR (month 12) 12-month revenue
Conservative 22 $9,900 ~$52.7K
Realistic 48 $21,600 ~$107.6K
Aggressive 90 $40,500 ~$198.9K

10.5 Breakeven

  • Monthly operating breakeven: ~9 tenants (MRR covers $4,000/mo fixed opex + COGS). Every tenant past 9 is ~$443/mo pure contribution.
  • Full build-cost payback: ~13 months realistic, ~10 aggressive, ~23 conservative.

10.6 Verdict: is it worth building?

Yes - build it. The math is unambiguous because three things compound:

  1. Near-zero marginal cost (~$0.50 per 1,000 queries) on self-hosted infra and open components - no license fees, no per-seat third-party cost.
  2. ~95-98% gross margin with premium pricing in a market that already has real floors (Glean ~$60K/yr minimum, Moveworks $100K+). Scirium at $199-$2,000/mo is dramatically cheaper to the customer yet ~98% margin to us.
  3. Owned distribution: ITPP already has the MSP client base and trust, so CAC is ~$1,000/tenant instead of a paid-acquisition crawl.

Full payback in ~10-13 months realistic, profitable even in the conservative case within ~2 years. The honest caveats: the M365 connector is the riskiest engineering line (fund and pilot it first), DeepSeek may raise prices (LiteLLM abstraction absorbs it), and churn/adoption are unproven in a new category (activation, queries-per-active-user, is the KPI to watch).


11. The Ask (infrastructure, decisions, success)

11.1 Necessary infrastructure (already spec'd)

The build requires, and the design docs already specify:

Infrastructure Status Detail
netcup host for Scirium To provision Docker host, Caddy TLS edge, one Rocket.Chat workspace + MongoDB per tenant
Orchestrator (FastAPI) To build Multi-tenant, owns tenancy/agents/scopes/retrieval/LLM
Postgres + pgvector To build Tenancy tables, chunks + HNSW vector index
admin-ai (LiteLLM) Existing DeepSeek V4 Pro primary + fallback chain
Wasabi S3 Existing Sync staging, backups, assets, audit exports
Vaultwarden Existing All secrets by reference, never inline
Hexclave/Stack Auth (app3) Existing Customer-facing SaaS auth; Entra OIDC for O365 pilot
White-label FOSS build To build (Phase 1 gate) fossify script + CI image build, never ship stock EE image

11.2 Deployment options (standard ITPP framing)

Option A - ITPP-INFRA Shared: runs on the existing netcup estate alongside ITPP operations, backed by the same Wasabi S3 backup pipeline. Lowest cost, fastest start, ITPP manages everything below the app layer.

Option B - Dedicated: dedicated netcup/Hetzner instances with a dedicated S3 bucket, managed by ITPP, for tenants that demand physical separation or for the multi-tenant production fleet at scale.

Shared responsibility is explicit: ITPP manages everything below the app layer (servers, Docker, TLS, Postgres, backups, monitoring); the customer owns their documents and how they use the assistant.

11.3 Decisions needed

  1. Approve the v1 answer-engine build (Phase 0 proof slice through Phase 1 FOSS white-label).
  2. Confirm the pricing tiers after the financial model returns.
  3. Confirm beachhead scope for Wall Orthodontics (channel count, white-label depth).
  4. Schedule the critical review of this proposal (in progress).

11.4 Success definition

  • v1: a Wall Orthodontics staff member asks a policy question in a channel and gets a cited, correct answer - with zero hallucinations across a week of real use.
  • Business: 3 paying tenants with demonstrated ROI (measured time-saved) within 6 months of v1 launch.
  • Technical: FOSS white-label build in CI before the first non-pilot customer; no PHI ever ingested; isolation verified by an adversarial test.