Files

12 KiB

itpp-infrastructure — CHANGELOG

2026-09-11 — Anita's Hermes Profile Moved to a Dedicated Box (anita-mnz)

  • Infra move: Anita's assistant profile moved off shared Core (152.53.241.111) to a dedicated box anita-mnz 159.195.16.30 (netcup, Manassas VA, 8 vCPU / 15 GB / 503 GB). Cutover 15:53 EDT, ~2 minutes dark, zero messages lost. She keeps the same Telegram bot and chat.
  • Why: repeated state.db corruption on Core under co-tenant memory pressure (4th recurrence, system_prompts then sessions pages). A dedicated box removes the co-tenancy.
  • Method: clean staged base + row-tolerant live-tail graft, so the corrupt live file never transferred. Signatures matched exactly: 99283/99283/166244414/157 (messages/maxid/contentbytes/sessions), integrity_check ok.
  • Post-move state: Core hermes-gateway-anita.service stopped and disabled (no double-poller risk; standby app1-bu carries no anita unit); target gateway active+enabled, telegram connected, 0 getUpdates conflicts.
  • Backup: her own box runs the 3 AM root-essentials backup → s3://hermes-vps-backups/root-backup/anita-mnz/. Full run tested end-to-end 2026-09-11 15:56 (170 MB, upload + download/extract verify OK).
  • ⚠ Correction to the line above (17:35, same day): that archive is config only. root-essentials-backup.sh excludes *.db by design, and nothing on the new box replaced that exclusion, so her conversation store, cron execution DB, notepad, wisdom and verification DBs had zero backup coverage. Fixed the same evening — details below.
  • Core change: hermes-live-sync no longer snapshots the frozen profiles/anita copy (it would advertise a stale db as "live").
  • MCP servers: stripped from her profile (decision 2026-09-11: "She doesn't need access to those"). The mcp_servers block (dre, osint-person, super-search127.0.0.1:8900/8902/8899) was removed from anita/config.yaml, which is why her log had been retry-parking those three every ~5 minutes since the move. Backed up first as config.yaml.bak-mcpstrip-*; 13 top-level keys verified intact. Takes effect on her next gateway restart.
  • Retired / order cancelled: Nuremberg 89.58.44.96 (v2202609377162518632.nicesrv.de). Correction to how this was first written here: it does not "hold nothing". It holds a bare default-profile Hermes install: no state.db, empty sessions/ and cron/, no systemd user units, no gateway process (only node_exporter, containerd, sshd). No unique data, so nothing needs preserving before cancellation. Recorded in /root/.hermes/references/decommissioned-hosts.json.
  • Monitoring fix: health-master-watchdog.py watch-listed hermes-gateway-anita.service as a LOCAL user unit on Core, so it would have alerted forever once that unit was disabled. Local check removed; anita-mnz (159.195.16.30) added to REMOTE_SERVERS and a new REMOTE_USER_UNITS remote user-unit check added (SSH + XDG_RUNTIME_DIR). Verified live: no false alert for her gateway, and her box answers as active.
  • Docs: full incident + pitfalls in /root/.hermes/references/dr-issue-log.md; transferable procedure in the hermes-migration skill.
  • Backup coverage gap (found 17:20, fixed 17:35): the nightly archive on her box was green and 170 MB, but a restore test showed it contains no .db file at all. root-essentials-backup.sh excludes *.db (correct — never tar a live SQLite file), which on Core is backstopped by hermes-backup.sh. The migrated box got the script set without the backstop. Silent failure mode: every indicator said "backed up" while her 974 MB store had no copy anywhere except the frozen Core dir being deleted.
  • Fix: new hermes-db-backup.sh on anita-mnz (/root/hermes-db-backup.sh, chmod 750), cron 10 3 * * *. Per-DB sqlite3 .backup (safe on a live DB), PRAGMA quick_check on every snapshot before it is accepted, one dated tarball, upload to s3://hermes-vps-backups/root-backup/anita-mnz/db/, then downloads the object back, extracts and re-verifies. 14-day retention.
  • Verified: first run 17:04 EDT, 291 MB uploaded; restore test passed — extracted, quick_check=ok, 99,285 messages.
  • Migration completeness proven before deleting the Core copy: memories/MEMORY.md, memories/USER.md and .env are md5-identical between the frozen Core copy and anita-mnz; cron/jobs.json holds the same six job IDs; skills 125 on MNZ vs 124 frozen; store 99,285 vs 99,283 messages. config.yaml differs only by the deliberate MCP strip.
  • Incident report written: docs/incidents/2026-09-11-core-state-db-corruption.md documents the store corruption (header destroyed at 12:49:22), the recovery that built a new store from the clean 01:00 archive and grafted 985 newer messages into it (108,573 messages, max id 322,530, 242 sessions, quick_check and integrity_check both ok), the eight day backup monitor false positive streak, and the prevention list (prune the store, snapshot instead of tar, restore-test every archive, integrity check the newest snapshot, record who restarts the gateway).
  • Frozen Core copy removed (17:41): /root/.hermes/profiles/anita (8.0 GB) was 7.9 GB of corrupt-DB corpses (corrupt-20260903/09/10, pre-rebuild-20260910, recovered-20260910). Archived to s3://hermes-vps-backups/decommissioned/anita-core-frozen-profile-20260911-1732.tar.gz (2,786,082,561 B, 21,206 entries), verified by downloading the object back and matching sha256 b457fa7f0cb6f546f36e53a338e90b7b940f51cf29d5a73d0cceb97f49c9b3dc against the local tarball, then deleted. Nothing unique was destroyed: memories, .env and the six cron job IDs were identical on anita-mnz, which carries one more skill and two more messages. Core disk 125 GB → 117 GB used; /root/.hermes/profiles is now empty. Also removes that copy from Core's nightly archive, which is why Core's backup had grown to 2.97 GB.
  • Stale S3 copy purged (17:41): s3://hermes-vps-backups/live/profiles/anita/ held 20,551 objects / 5.0 GB that the then-running 15-minute sync had pushed before it was paused on Sep 3, including a stale-looking state.db that would have been advertised as "live". Superseded by the archive above and removed; live/profiles/ is now empty.
  • Backup monitor checked (17:38): its single CRITICAL was hermes-live-sync: DISABLED/PAUSED, which is the pause from Sep 3 and explains the monitor's 8-day exit-1 streak (it exits 1 only on CRITICAL, never on warnings). Its three WARNINGs are false positives, verified by content rather than size: Wazuh manager tarballs are sha256-identical for 3 days (static config, job ran 03:15 today), LiteLLM's flagged object is a 333 B config yaml that never changes (the real backup is a 37.9 MB litellm-backup-*.tar.gz from 03:30 today, and there are 47 of them), and the voipsimplicity dump is the same 6,834,636 B each day but a different sha256 each day (valid gzip, 73 tables). The size-uniqueness heuristic cannot tell static-but-fine from stalled; the three entries should be reclassified as "unchanged content" rather than SUSPICIOUS.
  • Same-day correction on job status: "Germaine inbox watch" is enabled=false on anita-mnz. It was enabled=false in the frozen Core copy too — the migration did not disable it.

2026-08-17 - Scirium v2 Proposal Deployed to /scirium/

  • v2 proposal deployed to proposals.itpropartner.com/scirium/ (index.html + 04-business-proposal-v2.md + critical-review.html). Assembled from 4 parallel remediation teams (marketing, technical, financial, legal), SOM reconciled with Financial as authority, build cost corrected to ~$167K (was $68K), verdict: GO with conditions (churn gate, acquisition-maturation gate, trademark clearance, Phase 0 DLP spike). Status: DRAFT FOR REVIEW pending Germaine review before VerdictTank resubmission.
  • URL rename completed: v1 (codename Wall-O) frozen at proposals.itpropartner.com/wall-o/ with a SUPERSEDED banner pointing to /scirium/. v2 is live at /scirium/. This closes the pending item from the 2026-08-16 changelog entry.
  • Sources: v1 at projects/scirium/04-business-proposal.md; v2 at projects/scirium/04-business-proposal-v2.md; team remediation sections under /tmp/scirium-v2/output/ (not repo-bound).

2026-08-16 — Wall-O Renamed to Scirium

  • Product renamed Wall-O → Scirium (coined from Latin "scire" = to know). Applies going forward; "Wall-O" retired to internal codename history only.
  • Domain: scirium.com selected. .com/.io/.ai/.co/.app all available (RDAP 404 + empty NS cross-check). Trademarkia: 0 results for "scirium".
  • Cloudflare at-cost pricing (verified 2026-08-16): .com $10.44/yr, .io $50/yr (renewal ~$51.75), .ai $70/yr (min 2-year term = $140; rising to $80/yr on 2026-03-05), .co $15 first yr / $30 renewal, .app $14.20/yr, .dev $10.18/yr.
  • Source folder moved projects/wall-o/projects/scirium/. Legacy 4 docs still carry "Wall-O" internally; rebranded by the docs team as part of the v1/v2 documentation package.
  • Deployed proposal URL (proposals.itpropartner.com/wall-o/) unchanged pending redeploy under /scirium/.

2026-08-12 — app3 Web Docroot Migration to Per-Site Users

  • Change: every app3 nginx vhost moved off the shared /home/ippadmin/htdocs/ root to a per-site dedicated Linux user with docroot /home/<site-user>/htdocs/<domain> (security hardening — no more single-owner web tree).
  • Verified mappings (live nginx configs, 2026-08-14): mockups → /home/mockups, proposals → /home/proposals, docs → /home/docs, support → /home/support, my.verdicttank.com → /home/myverdicttank, verdicttank.com → /home/gmb, my.transitpin.com → /home/transitpin-dash.
  • Consequence: 10+ skills and their reference/script files still referenced the old /home/ippadmin/htdocs/ paths, causing a wrong-tree deploy on 2026-08-14. Remediated across SKILL.md, references/, and scripts/ (28 files, incl. singular mockup/proposal domain typos).
  • Rule: always read /etc/nginx/sites-enabled/<domain>.conf to confirm the real docroot before deploying. Never assume ippadmin owns a site's files.

2026-08-08 — Hexclave Renamed → Stack Auth

  • Hexclave renamed to Stack Auth. Now running at auth2.itpropartner.com on app3.
  • This is the same service (customer-facing authentication), same server, same Docker stack — only the name changed.
  • Old references to "Hexclave" in scripts, docs, and backups should be updated to "Stack Auth" / stack-auth.
  • Rule going forward: any rename of critical infrastructure gets a changelog entry at the time of the rename, not discovered later.

2026-08-06 — Fallback Chain Overhaul & Two-Key Strategy

  • Root cause: Aug 5 admin-ai budget cap + 4 dead fallback legs = $45 Anthropic burn in 10 hours
  • Rotated all 5 fallback provider keys (new keys for deepseek, google, xai, anthropic, openai)
  • Added F5: gpt-4.1-nano via OpenAI direct (independent infrastructure)
  • Fixed F3: grok-4.6 → grok-4.5 (grok-4.6 never existed — LiteLLM catalog ghost)
  • Documented two-key strategy: operational keys (admin-ai only) vs fallback keys (direct, daily-capped)
  • Added to operational chain: claude-haiku-4-5 (lightweight), grok-4.5 (auditor 2), deepseek-v4-flash (batch)
  • Synced Anita profile with identical fallback chain + provider keys
  • Admin-ai budget raised: $20 → $30/day
  • Updated: model-chain.md, operational-models.md

2026-07-16 — Audit Remediation

  • Created CHANGELOG.md (missing per project documentation standard)
  • Project directory: /root/projects/itpp-infrastructure