Files
itpp-infrastructure/CHANGELOG.md
T

19 KiB

itpp-infrastructure — CHANGELOG

2026-09-15 — core-bu Warm Standby BUILT and proven (DISARMED; arming is a posture decision)

  • core-bu (159.195.204.203, netcup Nuremberg) is now a working Hermes warm standby. Every step was executed on the box and verified by reading the result back.
  • Hermes parity, proven not assumed: /usr/local/lib/hermes-agent mirrored from Core. core-bu reports v0.21.0 (2026.8.31) · local 8dbf07e9 (+1 carried commit), Python 3.11.15, method git — identical to Core. A real one-shot agent turn returned STANDBY-SMOKE-OK, so the box can serve, not just install.
  • State parity verified against the live box: state.db 140,201,984 bytes on both, quick_check=ok; sessions 151/151; cron jobs.json 89 jobs / 86 enabled on both; 99 references; memories incl. MEMORY.md. The standby store sits one 11-minute checkpoint behind (14,222 vs 14,262 messages) by design — the failover path re-syncs before it serves.
  • Probe path proven: core-bu -> Core over Tailscale (100.71.155.7) returns core / active non-interactively. core-bu joined the tailnet as 100.113.119.108 with RunSSH=false, because Tailscale SSH intercepts non-interactive key auth and would have broken the probe.
  • All three failover action paths exercised under DRYRUN (gate lifted for the test, restored after):
    • health-bad: BAD <t> 1 <t> -> BAD <t> 2 <t+2> -> [dryrun] decision reached: would fence live box and take over (nothing done);
    • host-down: [dryrun] decision reached: host-down path would sync from S3 and start the standby (nothing done; reason: Live host not responding to pings);
    • failback: [dryrun] failback would notify and stop this box's gateway (nothing done), with a decoy process proving DRYRUN does not kill and the real kill command does.
  • Alert channels verified: Telegram getMe -> ok=True shonuff_is_a_bot; SMTP login OK on port 2525.
  • Six defects found and fixed during the build:
    1. No outbound itpp-infra private key on core-bu — the health probe could not reach Core at all. The provisioning record conflates the inbound authorized_keys entry with an outbound key. Installed; fingerprint SHA256:Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ (matches Core).
    2. DRYRUN=1 was not a global no-op (three separate action paths). The host-down path fell through to the S3 sync and the gateway start; the failback path notified and ran its kill. A "safe" dry-run against an unreachable Core would have started a second poller on the same bot token. All paths now guarded.
    3. The failover's final step could not have worked. hermes gateway start requires an installed unit and exits 1 without one. Unit installed with --no-start-now --no-start-on-login (present, disabled, inactive, 0 processes).
    4. SMTP port was wrong for netcup. Shipped SMTP_PORT=587; netcup blocks outbound 25/465/587 and 587 times out from core-bu, so email alerts would have failed silently forever. Now 2525. app1-bu is on Hetzner and works on both, so it was left alone.
    5. The failover sync would clobber the standby's own scripts. Core's ~/.hermes/scripts/ is inside the synced tree, so live/scripts/ holds app1-bu's copies (4,429-byte watchdog vs core-bu's host-specific 19,852-byte one). A wholesale failover sync would have overwritten PROBE_HOST / PROBE_SSH_KEY / the disarm gate at the worst moment. Added --exclude "scripts/hermes-standby-*".
    6. app1-bu — the previously armed standby — had no /root/.config/systemd/user/ at all while its watchdog also calls hermes gateway start. Its failover could not complete its final step either. Unit installed there too; its email path independently verified good on both ports.
  • Disarm gate verified three ways: watchdog logs the notice once per day then exits 0; sync logs DISARMED, sync skipped; boot restore exits before starting anything.
  • Not armed, deliberately. Arming core-bu requires disarming app1-bu first (the single-armed rule is enforced in code by DISARM_FILE), and it trades provider diversity away: Core and core-bu are both netcup, while app1-bu is the only non-netcup box. That is a posture decision for Germaine, not a build step. Stage C (live failover + failback drill, which fences Core's real gateway) also remains, and needs a scheduled window.
  • Bloat note: the wholesale pull brought 2.9 GB of Core's docker/ and 1.1 GB of data/ that core-bu does not run; left in place (472 GB free) but the failover sync should probably exclude them.
  • Full build receipt and evidence: docs/infrastructure/core-bu-standby-package-2026-09-15.md section 9.

2026-09-15 — app4 + core-bu Provisioned (Nuremberg)

  • Renamed by Germaine (recorded at rename time per the changelog mandate): the two netcup Nuremberg boxes ordered 2026-09-14 are now app4 and core-bu.
  • core-bu = v2202609377162521279.megasrv.de, 159.195.204.203/22, IPv6 2a0a:4cc0:c2:b3e0:9409:42ff:fe3b:5397 — netcup RS 2000 G12 (8 vCPU / 15 GB / 503 GB), exact twin of Core. Role: Core's warm standby.
  • app4 = v2202609377162521278.quicksrv.de, 159.195.205.80/22, IPv6 2a0a:4cc0:c2:bcbf:34b3:8cff:fea2:2892 — netcup RS 4000 G12 (12 vCPU / 32 GB / 1007 GB). Role: Core's customer-facing services (DocuSeal, TimeTrex, microbin, Uptime Kuma, Ops Portal backend, Postgres, Redis, customer Caddy routes).
  • Location deviation recorded: both boxes are in Nuremberg (NBG), measured 100.5 ms RTT from Core in Manassas (app2 Manassas = 0.5 ms, app1-bu Ashburn = 1.6 ms). The ITPP ordering standard specifies Manassas. Consequences and mitigations are in the migration plan; benefit is EU/US geographic separation between Core and its standby. Provider diversity is still NOT met (Core, app1-3, app4, core-bu, anita-mnz are all netcup); app1-bu (Hetzner) is the only other provider.
  • Provisioned to standard: Debian 13; ippadmin + NOPASSWD sudo; itpp-infra key for root and ippadmin; ufw active (22/80/443, plus 9100 from Core only); fail2ban; unattended-upgrades; 8 GB swap (9 GB on app4); Docker CE 29.8.0 + Compose v5.5.1; node_exporter; awscli + Wasabi credentials.
  • sshd hardened to fleet convention: PermitRootLogin without-password, PasswordAuthentication no, AllowUsers ippadmin root. Verified four ways per box (root key login, ippadmin key login + sudo -n, password auth refused, non-allowlisted user refused) with the self-reverting lockout guard armed during the change.
  • Backups: root-essentials-backup.sh + cron — app4 04:45 ET, core-bu 05:15 ETs3://hermes-vps-backups/root-backup/{app4,core-bu}/. First run on each tested end to end; script's own download+extract verify passed.
  • Monitoring: a node_exporter job was added to the live Prometheus config (/root/docker/monitoring/prometheus/prometheus.yml) — the job did not exist before. up=1 verified for core, app4, core-bu.
  • Pre-existing finding: node_exporter is not running on app1, app2, app3 or app1-bu, and the node_exporter target list in /opt/prometheus/prometheus.yml sits in an unmounted file that Prometheus never loaded (it still names decommissioned wphost02 and 178.156.131.57). Host metrics for the existing fleet were therefore never collected; tracked as a follow-up, not fixed here.

2026-09-11 — Anita's Hermes Profile Moved to a Dedicated Box (anita-mnz)

  • Infra move: Anita's assistant profile moved off shared Core (152.53.241.111) to a dedicated box anita-mnz 159.195.16.30 (netcup, Manassas VA, 8 vCPU / 15 GB / 503 GB). Cutover 15:53 EDT, ~2 minutes dark, zero messages lost. She keeps the same Telegram bot and chat.
  • Why: repeated state.db corruption on Core under co-tenant memory pressure (4th recurrence, system_prompts then sessions pages). A dedicated box removes the co-tenancy.
  • Method: clean staged base + row-tolerant live-tail graft, so the corrupt live file never transferred. Signatures matched exactly: 99283/99283/166244414/157 (messages/maxid/contentbytes/sessions), integrity_check ok.
  • Post-move state: Core hermes-gateway-anita.service stopped and disabled (no double-poller risk; standby app1-bu carries no anita unit); target gateway active+enabled, telegram connected, 0 getUpdates conflicts.
  • Backup: her own box runs the 3 AM root-essentials backup → s3://hermes-vps-backups/root-backup/anita-mnz/. Full run tested end-to-end 2026-09-11 15:56 (170 MB, upload + download/extract verify OK).
  • ⚠ Correction to the line above (17:35, same day): that archive is config only. root-essentials-backup.sh excludes *.db by design, and nothing on the new box replaced that exclusion, so her conversation store, cron execution DB, notepad, wisdom and verification DBs had zero backup coverage. Fixed the same evening — details below.
  • Core change: hermes-live-sync no longer snapshots the frozen profiles/anita copy (it would advertise a stale db as "live").
  • MCP servers: stripped from her profile (decision 2026-09-11: "She doesn't need access to those"). The mcp_servers block (dre, osint-person, super-search127.0.0.1:8900/8902/8899) was removed from anita/config.yaml, which is why her log had been retry-parking those three every ~5 minutes since the move. Backed up first as config.yaml.bak-mcpstrip-*; 13 top-level keys verified intact. Takes effect on her next gateway restart.
  • Retired / order cancelled: Nuremberg 89.58.44.96 (v2202609377162518632.nicesrv.de). Correction to how this was first written here: it does not "hold nothing". It holds a bare default-profile Hermes install: no state.db, empty sessions/ and cron/, no systemd user units, no gateway process (only node_exporter, containerd, sshd). No unique data, so nothing needs preserving before cancellation. Recorded in /root/.hermes/references/decommissioned-hosts.json.
  • Monitoring fix: health-master-watchdog.py watch-listed hermes-gateway-anita.service as a LOCAL user unit on Core, so it would have alerted forever once that unit was disabled. Local check removed; anita-mnz (159.195.16.30) added to REMOTE_SERVERS and a new REMOTE_USER_UNITS remote user-unit check added (SSH + XDG_RUNTIME_DIR). Verified live: no false alert for her gateway, and her box answers as active.
  • Docs: full incident + pitfalls in /root/.hermes/references/dr-issue-log.md; transferable procedure in the hermes-migration skill.
  • Backup coverage gap (found 17:20, fixed 17:35): the nightly archive on her box was green and 170 MB, but a restore test showed it contains no .db file at all. root-essentials-backup.sh excludes *.db (correct — never tar a live SQLite file), which on Core is backstopped by hermes-backup.sh. The migrated box got the script set without the backstop. Silent failure mode: every indicator said "backed up" while her 974 MB store had no copy anywhere except the frozen Core dir being deleted.
  • Fix: new hermes-db-backup.sh on anita-mnz (/root/hermes-db-backup.sh, chmod 750), cron 10 3 * * *. Per-DB sqlite3 .backup (safe on a live DB), PRAGMA quick_check on every snapshot before it is accepted, one dated tarball, upload to s3://hermes-vps-backups/root-backup/anita-mnz/db/, then downloads the object back, extracts and re-verifies. 14-day retention.
  • Verified: first run 17:04 EDT, 291 MB uploaded; restore test passed — extracted, quick_check=ok, 99,285 messages.
  • Migration completeness proven before deleting the Core copy: memories/MEMORY.md, memories/USER.md and .env are md5-identical between the frozen Core copy and anita-mnz; cron/jobs.json holds the same six job IDs; skills 125 on MNZ vs 124 frozen; store 99,285 vs 99,283 messages. config.yaml differs only by the deliberate MCP strip.
  • Incident report written: docs/incidents/2026-09-11-core-state-db-corruption.md documents the store corruption (header destroyed at 12:49:22), the recovery that built a new store from the clean 01:00 archive and grafted 985 newer messages into it (108,573 messages, max id 322,530, 242 sessions, quick_check and integrity_check both ok), the eight day backup monitor false positive streak, and the prevention list (prune the store, snapshot instead of tar, restore-test every archive, integrity check the newest snapshot, record who restarts the gateway).
  • Frozen Core copy removed (17:41): /root/.hermes/profiles/anita (8.0 GB) was 7.9 GB of corrupt-DB corpses (corrupt-20260903/09/10, pre-rebuild-20260910, recovered-20260910). Archived to s3://hermes-vps-backups/decommissioned/anita-core-frozen-profile-20260911-1732.tar.gz (2,786,082,561 B, 21,206 entries), verified by downloading the object back and matching sha256 b457fa7f0cb6f546f36e53a338e90b7b940f51cf29d5a73d0cceb97f49c9b3dc against the local tarball, then deleted. Nothing unique was destroyed: memories, .env and the six cron job IDs were identical on anita-mnz, which carries one more skill and two more messages. Core disk 125 GB → 117 GB used; /root/.hermes/profiles is now empty. Also removes that copy from Core's nightly archive, which is why Core's backup had grown to 2.97 GB.
  • Stale S3 copy purged (17:41): s3://hermes-vps-backups/live/profiles/anita/ held 20,551 objects / 5.0 GB that the then-running 15-minute sync had pushed before it was paused on Sep 3, including a stale-looking state.db that would have been advertised as "live". Superseded by the archive above and removed; live/profiles/ is now empty.
  • Backup monitor checked (17:38): its single CRITICAL was hermes-live-sync: DISABLED/PAUSED, which is the pause from Sep 3 and explains the monitor's 8-day exit-1 streak (it exits 1 only on CRITICAL, never on warnings). Its three WARNINGs are false positives, verified by content rather than size: Wazuh manager tarballs are sha256-identical for 3 days (static config, job ran 03:15 today), LiteLLM's flagged object is a 333 B config yaml that never changes (the real backup is a 37.9 MB litellm-backup-*.tar.gz from 03:30 today, and there are 47 of them), and the voipsimplicity dump is the same 6,834,636 B each day but a different sha256 each day (valid gzip, 73 tables). The size-uniqueness heuristic cannot tell static-but-fine from stalled; the three entries should be reclassified as "unchanged content" rather than SUSPICIOUS.
  • Same-day correction on job status: "Germaine inbox watch" is enabled=false on anita-mnz. It was enabled=false in the frozen Core copy too — the migration did not disable it.

2026-08-17 - Scirium v2 Proposal Deployed to /scirium/

  • v2 proposal deployed to proposals.itpropartner.com/scirium/ (index.html + 04-business-proposal-v2.md + critical-review.html). Assembled from 4 parallel remediation teams (marketing, technical, financial, legal), SOM reconciled with Financial as authority, build cost corrected to ~$167K (was $68K), verdict: GO with conditions (churn gate, acquisition-maturation gate, trademark clearance, Phase 0 DLP spike). Status: DRAFT FOR REVIEW pending Germaine review before VerdictTank resubmission.
  • URL rename completed: v1 (codename Wall-O) frozen at proposals.itpropartner.com/wall-o/ with a SUPERSEDED banner pointing to /scirium/. v2 is live at /scirium/. This closes the pending item from the 2026-08-16 changelog entry.
  • Sources: v1 at projects/scirium/04-business-proposal.md; v2 at projects/scirium/04-business-proposal-v2.md; team remediation sections under /tmp/scirium-v2/output/ (not repo-bound).

2026-08-16 — Wall-O Renamed to Scirium

  • Product renamed Wall-O → Scirium (coined from Latin "scire" = to know). Applies going forward; "Wall-O" retired to internal codename history only.
  • Domain: scirium.com selected. .com/.io/.ai/.co/.app all available (RDAP 404 + empty NS cross-check). Trademarkia: 0 results for "scirium".
  • Cloudflare at-cost pricing (verified 2026-08-16): .com $10.44/yr, .io $50/yr (renewal ~$51.75), .ai $70/yr (min 2-year term = $140; rising to $80/yr on 2026-03-05), .co $15 first yr / $30 renewal, .app $14.20/yr, .dev $10.18/yr.
  • Source folder moved projects/wall-o/projects/scirium/. Legacy 4 docs still carry "Wall-O" internally; rebranded by the docs team as part of the v1/v2 documentation package.
  • Deployed proposal URL (proposals.itpropartner.com/wall-o/) unchanged pending redeploy under /scirium/.

2026-08-12 — app3 Web Docroot Migration to Per-Site Users

  • Change: every app3 nginx vhost moved off the shared /home/ippadmin/htdocs/ root to a per-site dedicated Linux user with docroot /home/<site-user>/htdocs/<domain> (security hardening — no more single-owner web tree).
  • Verified mappings (live nginx configs, 2026-08-14): mockups → /home/mockups, proposals → /home/proposals, docs → /home/docs, support → /home/support, my.verdicttank.com → /home/myverdicttank, verdicttank.com → /home/gmb, my.transitpin.com → /home/transitpin-dash.
  • Consequence: 10+ skills and their reference/script files still referenced the old /home/ippadmin/htdocs/ paths, causing a wrong-tree deploy on 2026-08-14. Remediated across SKILL.md, references/, and scripts/ (28 files, incl. singular mockup/proposal domain typos).
  • Rule: always read /etc/nginx/sites-enabled/<domain>.conf to confirm the real docroot before deploying. Never assume ippadmin owns a site's files.

2026-08-08 — Hexclave Renamed → Stack Auth

  • Hexclave renamed to Stack Auth. Now running at auth2.itpropartner.com on app3.
  • This is the same service (customer-facing authentication), same server, same Docker stack — only the name changed.
  • Old references to "Hexclave" in scripts, docs, and backups should be updated to "Stack Auth" / stack-auth.
  • Rule going forward: any rename of critical infrastructure gets a changelog entry at the time of the rename, not discovered later.

2026-08-06 — Fallback Chain Overhaul & Two-Key Strategy

  • Root cause: Aug 5 admin-ai budget cap + 4 dead fallback legs = $45 Anthropic burn in 10 hours
  • Rotated all 5 fallback provider keys (new keys for deepseek, google, xai, anthropic, openai)
  • Added F5: gpt-4.1-nano via OpenAI direct (independent infrastructure)
  • Fixed F3: grok-4.6 → grok-4.5 (grok-4.6 never existed — LiteLLM catalog ghost)
  • Documented two-key strategy: operational keys (admin-ai only) vs fallback keys (direct, daily-capped)
  • Added to operational chain: claude-haiku-4-5 (lightweight), grok-4.5 (auditor 2), deepseek-v4-flash (batch)
  • Synced Anita profile with identical fallback chain + provider keys
  • Admin-ai budget raised: $20 → $30/day
  • Updated: model-chain.md, operational-models.md

2026-07-16 — Audit Remediation

  • Created CHANGELOG.md (missing per project documentation standard)
  • Project directory: /root/projects/itpp-infrastructure