## 2026-09-15 (later) - core-bu ARMED, app1-bu disarmed Single-armed rule enforced in the safe order: app1-bu disarmed FIRST, then core-bu armed, so there was never a window with two armed standbys. - **app1-bu disarmed:** its two standby crons (`*/5` watchdog, `*/10` sync) commented with marker `#DISARMED-2026-09-15 (armed node is now core-bu; re-arm: remove this marker)`. Its unrelated 04:37 S3-restorability cron left intact. - **core-bu armed:** `/root/.hermes/standby.disarmed` removed (disarm gate gone). Re-disarm by recreating that file; a copy is at `/root/standby.disarmed.rearm-backup`. Verified after the swap, not assumed: - core-bu real non-dry-run watchdog cycle probed Core: `HEALTH OK unit=active pid=100998 msgs=14262`, then stayed dormant. Zero gateway processes, user unit inactive/disabled. - A real armed sync ran with the gate open and the standby scripts survived untouched (watchdog sha256 prefix unchanged, `63bea8b7557f634f`), so the clobber guard holds under live conditions, not just by reasoning. - app1-bu: both standby crons commented, non-standby crons intact, no gateway process, unit inactive/disabled. **Finding on app1-bu (recorded while disarming):** it has NO boot-time restore unit. `systemctl is-enabled hermes-standby.service` returns `not-found`, so a reboot of app1-bu would never have restored or failed over. Its two crons were the entire failover mechanism. Re-arming app1-bu requires installing that unit first. **Trade-off accepted:** both DR nodes (Core and core-bu) are now netcup. Provider diversity is gone until the US replacement lands. app1-bu remains the only non-netcup node but is no longer armed. **Operational note:** with core-bu armed, a Core Hermes outage lasting 2 consecutive 5-minute cycles (about 10 minutes) fences Core and triggers takeover. Planned gateway work expected to exceed about 10 minutes will cause a failover, then an automatic failback when Core returns. Re-disarm core-bu (`touch /root/.hermes/standby.disarmed`) before any long maintenance window. --- # itpp-infrastructure — CHANGELOG ## 2026-09-15 — core-bu Warm Standby BUILT and proven (armed later the same day, see the entry above) - **core-bu (`159.195.204.203`, netcup Nuremberg) is now a working Hermes warm standby.** Every step was executed on the box and verified by reading the result back. - **Hermes parity, proven not assumed:** `/usr/local/lib/hermes-agent` mirrored from Core. core-bu reports `v0.21.0 (2026.8.31) · local 8dbf07e9 (+1 carried commit)`, Python `3.11.15`, method `git` — identical to Core. A real one-shot agent turn returned `STANDBY-SMOKE-OK`, so the box can serve, not just install. - **State parity verified against the live box:** `state.db` **140,201,984 bytes on both**, `quick_check=ok`; `sessions` 151/151; cron `jobs.json` **89 jobs / 86 enabled on both**; 99 references; memories incl. `MEMORY.md`. The standby store sits one 11-minute checkpoint behind (14,222 vs 14,262 messages) by design — the failover path re-syncs before it serves. - **Probe path proven:** core-bu -> Core over Tailscale (`100.71.155.7`) returns `core` / `active` non-interactively. core-bu joined the tailnet as `100.113.119.108` with `RunSSH=false`, because Tailscale SSH intercepts non-interactive key auth and would have broken the probe. - **All three failover action paths exercised under DRYRUN (gate lifted for the test, restored after):** - health-bad: `BAD 1 ` -> `BAD 2 ` -> `[dryrun] decision reached: would fence live box and take over (nothing done)`; - host-down: `[dryrun] decision reached: host-down path would sync from S3 and start the standby (nothing done; reason: Live host not responding to pings)`; - failback: `[dryrun] failback would notify and stop this box's gateway (nothing done)`, with a decoy process proving DRYRUN does not kill and the real kill command does. - **Alert channels verified:** Telegram `getMe` -> `ok=True shonuff_is_a_bot`; SMTP login OK on port **2525**. - **Six defects found and fixed during the build:** 1. **No outbound `itpp-infra` private key on core-bu** — the health probe could not reach Core at all. The provisioning record conflates the *inbound* `authorized_keys` entry with an outbound key. Installed; fingerprint `SHA256:Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ` (matches Core). 2. **`DRYRUN=1` was not a global no-op** (three separate action paths). The host-down path fell through to the S3 sync and the gateway start; the failback path notified and ran its kill. A "safe" dry-run against an unreachable Core would have started a **second poller on the same bot token**. All paths now guarded. 3. **The failover's final step could not have worked.** `hermes gateway start` requires an installed unit and exits 1 without one. Unit installed with `--no-start-now --no-start-on-login` (present, `disabled`, `inactive`, 0 processes). 4. **SMTP port was wrong for netcup.** Shipped `SMTP_PORT=587`; netcup blocks outbound 25/465/587 and 587 **times out** from core-bu, so email alerts would have failed silently forever. Now `2525`. app1-bu is on Hetzner and works on both, so it was left alone. 5. **The failover sync would clobber the standby's own scripts.** Core's `~/.hermes/scripts/` is inside the synced tree, so `live/scripts/` holds **app1-bu's** copies (4,429-byte watchdog vs core-bu's host-specific 19,852-byte one). A wholesale failover sync would have overwritten `PROBE_HOST` / `PROBE_SSH_KEY` / the disarm gate at the worst moment. Added `--exclude "scripts/hermes-standby-*"`. 6. **`app1-bu` — the previously armed standby — had no `/root/.config/systemd/user/` at all** while its watchdog also calls `hermes gateway start`. **Its failover could not complete its final step either.** Unit installed there too; its email path independently verified good on both ports. - **Disarm gate verified three ways:** watchdog logs the notice once per day then exits 0; sync logs `DISARMED, sync skipped`; boot restore exits before starting anything. - **Not armed, deliberately.** Arming core-bu requires disarming app1-bu first (the single-armed rule is enforced in code by `DISARM_FILE`), and it **trades provider diversity away**: Core and core-bu are both netcup, while app1-bu is the only non-netcup box. That is a posture decision for Germaine, not a build step. Stage C (live failover + failback drill, which fences Core's real gateway) also remains, and needs a scheduled window. - **Bloat note:** the wholesale pull brought 2.9 GB of Core's `docker/` and 1.1 GB of `data/` that core-bu does not run; left in place (472 GB free) but the failover sync should probably exclude them. - Full build receipt and evidence: `docs/infrastructure/core-bu-standby-package-2026-09-15.md` section 9. 2026-09-15 — app4 + core-bu Provisioned (Nuremberg) - **Renamed by Germaine (recorded at rename time per the changelog mandate):** the two netcup Nuremberg boxes ordered 2026-09-14 are now **app4** and **core-bu**. - **core-bu** = `v2202609377162521279.megasrv.de`, `159.195.204.203/22`, IPv6 `2a0a:4cc0:c2:b3e0:9409:42ff:fe3b:5397` — netcup **RS 2000 G12** (8 vCPU / 15 GB / 503 GB), exact twin of Core. Role: **Core's warm standby**. - **app4** = `v2202609377162521278.quicksrv.de`, `159.195.205.80/22`, IPv6 `2a0a:4cc0:c2:bcbf:34b3:8cff:fea2:2892` — netcup **RS 4000 G12** (12 vCPU / 32 GB / 1007 GB). Role: **Core's customer-facing services** (DocuSeal, TimeTrex, microbin, Uptime Kuma, Ops Portal backend, Postgres, Redis, customer Caddy routes). - **Location deviation recorded:** both boxes are in **Nuremberg (NBG)**, measured **100.5 ms RTT** from Core in Manassas (app2 Manassas = 0.5 ms, app1-bu Ashburn = 1.6 ms). The ITPP ordering standard specifies Manassas. Consequences and mitigations are in the migration plan; benefit is EU/US geographic separation between Core and its standby. **Provider diversity is still NOT met** (Core, app1-3, app4, core-bu, anita-mnz are all netcup); app1-bu (Hetzner) is the only other provider. - **Provisioned to standard:** Debian 13; `ippadmin` + NOPASSWD sudo; `itpp-infra` key for root and ippadmin; ufw active (22/80/443, plus 9100 from Core only); fail2ban; unattended-upgrades; 8 GB swap (9 GB on app4); Docker CE 29.8.0 + Compose v5.5.1; node_exporter; awscli + Wasabi credentials. - **sshd hardened to fleet convention:** `PermitRootLogin without-password`, `PasswordAuthentication no`, `AllowUsers ippadmin root`. Verified four ways per box (root key login, ippadmin key login + `sudo -n`, password auth refused, non-allowlisted user refused) with the self-reverting lockout guard armed during the change. - **Backups:** `root-essentials-backup.sh` + cron — app4 **04:45 ET**, core-bu **05:15 ET** → `s3://hermes-vps-backups/root-backup/{app4,core-bu}/`. First run on each tested end to end; script's own download+extract verify passed. - **Monitoring:** a `node_exporter` job was **added to the live Prometheus config** (`/root/docker/monitoring/prometheus/prometheus.yml`) — the job did not exist before. `up=1` verified for core, app4, core-bu. - **Pre-existing finding:** node_exporter is **not running** on app1, app2, app3 or app1-bu, and the node_exporter target list in `/opt/prometheus/prometheus.yml` sits in an **unmounted file that Prometheus never loaded** (it still names decommissioned `wphost02` and `178.156.131.57`). Host metrics for the existing fleet were therefore never collected; tracked as a follow-up, not fixed here. ## 2026-09-11 — Anita's Hermes Profile Moved to a Dedicated Box (anita-mnz) - **Infra move:** Anita's assistant profile moved off shared Core (`152.53.241.111`) to a dedicated box **anita-mnz `159.195.16.30`** (netcup, Manassas VA, 8 vCPU / 15 GB / 503 GB). Cutover 15:53 EDT, ~2 minutes dark, **zero messages lost**. She keeps the same Telegram bot and chat. - **Why:** repeated `state.db` corruption on Core under co-tenant memory pressure (4th recurrence, `system_prompts` then `sessions` pages). A dedicated box removes the co-tenancy. - **Method:** clean staged base + row-tolerant live-tail graft, so the corrupt live file never transferred. Signatures matched exactly: `99283/99283/166244414/157` (messages/maxid/contentbytes/sessions), `integrity_check ok`. - **Post-move state:** Core `hermes-gateway-anita.service` stopped and **disabled** (no double-poller risk; standby app1-bu carries no anita unit); target gateway `active`+`enabled`, telegram `connected`, 0 `getUpdates` conflicts. - **Backup:** her own box runs the 3 AM root-essentials backup → `s3://hermes-vps-backups/root-backup/anita-mnz/`. Full run tested end-to-end 2026-09-11 15:56 (170 MB, upload + download/extract verify OK). - **⚠ Correction to the line above (17:35, same day):** that archive is config only. `root-essentials-backup.sh` excludes `*.db` by design, and nothing on the new box replaced that exclusion, so her conversation store, cron execution DB, notepad, wisdom and verification DBs had **zero** backup coverage. Fixed the same evening — details below. - **Core change:** `hermes-live-sync` no longer snapshots the frozen `profiles/anita` copy (it would advertise a stale db as "live"). - **MCP servers:** stripped from her profile (decision 2026-09-11: "She doesn't need access to those"). The `mcp_servers` block (`dre`, `osint-person`, `super-search` → `127.0.0.1:8900/8902/8899`) was removed from `anita/config.yaml`, which is why her log had been retry-parking those three every ~5 minutes since the move. Backed up first as `config.yaml.bak-mcpstrip-*`; 13 top-level keys verified intact. Takes effect on her next gateway restart. - **Retired / order cancelled:** Nuremberg `89.58.44.96` (`v2202609377162518632.nicesrv.de`). **Correction to how this was first written here:** it does not "hold nothing". It holds a bare default-profile Hermes install: no `state.db`, empty `sessions/` and `cron/`, no systemd user units, no gateway process (only node_exporter, containerd, sshd). No unique data, so nothing needs preserving before cancellation. Recorded in `/root/.hermes/references/decommissioned-hosts.json`. - **Monitoring fix:** `health-master-watchdog.py` watch-listed `hermes-gateway-anita.service` as a LOCAL user unit on Core, so it would have alerted forever once that unit was disabled. Local check removed; `anita-mnz` (`159.195.16.30`) added to `REMOTE_SERVERS` and a new `REMOTE_USER_UNITS` remote user-unit check added (SSH + `XDG_RUNTIME_DIR`). Verified live: no false alert for her gateway, and her box answers as active. - **Docs:** full incident + pitfalls in `/root/.hermes/references/dr-issue-log.md`; transferable procedure in the `hermes-migration` skill. - **Backup coverage gap (found 17:20, fixed 17:35):** the nightly archive on her box was green and 170 MB, but a restore test showed it contains no `.db` file at all. `root-essentials-backup.sh` excludes `*.db` (correct — never `tar` a live SQLite file), which on Core is backstopped by `hermes-backup.sh`. The migrated box got the script set without the backstop. **Silent failure mode**: every indicator said "backed up" while her 974 MB store had no copy anywhere except the frozen Core dir being deleted. - **Fix:** new `hermes-db-backup.sh` on `anita-mnz` (`/root/hermes-db-backup.sh`, chmod 750), cron `10 3 * * *`. Per-DB `sqlite3 .backup` (safe on a live DB), `PRAGMA quick_check` on every snapshot before it is accepted, one dated tarball, upload to `s3://hermes-vps-backups/root-backup/anita-mnz/db/`, then **downloads the object back, extracts and re-verifies**. 14-day retention. - **Verified:** first run 17:04 EDT, 291 MB uploaded; restore test passed — extracted, `quick_check=ok`, 99,285 messages. - **Migration completeness proven before deleting the Core copy:** `memories/MEMORY.md`, `memories/USER.md` and `.env` are md5-identical between the frozen Core copy and `anita-mnz`; `cron/jobs.json` holds the same six job IDs; skills 125 on MNZ vs 124 frozen; store 99,285 vs 99,283 messages. `config.yaml` differs only by the deliberate MCP strip. - **Incident report written:** `docs/incidents/2026-09-11-core-state-db-corruption.md` documents the store corruption (header destroyed at 12:49:22), the recovery that built a new store from the clean 01:00 archive and grafted 985 newer messages into it (108,573 messages, max id 322,530, 242 sessions, `quick_check` and `integrity_check` both `ok`), the eight day backup monitor false positive streak, and the prevention list (prune the store, snapshot instead of tar, restore-test every archive, integrity check the newest snapshot, record who restarts the gateway). - **Frozen Core copy removed (17:41):** `/root/.hermes/profiles/anita` (8.0 GB) was 7.9 GB of corrupt-DB corpses (`corrupt-20260903/09/10`, `pre-rebuild-20260910`, `recovered-20260910`). Archived to `s3://hermes-vps-backups/decommissioned/anita-core-frozen-profile-20260911-1732.tar.gz` (2,786,082,561 B, 21,206 entries), **verified by downloading the object back and matching sha256** `b457fa7f0cb6f546f36e53a338e90b7b940f51cf29d5a73d0cceb97f49c9b3dc` against the local tarball, then deleted. Nothing unique was destroyed: memories, `.env` and the six cron job IDs were identical on `anita-mnz`, which carries one more skill and two more messages. Core disk 125 GB → 117 GB used; `/root/.hermes/profiles` is now empty. Also removes that copy from Core's nightly archive, which is why Core's backup had grown to 2.97 GB. - **Stale S3 copy purged (17:41):** `s3://hermes-vps-backups/live/profiles/anita/` held 20,551 objects / 5.0 GB that the then-running 15-minute sync had pushed before it was paused on Sep 3, including a stale-looking `state.db` that would have been advertised as "live". Superseded by the archive above and removed; `live/profiles/` is now empty. - **Backup monitor checked (17:38):** its single CRITICAL was `hermes-live-sync: DISABLED/PAUSED`, which is the pause from Sep 3 and explains the monitor's 8-day exit-1 streak (it exits 1 only on CRITICAL, never on warnings). Its three WARNINGs are false positives, verified by content rather than size: Wazuh manager tarballs are sha256-identical for 3 days (static config, job ran 03:15 today), LiteLLM's flagged object is a 333 B config yaml that never changes (the real backup is a 37.9 MB `litellm-backup-*.tar.gz` from 03:30 today, and there are 47 of them), and the voipsimplicity dump is the same 6,834,636 B each day but a **different** sha256 each day (valid gzip, 73 tables). The size-uniqueness heuristic cannot tell static-but-fine from stalled; the three entries should be reclassified as "unchanged content" rather than SUSPICIOUS. - **Same-day correction on job status:** "Germaine inbox watch" is `enabled=false` on `anita-mnz`. It was `enabled=false` in the frozen Core copy too — the migration did not disable it. ## 2026-08-17 - Scirium v2 Proposal Deployed to /scirium/ - **v2 proposal deployed** to `proposals.itpropartner.com/scirium/` (index.html + 04-business-proposal-v2.md + critical-review.html). Assembled from 4 parallel remediation teams (marketing, technical, financial, legal), SOM reconciled with Financial as authority, build cost corrected to ~$167K (was $68K), verdict: GO with conditions (churn gate, acquisition-maturation gate, trademark clearance, Phase 0 DLP spike). Status: DRAFT FOR REVIEW pending Germaine review before VerdictTank resubmission. - **URL rename completed:** v1 (codename Wall-O) frozen at `proposals.itpropartner.com/wall-o/` with a SUPERSEDED banner pointing to `/scirium/`. v2 is live at `/scirium/`. This closes the pending item from the 2026-08-16 changelog entry. - **Sources:** v1 at `projects/scirium/04-business-proposal.md`; v2 at `projects/scirium/04-business-proposal-v2.md`; team remediation sections under `/tmp/scirium-v2/output/` (not repo-bound). ## 2026-08-16 — Wall-O Renamed to Scirium - **Product renamed Wall-O → Scirium** (coined from Latin "scire" = to know). Applies going forward; "Wall-O" retired to internal codename history only. - **Domain:** `scirium.com` selected. `.com`/`.io`/`.ai`/`.co`/`.app` all available (RDAP 404 + empty NS cross-check). Trademarkia: 0 results for "scirium". - **Cloudflare at-cost pricing (verified 2026-08-16):** `.com` $10.44/yr, `.io` $50/yr (renewal ~$51.75), `.ai` $70/yr (min 2-year term = $140; rising to $80/yr on 2026-03-05), `.co` $15 first yr / $30 renewal, `.app` $14.20/yr, `.dev` $10.18/yr. - **Source folder moved** `projects/wall-o/` → `projects/scirium/`. Legacy 4 docs still carry "Wall-O" internally; rebranded by the docs team as part of the v1/v2 documentation package. - **Deployed proposal URL** (`proposals.itpropartner.com/wall-o/`) unchanged pending redeploy under `/scirium/`. ## 2026-08-12 — app3 Web Docroot Migration to Per-Site Users - **Change:** every app3 nginx vhost moved off the shared `/home/ippadmin/htdocs/` root to a per-site dedicated Linux user with docroot `/home//htdocs/` (security hardening — no more single-owner web tree). - **Verified mappings (live nginx configs, 2026-08-14):** mockups → `/home/mockups`, proposals → `/home/proposals`, docs → `/home/docs`, support → `/home/support`, my.verdicttank.com → `/home/myverdicttank`, verdicttank.com → `/home/gmb`, my.transitpin.com → `/home/transitpin-dash`. - **Consequence:** 10+ skills and their reference/script files still referenced the old `/home/ippadmin/htdocs/` paths, causing a wrong-tree deploy on 2026-08-14. Remediated across SKILL.md, references/, and scripts/ (28 files, incl. singular `mockup`/`proposal` domain typos). - **Rule:** always read `/etc/nginx/sites-enabled/.conf` to confirm the real docroot before deploying. Never assume `ippadmin` owns a site's files. ## 2026-08-08 — Hexclave Renamed → Stack Auth - **Hexclave** renamed to **Stack Auth**. Now running at `auth2.itpropartner.com` on app3. - This is the same service (customer-facing authentication), same server, same Docker stack — only the name changed. - Old references to "Hexclave" in scripts, docs, and backups should be updated to "Stack Auth" / `stack-auth`. - **Rule going forward:** any rename of critical infrastructure gets a changelog entry at the time of the rename, not discovered later. ## 2026-08-06 — Fallback Chain Overhaul & Two-Key Strategy - **Root cause:** Aug 5 admin-ai budget cap + 4 dead fallback legs = $45 Anthropic burn in 10 hours - Rotated all 5 fallback provider keys (new keys for deepseek, google, xai, anthropic, openai) - Added F5: gpt-4.1-nano via OpenAI direct (independent infrastructure) - Fixed F3: grok-4.6 → grok-4.5 (grok-4.6 never existed — LiteLLM catalog ghost) - Documented two-key strategy: operational keys (admin-ai only) vs fallback keys (direct, daily-capped) - Added to operational chain: claude-haiku-4-5 (lightweight), grok-4.5 (auditor 2), deepseek-v4-flash (batch) - Synced Anita profile with identical fallback chain + provider keys - Admin-ai budget raised: $20 → $30/day - Updated: model-chain.md, operational-models.md ## 2026-07-16 — Audit Remediation - Created CHANGELOG.md (missing per project documentation standard) - Project directory: `/root/projects/itpp-infrastructure`