25 KiB
2026-09-15 (later) - core-bu ARMED, app1-bu disarmed
Single-armed rule enforced in the safe order: app1-bu disarmed FIRST, then core-bu armed, so there was never a window with two armed standbys.
- app1-bu disarmed: its two standby crons (
*/5watchdog,*/10sync) commented with marker#DISARMED-2026-09-15 (armed node is now core-bu; re-arm: remove this marker). Its unrelated 04:37 S3-restorability cron left intact. - core-bu armed:
/root/.hermes/standby.disarmedremoved (disarm gate gone). Re-disarm by recreating that file; a copy is at/root/standby.disarmed.rearm-backup.
Verified after the swap, not assumed:
- core-bu real non-dry-run watchdog cycle probed Core:
HEALTH OK unit=active pid=100998 msgs=14262, then stayed dormant. Zero gateway processes, user unit inactive/disabled. - A real armed sync ran with the gate open and the standby scripts survived untouched (watchdog sha256 prefix unchanged,
63bea8b7557f634f), so the clobber guard holds under live conditions, not just by reasoning. - app1-bu: both standby crons commented, non-standby crons intact, no gateway process, unit inactive/disabled.
Finding on app1-bu (recorded while disarming): it has NO boot-time restore unit. systemctl is-enabled hermes-standby.service returns not-found, so a reboot of app1-bu would never have restored or failed over. Its two crons were the entire failover mechanism. Re-arming app1-bu requires installing that unit first.
Trade-off accepted: both DR nodes (Core and core-bu) are now netcup. Provider diversity is gone until the US replacement lands. app1-bu remains the only non-netcup node but is no longer armed.
Sync fixes found only after arming (the arm is what exposed them):
- Two
aws s3 syncinstances were running concurrently. The first full pull took longer than the 10-minute cron interval, so every tick started another one. Fixed with anflocksingle-instance lock. Proven: a second launch logsAnother sync is already running (lock held), exiting. - The sync pulled
live/cron/as a whole prefix, which means LISTING 190k+ objects undercron/output/(Core's job-log history, 192,716 files on disk at last count). Even with--exclude, the listing is the cost, and that listing is what blew past the interval. Fixed by dropping cron from the prefix sync and copying onlycron/jobs.jsonandcron/executions.db. - The sync log had reached 30 MB of per-object download lines, and aws's partial progress lines merged into the following log line. That corrupted an audit read: a filename containing "error" (e.g.
errors.py) looked like a real error, and a merged line made a single run look like it downloaded 17,927 objects. Switched to--only-show-errors.
Result: a full cycle now completes in 13 seconds with a 9-line log and zero cron/output files touched (verified by mtime), versus 10+ minutes and overlapping runs before. Job definitions still refresh: cron/jobs.json advanced from 12:49:37 to 13:19:40 on the test run.
Operational note: with core-bu armed, a Core Hermes outage lasting 2 consecutive 5-minute cycles (about 10 minutes) fences Core and triggers takeover. Planned gateway work expected to exceed about 10 minutes will cause a failover, then an automatic failback when Core returns. Re-disarm core-bu (touch /root/.hermes/standby.disarmed) before any long maintenance window.
itpp-infrastructure — CHANGELOG
2026-09-15 (third) — Core's stale job output: root cause, cleanup, allow-list fix
Root cause. hermes-live-sync.sh mirrors ~/.hermes/ to s3://hermes-vps-backups/live/ every 11 minutes and had no cron/output exclusion, while aws s3 sync only ever adds and updates, never deletes. So every cron job-run output Core has ever produced is still in S3 (July 5 onward), long after Core's own 50-runs-per-job retention pruned it locally. Measured 2026-09-15: 270,262 objects / 624 MB under live/cron/output/ against 16,718 files on Core. The cost was never the storage, it was the listing: every restore and every standby sync had to enumerate 270k objects, which is what pushed the standby sync past its own 10-minute interval and made two run concurrently.
Cleanup
- Core local
~/.hermes/cron/output: 88 MB / 16,718 files -> 36 MB / 3,921 files. Removed three output dirs belonging to jobs that no longer exist injobs.json(0137c3505b72,677fc4199392,77b99f6e2ff6; 12,733 files, output stopped mid-July) plus 112 stale flat-file run logs older than 30 days. All eight live state files verified md5-identical before and after. - core-bu: deleted the frozen 1.2 GB / 194,819-file pulled copy. A standby never reads
cron/outputand the sync no longer pulls it. - S3
live/cron/output/: 270,262 objects / 624 MB -> 8 objects / 74 KB. Deleted 136 per-job prefixes and 112 loose run logs. No history lost: thehermes-full-backuptarballs (70 of them, back to 2026-07-04) includecron/output.
Fix — allow-list, not a wildcard exclude. Applied to Core hermes-live-sync.sh, hermes-standby-restore.sh (Core archival copy, core-bu's live copy, S3 standby/), and core-bu hermes-standby-sync.sh:
--exclude "cron/output/*"
--include "cron/output/osint-discovery.json" (plus 7 more, named exactly, per script)
A wildcard include such as --include "cron/output/*.md" does not work: * also matches /, so it re-admits files nested inside job subdirs. This was proven on a synthetic tree before deploying. The eight allow-listed files are the shared job state jobs read for dedupe (so a takeover does not re-alert everything); run output has no failover value.
Verification. After the prune, the 3,921 job-output files still existed locally but were gone from S3, so any failure to exclude would visibly re-upload them. A full live-sync then ran to completion and live/cron/output/ still held exactly 8 objects. Rollback: hermes-live-sync.sh.bak-20260915-cronout, hermes-standby-restore.sh.bak-20260915-cronout, hermes-standby-sync.sh.bak2-20260915 on core-bu.
2026-09-15 — core-bu Warm Standby BUILT and proven (armed later the same day, see the entry above)
- core-bu (
159.195.204.203, netcup Nuremberg) is now a working Hermes warm standby. Every step was executed on the box and verified by reading the result back. - Hermes parity, proven not assumed:
/usr/local/lib/hermes-agentmirrored from Core. core-bu reportsv0.21.0 (2026.8.31) · local 8dbf07e9 (+1 carried commit), Python3.11.15, methodgit— identical to Core. A real one-shot agent turn returnedSTANDBY-SMOKE-OK, so the box can serve, not just install. - State parity verified against the live box:
state.db140,201,984 bytes on both,quick_check=ok;sessions151/151; cronjobs.json89 jobs / 86 enabled on both; 99 references; memories incl.MEMORY.md. The standby store sits one 11-minute checkpoint behind (14,222 vs 14,262 messages) by design — the failover path re-syncs before it serves. - Probe path proven: core-bu -> Core over Tailscale (
100.71.155.7) returnscore/activenon-interactively. core-bu joined the tailnet as100.113.119.108withRunSSH=false, because Tailscale SSH intercepts non-interactive key auth and would have broken the probe. - All three failover action paths exercised under DRYRUN (gate lifted for the test, restored after):
- health-bad:
BAD <t> 1 <t>->BAD <t> 2 <t+2>->[dryrun] decision reached: would fence live box and take over (nothing done); - host-down:
[dryrun] decision reached: host-down path would sync from S3 and start the standby (nothing done; reason: Live host not responding to pings); - failback:
[dryrun] failback would notify and stop this box's gateway (nothing done), with a decoy process proving DRYRUN does not kill and the real kill command does.
- health-bad:
- Alert channels verified: Telegram
getMe->ok=True shonuff_is_a_bot; SMTP login OK on port 2525. - Six defects found and fixed during the build:
- No outbound
itpp-infraprivate key on core-bu — the health probe could not reach Core at all. The provisioning record conflates the inboundauthorized_keysentry with an outbound key. Installed; fingerprintSHA256:Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ(matches Core). DRYRUN=1was not a global no-op (three separate action paths). The host-down path fell through to the S3 sync and the gateway start; the failback path notified and ran its kill. A "safe" dry-run against an unreachable Core would have started a second poller on the same bot token. All paths now guarded.- The failover's final step could not have worked.
hermes gateway startrequires an installed unit and exits 1 without one. Unit installed with--no-start-now --no-start-on-login(present,disabled,inactive, 0 processes). - SMTP port was wrong for netcup. Shipped
SMTP_PORT=587; netcup blocks outbound 25/465/587 and 587 times out from core-bu, so email alerts would have failed silently forever. Now2525. app1-bu is on Hetzner and works on both, so it was left alone. - The failover sync would clobber the standby's own scripts. Core's
~/.hermes/scripts/is inside the synced tree, solive/scripts/holds app1-bu's copies (4,429-byte watchdog vs core-bu's host-specific 19,852-byte one). A wholesale failover sync would have overwrittenPROBE_HOST/PROBE_SSH_KEY/ the disarm gate at the worst moment. Added--exclude "scripts/hermes-standby-*". app1-bu— the previously armed standby — had no/root/.config/systemd/user/at all while its watchdog also callshermes gateway start. Its failover could not complete its final step either. Unit installed there too; its email path independently verified good on both ports.
- No outbound
- Disarm gate verified three ways: watchdog logs the notice once per day then exits 0; sync logs
DISARMED, sync skipped; boot restore exits before starting anything. - Not armed, deliberately. Arming core-bu requires disarming app1-bu first (the single-armed rule is enforced in code by
DISARM_FILE), and it trades provider diversity away: Core and core-bu are both netcup, while app1-bu is the only non-netcup box. That is a posture decision for Germaine, not a build step. Stage C (live failover + failback drill, which fences Core's real gateway) also remains, and needs a scheduled window. - Bloat note: the wholesale pull brought 2.9 GB of Core's
docker/and 1.1 GB ofdata/that core-bu does not run; left in place (472 GB free) but the failover sync should probably exclude them. - Full build receipt and evidence:
docs/infrastructure/core-bu-standby-package-2026-09-15.mdsection 9.
2026-09-15 — app4 + core-bu Provisioned (Nuremberg)
- Renamed by Germaine (recorded at rename time per the changelog mandate): the two netcup Nuremberg boxes ordered 2026-09-14 are now app4 and core-bu.
- core-bu =
v2202609377162521279.megasrv.de,159.195.204.203/22, IPv62a0a:4cc0:c2:b3e0:9409:42ff:fe3b:5397— netcup RS 2000 G12 (8 vCPU / 15 GB / 503 GB), exact twin of Core. Role: Core's warm standby. - app4 =
v2202609377162521278.quicksrv.de,159.195.205.80/22, IPv62a0a:4cc0:c2:bcbf:34b3:8cff:fea2:2892— netcup RS 4000 G12 (12 vCPU / 32 GB / 1007 GB). Role: Core's customer-facing services (DocuSeal, TimeTrex, microbin, Uptime Kuma, Ops Portal backend, Postgres, Redis, customer Caddy routes). - Location deviation recorded: both boxes are in Nuremberg (NBG), measured 100.5 ms RTT from Core in Manassas (app2 Manassas = 0.5 ms, app1-bu Ashburn = 1.6 ms). The ITPP ordering standard specifies Manassas. Consequences and mitigations are in the migration plan; benefit is EU/US geographic separation between Core and its standby. Provider diversity is still NOT met (Core, app1-3, app4, core-bu, anita-mnz are all netcup); app1-bu (Hetzner) is the only other provider.
- Provisioned to standard: Debian 13;
ippadmin+ NOPASSWD sudo;itpp-infrakey for root and ippadmin; ufw active (22/80/443, plus 9100 from Core only); fail2ban; unattended-upgrades; 8 GB swap (9 GB on app4); Docker CE 29.8.0 + Compose v5.5.1; node_exporter; awscli + Wasabi credentials. - sshd hardened to fleet convention:
PermitRootLogin without-password,PasswordAuthentication no,AllowUsers ippadmin root. Verified four ways per box (root key login, ippadmin key login +sudo -n, password auth refused, non-allowlisted user refused) with the self-reverting lockout guard armed during the change. - Backups:
root-essentials-backup.sh+ cron — app4 04:45 ET, core-bu 05:15 ET →s3://hermes-vps-backups/root-backup/{app4,core-bu}/. First run on each tested end to end; script's own download+extract verify passed. - Monitoring: a
node_exporterjob was added to the live Prometheus config (/root/docker/monitoring/prometheus/prometheus.yml) — the job did not exist before.up=1verified for core, app4, core-bu. - Pre-existing finding: node_exporter is not running on app1, app2, app3 or app1-bu, and the node_exporter target list in
/opt/prometheus/prometheus.ymlsits in an unmounted file that Prometheus never loaded (it still names decommissionedwphost02and178.156.131.57). Host metrics for the existing fleet were therefore never collected; tracked as a follow-up, not fixed here.
2026-09-11 — Anita's Hermes Profile Moved to a Dedicated Box (anita-mnz)
- Infra move: Anita's assistant profile moved off shared Core (
152.53.241.111) to a dedicated box anita-mnz159.195.16.30(netcup, Manassas VA, 8 vCPU / 15 GB / 503 GB). Cutover 15:53 EDT, ~2 minutes dark, zero messages lost. She keeps the same Telegram bot and chat. - Why: repeated
state.dbcorruption on Core under co-tenant memory pressure (4th recurrence,system_promptsthensessionspages). A dedicated box removes the co-tenancy. - Method: clean staged base + row-tolerant live-tail graft, so the corrupt live file never transferred. Signatures matched exactly:
99283/99283/166244414/157(messages/maxid/contentbytes/sessions),integrity_check ok. - Post-move state: Core
hermes-gateway-anita.servicestopped and disabled (no double-poller risk; standby app1-bu carries no anita unit); target gatewayactive+enabled, telegramconnected, 0getUpdatesconflicts. - Backup: her own box runs the 3 AM root-essentials backup →
s3://hermes-vps-backups/root-backup/anita-mnz/. Full run tested end-to-end 2026-09-11 15:56 (170 MB, upload + download/extract verify OK). - ⚠ Correction to the line above (17:35, same day): that archive is config only.
root-essentials-backup.shexcludes*.dbby design, and nothing on the new box replaced that exclusion, so her conversation store, cron execution DB, notepad, wisdom and verification DBs had zero backup coverage. Fixed the same evening — details below. - Core change:
hermes-live-syncno longer snapshots the frozenprofiles/anitacopy (it would advertise a stale db as "live"). - MCP servers: stripped from her profile (decision 2026-09-11: "She doesn't need access to those"). The
mcp_serversblock (dre,osint-person,super-search→127.0.0.1:8900/8902/8899) was removed fromanita/config.yaml, which is why her log had been retry-parking those three every ~5 minutes since the move. Backed up first asconfig.yaml.bak-mcpstrip-*; 13 top-level keys verified intact. Takes effect on her next gateway restart. - Retired / order cancelled: Nuremberg
89.58.44.96(v2202609377162518632.nicesrv.de). Correction to how this was first written here: it does not "hold nothing". It holds a bare default-profile Hermes install: nostate.db, emptysessions/andcron/, no systemd user units, no gateway process (only node_exporter, containerd, sshd). No unique data, so nothing needs preserving before cancellation. Recorded in/root/.hermes/references/decommissioned-hosts.json. - Monitoring fix:
health-master-watchdog.pywatch-listedhermes-gateway-anita.serviceas a LOCAL user unit on Core, so it would have alerted forever once that unit was disabled. Local check removed;anita-mnz(159.195.16.30) added toREMOTE_SERVERSand a newREMOTE_USER_UNITSremote user-unit check added (SSH +XDG_RUNTIME_DIR). Verified live: no false alert for her gateway, and her box answers as active. - Docs: full incident + pitfalls in
/root/.hermes/references/dr-issue-log.md; transferable procedure in thehermes-migrationskill. - Backup coverage gap (found 17:20, fixed 17:35): the nightly archive on her box was green and 170 MB, but a restore test showed it contains no
.dbfile at all.root-essentials-backup.shexcludes*.db(correct — nevertara live SQLite file), which on Core is backstopped byhermes-backup.sh. The migrated box got the script set without the backstop. Silent failure mode: every indicator said "backed up" while her 974 MB store had no copy anywhere except the frozen Core dir being deleted. - Fix: new
hermes-db-backup.shonanita-mnz(/root/hermes-db-backup.sh, chmod 750), cron10 3 * * *. Per-DBsqlite3 .backup(safe on a live DB),PRAGMA quick_checkon every snapshot before it is accepted, one dated tarball, upload tos3://hermes-vps-backups/root-backup/anita-mnz/db/, then downloads the object back, extracts and re-verifies. 14-day retention. - Verified: first run 17:04 EDT, 291 MB uploaded; restore test passed — extracted,
quick_check=ok, 99,285 messages. - Migration completeness proven before deleting the Core copy:
memories/MEMORY.md,memories/USER.mdand.envare md5-identical between the frozen Core copy andanita-mnz;cron/jobs.jsonholds the same six job IDs; skills 125 on MNZ vs 124 frozen; store 99,285 vs 99,283 messages.config.yamldiffers only by the deliberate MCP strip. - Incident report written:
docs/incidents/2026-09-11-core-state-db-corruption.mddocuments the store corruption (header destroyed at 12:49:22), the recovery that built a new store from the clean 01:00 archive and grafted 985 newer messages into it (108,573 messages, max id 322,530, 242 sessions,quick_checkandintegrity_checkbothok), the eight day backup monitor false positive streak, and the prevention list (prune the store, snapshot instead of tar, restore-test every archive, integrity check the newest snapshot, record who restarts the gateway). - Frozen Core copy removed (17:41):
/root/.hermes/profiles/anita(8.0 GB) was 7.9 GB of corrupt-DB corpses (corrupt-20260903/09/10,pre-rebuild-20260910,recovered-20260910). Archived tos3://hermes-vps-backups/decommissioned/anita-core-frozen-profile-20260911-1732.tar.gz(2,786,082,561 B, 21,206 entries), verified by downloading the object back and matching sha256b457fa7f0cb6f546f36e53a338e90b7b940f51cf29d5a73d0cceb97f49c9b3dcagainst the local tarball, then deleted. Nothing unique was destroyed: memories,.envand the six cron job IDs were identical onanita-mnz, which carries one more skill and two more messages. Core disk 125 GB → 117 GB used;/root/.hermes/profilesis now empty. Also removes that copy from Core's nightly archive, which is why Core's backup had grown to 2.97 GB. - Stale S3 copy purged (17:41):
s3://hermes-vps-backups/live/profiles/anita/held 20,551 objects / 5.0 GB that the then-running 15-minute sync had pushed before it was paused on Sep 3, including a stale-lookingstate.dbthat would have been advertised as "live". Superseded by the archive above and removed;live/profiles/is now empty. - Backup monitor checked (17:38): its single CRITICAL was
hermes-live-sync: DISABLED/PAUSED, which is the pause from Sep 3 and explains the monitor's 8-day exit-1 streak (it exits 1 only on CRITICAL, never on warnings). Its three WARNINGs are false positives, verified by content rather than size: Wazuh manager tarballs are sha256-identical for 3 days (static config, job ran 03:15 today), LiteLLM's flagged object is a 333 B config yaml that never changes (the real backup is a 37.9 MBlitellm-backup-*.tar.gzfrom 03:30 today, and there are 47 of them), and the voipsimplicity dump is the same 6,834,636 B each day but a different sha256 each day (valid gzip, 73 tables). The size-uniqueness heuristic cannot tell static-but-fine from stalled; the three entries should be reclassified as "unchanged content" rather than SUSPICIOUS. - Same-day correction on job status: "Germaine inbox watch" is
enabled=falseonanita-mnz. It wasenabled=falsein the frozen Core copy too — the migration did not disable it.
2026-08-17 - Scirium v2 Proposal Deployed to /scirium/
- v2 proposal deployed to
proposals.itpropartner.com/scirium/(index.html + 04-business-proposal-v2.md + critical-review.html). Assembled from 4 parallel remediation teams (marketing, technical, financial, legal), SOM reconciled with Financial as authority, build cost corrected to ~$167K (was $68K), verdict: GO with conditions (churn gate, acquisition-maturation gate, trademark clearance, Phase 0 DLP spike). Status: DRAFT FOR REVIEW pending Germaine review before VerdictTank resubmission. - URL rename completed: v1 (codename Wall-O) frozen at
proposals.itpropartner.com/wall-o/with a SUPERSEDED banner pointing to/scirium/. v2 is live at/scirium/. This closes the pending item from the 2026-08-16 changelog entry. - Sources: v1 at
projects/scirium/04-business-proposal.md; v2 atprojects/scirium/04-business-proposal-v2.md; team remediation sections under/tmp/scirium-v2/output/(not repo-bound).
2026-08-16 — Wall-O Renamed to Scirium
- Product renamed Wall-O → Scirium (coined from Latin "scire" = to know). Applies going forward; "Wall-O" retired to internal codename history only.
- Domain:
scirium.comselected..com/.io/.ai/.co/.appall available (RDAP 404 + empty NS cross-check). Trademarkia: 0 results for "scirium". - Cloudflare at-cost pricing (verified 2026-08-16):
.com$10.44/yr,.io$50/yr (renewal ~$51.75),.ai$70/yr (min 2-year term = $140; rising to $80/yr on 2026-03-05),.co$15 first yr / $30 renewal,.app$14.20/yr,.dev$10.18/yr. - Source folder moved
projects/wall-o/→projects/scirium/. Legacy 4 docs still carry "Wall-O" internally; rebranded by the docs team as part of the v1/v2 documentation package. - Deployed proposal URL (
proposals.itpropartner.com/wall-o/) unchanged pending redeploy under/scirium/.
2026-08-12 — app3 Web Docroot Migration to Per-Site Users
- Change: every app3 nginx vhost moved off the shared
/home/ippadmin/htdocs/root to a per-site dedicated Linux user with docroot/home/<site-user>/htdocs/<domain>(security hardening — no more single-owner web tree). - Verified mappings (live nginx configs, 2026-08-14): mockups →
/home/mockups, proposals →/home/proposals, docs →/home/docs, support →/home/support, my.verdicttank.com →/home/myverdicttank, verdicttank.com →/home/gmb, my.transitpin.com →/home/transitpin-dash. - Consequence: 10+ skills and their reference/script files still referenced the old
/home/ippadmin/htdocs/paths, causing a wrong-tree deploy on 2026-08-14. Remediated across SKILL.md, references/, and scripts/ (28 files, incl. singularmockup/proposaldomain typos). - Rule: always read
/etc/nginx/sites-enabled/<domain>.confto confirm the real docroot before deploying. Never assumeippadminowns a site's files.
2026-08-08 — Hexclave Renamed → Stack Auth
- Hexclave renamed to Stack Auth. Now running at
auth2.itpropartner.comon app3. - This is the same service (customer-facing authentication), same server, same Docker stack — only the name changed.
- Old references to "Hexclave" in scripts, docs, and backups should be updated to "Stack Auth" /
stack-auth. - Rule going forward: any rename of critical infrastructure gets a changelog entry at the time of the rename, not discovered later.
2026-08-06 — Fallback Chain Overhaul & Two-Key Strategy
- Root cause: Aug 5 admin-ai budget cap + 4 dead fallback legs = $45 Anthropic burn in 10 hours
- Rotated all 5 fallback provider keys (new keys for deepseek, google, xai, anthropic, openai)
- Added F5: gpt-4.1-nano via OpenAI direct (independent infrastructure)
- Fixed F3: grok-4.6 → grok-4.5 (grok-4.6 never existed — LiteLLM catalog ghost)
- Documented two-key strategy: operational keys (admin-ai only) vs fallback keys (direct, daily-capped)
- Added to operational chain: claude-haiku-4-5 (lightweight), grok-4.5 (auditor 2), deepseek-v4-flash (batch)
- Synced Anita profile with identical fallback chain + provider keys
- Admin-ai budget raised: $20 → $30/day
- Updated: model-chain.md, operational-models.md
2026-07-16 — Audit Remediation
- Created CHANGELOG.md (missing per project documentation standard)
- Project directory:
/root/projects/itpp-infrastructure