- migration-plan-app4-core-bu-2026-09-15.md: 8-phase plan (Nuremberg decision, provider-diversity gap, acceptance criteria, rollback, DNS/Caddy checklist) - core-service-inventory-2026-09-15: verified Core inventory, ~30 customer-facing services (the Aug 15 plan listed 5), 3 DocuSeal instances, TimeTrex Postgres, dead Caddy routes - reference-update-matrix-2026-09-15: 52 artifacts that name a host - fix naming collision: 6 files called the Hetzner box core-bu, the name core-bu now claims; app1-bu = 5.161.225.131, core-bu = 159.195.204.203 (netcup Nuremberg) - correct the false provider-diversity claim (the standby is now netcup too) - supersede app4-migration-plan.md (wrong region reported, silent on core-bu)
43 KiB
Reference Update Matrix — app4 / core-bu Go-Live and app1-bu Retirement
Date: 2026-09-15
Author: Sho'Nuff (read-only audit subagent, ran on Core)
Scope: Every document, script, config, monitoring check, and record found on Core
that names a host/service and must change when app4 (159.195.205.80) and the new
netcup core-bu (159.195.204.203) go live, and when the Hetzner standby app1-bu
(5.161.225.131) is eventually retired.
Method: Live grep/read of /root/projects/itpp-infrastructure, /root/.hermes/scripts,
/root/.hermes/references, /root/.hermes/cron/jobs.json, /root/.hermes/config.yaml,
/etc/caddy/Caddyfile, /etc/cron.d, root crontab, systemd units, /opt/ops-portal,
and read-only Gitea API calls to git.itpropartner.com. No file was edited. No service
was restarted. Secrets are masked below where shown.
Read-only constraint honored: all evidence below is quoted from files actually read
on 2026-09-15. Line numbers refer to the file state at read time.
IMPORTANT correction found during this audit: /root/.hermes/scripts/doc-live-verify.py
currently hardcodes app3.itpropartner.com specs as "RS 4000 (8 vCPU, 16 GB RAM, 320 GB SSD)"
— this is stale even before app4/core-bu (app3 is actually 12 vCPU/32GB/1TB per the live
server-architecture-plan skill and architecture.md). Flagged here since it lives in the
same SERVER_INVENTORY structure that needs the P1/P2 edits below.
How to read this matrix
- Phase (P1/P2/P3) — P1: edit at provisioning time (DNS/rDNS/inventory/monitoring/backup targets/decommissioned-hosts/SSH keys), before any service moves. P2: edit at the moment services actually cut over to app4 / core-bu takes over Core's standby role. P3: edit only once app1-bu is deleted from the Hetzner account.
- Current is the exact text/value read from the file (quoted or line-numbered).
- Required new value is what the artifact must say once the transition in that phase lands.
- Rows are grouped by artifact class: docs, scripts, cron/systemd, Caddy, Gitea/repo, ops-portal, monitoring/decommission tooling.
P1 — Must change at provisioning time (DNS/rDNS, inventory, monitoring watch lists, backup targets, decommissioned-hosts, SSH key distribution)
| # | Artifact | Current text/value (evidence) | Required new text/value | Why it matters |
|---|---|---|---|---|
| 1 | docs/infrastructure/key-inventory.md line 14 |
Universal server admin key... All servers (Core, app1, app2, app3, app1-bu, home router). |
Add app4 (159.195.205.80) and core-bu (159.195.204.203) to the deployed-to list for itpp-infra the moment the SSH key is injected at provisioning |
Key inventory is the audit trail for "who can SSH where" — a missing entry means the new boxes are invisible to a credential audit |
| 2 | docs/infrastructure/key-inventory.md §2 table (Server Root Passwords) |
Rows for Core/app1/app2/app3/app1-bu only | Add app4 and core-bu rows: IP, netcup, access method, Vaultwarden pointer | Same table used for DR "how do I get in" — missing rows = no documented root access path |
| 3 | /root/.hermes/references/decommissioned-hosts.json |
No entries for app4/core-bu (they don't exist yet); app1-bu is still listed as live nowhere in this file (correct — it's not decommissioned) | No P1 change needed to this file for app4/core-bu (they are new, not decommissioned). Flagged here only to confirm scope: this file is a P3 artifact for app1-bu, not P1 | Prevents accidentally graveyard-listing a box that is being born, not retired |
| 4 | /root/.hermes/scripts/stale-reference-verify.py (docstring + SCAN_ROOTS) |
Scans /etc/systemd/system/, /root/.hermes/scripts/, /etc/cron.d/, /var/spool/cron/crontabs/ for graveyard hits |
No functional change required at P1 (app4/core-bu aren't graveyarded). Verify after P1 that no stray reference to the old "core-bu = Hetzner" naming (see row 20) creates a false-clean scan | The script's correctness depends on decommissioned-hosts.json staying accurate; a naming collision (two things called "core-bu") could mask a real stale reference later |
| 5 | /root/.hermes/scripts/doc-live-verify.py lines 50-57 (SERVER_INVENTORY) |
Dict has entries for core.itpropartner.com, app1, app2, app3.itpropartner.com, app1-bu.itpropartner.com only — no app4/core-bu entries exist |
Add app4.itpropartner.com (159.195.205.80, netcup, RS 4000 G12, "customer-facing services") and a core-bu entry (159.195.204.203, netcup, RS 2000 G12 twin, "warm standby for Core") to SERVER_INVENTORY and domain_map (lines ~90-108) |
Without this, the scanner has no baseline for the new IPs and can't tell a correct new reference from a typo'd one |
| 6 | /root/.hermes/scripts/health-master-watchdog.py REMOTE_SERVERS (lines ~91-97) |
[("app1", "152.53.36.131"), ("app2", "152.53.39.202"), ("app3", "152.53.241.111"), ("app1-bu", "5.161.225.131"), ("anita-mnz", "159.195.16.30")] |
Add ("app4", "159.195.205.80") and ("core-bu", "159.195.204.203") at provisioning so ping/reachability + backup-freshness checks cover the new boxes from day one, per the precedent set for anita-mnz on 2026-09-11 |
This exact list is what the skill's own pitfall log calls out — a box not added here alerts never, a box removed from the wrong list (old app1-bu) alerts forever after retirement |
| 7 | /root/.hermes/scripts/health-master-watchdog.py REMOTE_DOCKER_CONTAINERS (lines ~108-124) |
Keys for app1, app2 only; comment: # app1-bu (Hetzner standby) has no Docker; app3 is CloudPanel-only. |
Add an app4 key with the moved containers' names once DocuSeal/TimeTrex/microbin/Uptime Kuma/Ops-Portal-backend containers exist there (P2 timing for the container list itself, but the dict key should exist and be empty-safe at P1 so the code path is proven) |
Container health checks are useless if the box hosting them was never added to the watch structure |
| 8 | /root/.hermes/scripts/vps-threshold-check.sh lines 23-27 |
`SERVERS=("core | 152.53.192.33 | 1" "app1 |
| 9 | /root/.hermes/scripts/security-compliance-check.sh line 6 |
SERVERS="152.53.192.33:core 152.53.36.131:app1 152.53.39.202:app2 152.53.241.111:app3 5.161.225.131:app1-bu" |
Append 159.195.205.80:app4 159.195.204.203:core-bu |
Nightly compliance sweep (patches, fail2ban, failed systemd units) silently skips any host not in this string |
| 10 | /root/.hermes/scripts/backup-health-monitor.sh line 774 |
for server_entry in "core:152.53.192.33" "app1:152.53.36.131" "app2:152.53.39.202" "app3:152.53.241.111"; do (Phase 3: System Crontabs SSH check) — note app1-bu is already absent from this specific loop |
Add "app4:159.195.205.80" and "core-bu:159.195.204.203" once those boxes have crontabs to audit |
This SSH-crontab-audit loop already under-covers (app1-bu missing); don't propagate the gap to the two new boxes |
| 11 | /root/.hermes/scripts/backup-failure-check.sh line 103 |
for server in "app1:152.53.36.131" "app2:152.53.39.202" "app3:152.53.241.111"; do |
Add app4 (and core-bu if it carries backup jobs of its own beyond the standby sync) | Same class of gap as row 10 — backup failure detection has a fixed server list |
| 12 | /root/.hermes/scripts/api-health-check.py lines 62-83 |
Hardcoded (name, mode, ip, port, path) tuples for app1/app2/app3 endpoints only |
Add app4 tuples for each service once it lands there (DocuSeal :3000, TimeTrex :8085, microbin :8260, Uptime Kuma :3001, Ops Portal :8090) — this is P2-timed for the actual endpoints, but the provisioning step should reserve the pattern now | API health checks won't exist for anything moved to app4 until this file is edited |
| 13 | /root/.hermes/references/reserved-ports.json |
_comment: "Reserved service ports on core (152.53.192.33)." Lists ports 8899, 8500, 8090, 8099, 8200, 8910, 8888, 3002, 3000 all scoped to Core |
When Ops Portal (:8090), DocuSeal (:3000) etc. move off Core to app4, this file's port-identity-guard becomes partially stale for Core and needs an equivalent file (or scope note) for app4 | port-identity-guard.py (referenced by this file) exists to catch port-squatting; if the service moves and the guard still watches Core's now-empty port, a real squat on app4 goes undetected |
| 14 | /root/.hermes/cron/jobs.json — hetzner-weekly-snapshots job (id faa6b8760e38, script snapshot-hetzner.py) |
Runs weekly Mon 5:00 AM, no_agent: true, targets Hetzner API broadly |
No P1 edit required (this job is Hetzner-API-driven, not IP-hardcoded in jobs.json itself — confirm the script's own server-ID list separately, see row 27) | Flagged for completeness; the actual IP/server-ID list lives inside snapshot-hetzner.py, not jobs.json |
| 15 | docs/infrastructure/app4-migration-plan.md line 70 |
**Recommendation: netcup RS 4000 G12 (12 vCPU / 32 GB DDR5 ECC / 1 TB NVMe), ~$44/mo, Manassas VA.** |
Correct to Nuremberg (Manassas RS was sold out per standby-host-replacement-2026-09-14.md context and the server-architecture-plan skill's 2026-09-15 update): app4 is ordered at Nuremberg, IP 159.195.205.80, rDNS v2202609377162521278.quicksrv.de |
The plan's own location assumption is now wrong; anyone following it to provision would target the wrong DC and the wrong ordering flow |
| 16 | docs/infrastructure/app4-migration-plan.md (no core-bu content at all — plan only discusses app4) |
Plan silent on the new core-bu box entirely; describes only Core→app4 | Add a section (or a companion doc) covering core-bu's provisioning as Core's new-generation standby, superseding the Hetzner-only DR references throughout hermes-standby-deployment skill and server-dr-plans.md |
The single biggest gap in the existing plan set — core-bu's arrival redefines the entire DR architecture (a second, same-provider-family warm standby) and nothing documents it yet |
| 17 | README.md lines 81-82, 91, 94 |
**Hostname:** core-bu / **IP:** 5.161.225.131 ... **Hetzner Cloud (current):** As of 2026-08-28, the Hetzner Cloud API returns exactly **one** server — **app1-bu / core-bu** (5.161.225.131...) |
This is the exact naming collision flagged by the server-architecture-plan skill's 2026-09-14 correction: "core-bu is a name for the standby role, not a reference to the Hetzner host." Update README to stop calling the Hetzner box core-bu and use app1-bu only, reserving core-bu for the new netcup box (159.195.204.203) |
This is a live, already-existing naming defect (not just a future one) — three files (README.md, docs/infrastructure/key-inventory.md §2, network-diagram.md, master-apps-services.md) currently call the Hetzner box "core-bu", which will directly collide with the new netcup box once it's named core-bu. Fix at P1, before the second "core-bu" exists, or every future reference is ambiguous |
| 18 | /root/.hermes/references/ip-dns-changes.md line 17 and server-inventory.md line 14 and master-apps-services.md line 19 and network-diagram.md line 54 |
All four call the Hetzner box (5.161.225.131) "core-bu" | Same fix as row 17 — rename all four to app1-bu for the Hetzner box, and add net-new rows for the netcup core-bu (159.195.204.203) once ordered |
Four separate documents currently reinforce the same naming collision; all four must be corrected together or drift resumes immediately |
| 19 | docs/monitoring/uptime-kuma-monitoring-plan.md §2 (Core/App1/App2/App3/wphost02 Caddy sections) and the A1-A4 Server Health monitor rows (lines 392-395) |
Monitors "Core", "app1", "app2", "app3" by IP only — no app1-bu, app4, or core-bu row exists in the Uptime Kuma monitor list | Add A5 (app4, 159.195.205.80) and A6 (core-bu, 159.195.204.203) TCP/22 health monitors at provisioning; keep app1-bu's own monitor (if any exists in live Kuma — not confirmed by this audit, Kuma admin UI not inspected) until P3 | Public-facing Uptime Kuma is the customer-visible signal; a new customer-facing host (app4) with zero monitors defeats the purpose of the migration |
| 20 | /root/.hermes/references/server-inventory.md (whole file, dated July 10 2026) |
Table has Core/app1/app2/app3/core-bu(Hetzner)/old-ai/docker-box/wphost02/UNMS/UniFi — badly stale even before this change (old-ai, docker box, wphost02 are all long decommissioned per decommissioned-hosts.json) |
This file needs a full refresh regardless of app4/core-bu; at minimum add app4 and core-bu rows and correct the core-bu/app1-bu naming per row 17 | Already the most stale document found in this sweep — nothing here should be trusted as current without cross-checking server-architecture-plan skill or live probes |
| 21 | /opt/ops-portal/server.py lines 370-374 and 649-653 and 1878-1881 |
Three separate hardcoded server lists: {"name": "Core", "ip": "152.53.192.33", ...} etc., each listing Core/app1/app2/app3/app1-bu only, no app4/core-bu |
Add app4 and core-bu entries to all three lists (they are NOT DRY — each must be edited independently, a known Caddyfile-style risk) | Ops Portal is the customer/staff-facing dashboard; three independently-maintained lists mean three places to miss, and the portal itself is moving to app4 (see row 24) which makes this doubly urgent |
| 22 | /opt/ops-portal/server.py line 706 and 1912 (REMOTE_DOCKER_CONTAINERS-equivalent dict, app1-bu key) |
"app1-bu": [{"name": "caddy", "type": "systemd"}, {"name": "docker", "type": "systemd"}] (two separate near-identical dicts at lines ~706 and ~1912) |
Add "app4" key with its own service list once services move; keep "app1-bu" key until P3 |
Same all-lists-must-be-edited risk as row 21 |
| 23 | /opt/ops-portal/static/dependency-diagram.html line 198 |
<text ...>app1-bu.itpropartner.com</text> (SVG dependency diagram) |
Add app4.itpropartner.com and core-bu nodes to the diagram at provisioning | Static SVG diagrams don't regenerate themselves — this is hand-edited HTML that silently goes stale |
P2 — Must change at service cutover (app4 takes over customer-facing services, core-bu becomes Core's warm standby)
| # | Artifact | Current text/value (evidence) | Required new text/value | Why it matters |
|---|---|---|---|---|
| 24 | /etc/caddy/Caddyfile (live, read 2026-09-15) — every customer-facing site block |
default_bind 152.53.192.33 (global block) plus individual reverse_proxy 127.0.0.1:PORT blocks for sign.itpropartner.com, ops.itpropartner.com, my.itpropartner.com, status.itpropartner.com, uptimekuma.itpropartner.com, voice.itpropartner.com, voice-open.itpropartner.com, auth.itpropartner.com, and the my.itpropartner.com block's reverse_proxy http://152.53.241.111:8090 calls to app3's backup-restore API |
Per app4-migration-plan.md §2.2/§4 Phase 4: after data migration and DNS cutover, remove these site blocks from Core's Caddyfile (targeted edits + caddy-audit hook, never full rewrite) and stand up an equivalent Caddyfile on app4 with default_bind 159.195.205.80. Core's Caddyfile keeps only core.itpropartner.com and Core-local routes |
This is the single largest, most consequential P2 change — every customer domain's TLS termination and backend routing moves. Getting the Caddyfile edit wrong breaks live customer traffic |
| 25 | dns-records.md §PRODUCTION table (lines 21-43) |
ops.itpropartner.com → 152.53.192.33 → Core, my.itpropartner.com → 152.53.192.33 → Core, sign.itpropartner.com → 152.53.192.33 → Core, uptimekuma.itpropartner.com → 152.53.192.33 → Core, status.itpropartner.com → 152.53.192.33 → Core |
Flip A records for each moved domain to 159.195.205.80 (app4) per the phased DNS cutover in app4-migration-plan.md §4/§6 (lower TTL first, soak, then cut). SiteGround manual panel for itpropartner.com — no API |
DNS is the actual cutover mechanism; the doc must track reality or the next person "fixing" a stale doc could break live routing |
| 26 | backup-plan.md §Core backup inventory (lines 12-19) and §Schedules Summary |
core-services-backup.sh targets Grafana/Uptime-Kuma/Docker-volumes/Prometheus all on Core (152.53.192.33); Ops Portal backend backup is implied under root-essentials-backup.sh/core-services-backup.sh on Core |
Per app4-migration-plan.md §7: repoint backup scripts and cron for every moved app (DocuSeal, TimeTrex, microbin, Uptime Kuma, Ops Portal backend, Postgres, Redis) from Core to app4. Add an "App4" backup section to backup-plan.md mirroring the App1/App2/App3 sections, with its own S3 prefix (s3://hermes-vps-backups/app4/...) |
If the app4 migration plan's own §7 instruction ("confirm app4 cron owns them before removing Core entries") isn't followed, moved services lose backup coverage silently — exactly the anita-mnz *.db-exclusion failure mode from 2026-09-11 |
| 27 | /root/.hermes/scripts/uptime-kuma backup script (core-services-backup.sh, per backup-plan.md row) and the Uptime-Kuma container itself |
Uptime Kuma runs on Core per server-architecture-plan skill ("Uptime-Kuma (Docker, port 3001) ... on Core") and is backed up via core-services-backup.sh |
Per app4-migration-plan.md §2.2/Phase 2 ("Move microbin and Uptime Kuma first"), Uptime Kuma is the first DNS flip of the whole project — update core-services-backup.sh to stop covering it and add it to a new app4-backup.sh, and update uptimekuma.itpropartner.com Caddy block + DNS together |
Uptime Kuma is the monitoring tool watching everything else — if its own migration breaks its backup or its DNS without coordinated verification, you lose visibility into the exact moment you need it most |
| 28 | hermes-standby-deployment skill and /root/.hermes/scripts/hermes-standby-watchdog.sh / hermes-standby-sync.sh / hermes-standby-restore.sh |
LIVE_HOST="152.53.192.33" hardcoded in all three scripts (watchdog line 8, sync line 11, restore line 17); watchdog alert text: Standby: app1-bu.itpropartner.com (5.161.225.131) |
When core-bu (159.195.204.203) becomes the standby role for Core (replacing/supplementing app1-bu), the design question is: does core-bu run these SAME scripts pointed at the same LIVE_HOST=152.53.192.33, or does app1-bu get retired and core-bu take over the whole watchdog/sync/restore stack? Either way, LIVE_HOST stays 152.53.192.33 (Core's IP doesn't change) but Standby: alert text and the deployment target change to core-bu's IP/hostname. This decision must be made explicit in a doc before P2 lands, not inferred from script diffs |
Two standbys with two independently-running watchdog/sync/fencing setups is a split-brain risk if both start Hermes on Core-down; the architecture decision (core-bu replaces app1-bu vs. runs alongside it) is currently undocumented anywhere in the swept files |
| 29 | hermes-standby-deployment skill "2026-09-15 correction" section — fence_core() |
fence_core() SSHes to the live box (PROBE_HOST/PROBE_SSH_KEY) and runs systemctl --user stop hermes-gateway |
If core-bu becomes the active standby, PROBE_HOST must point at Core (unchanged, 152.53.192.33) but the fencing SSH key and the box RUNNING fence_core() moves to core-bu. Update the skill's Testing/Numbers-that-matter section to reflect which box is standby-of-record |
The recently-fixed failover state machine (2026-09-15 correction) was hard-won; redeploying it on a new box without re-verifying TAKEOVER_AFTER/DRYRUN/fence behavior risks reintroducing the exact bug just fixed |
| 30 | docs/infrastructure/standby-host-replacement-2026-09-14.md (whole document) |
Entire document's recommendation is "move app1-bu to Hetzner fsn1/nbg1 for cost + regional diversity" — written one day before the decision to build a netcup core-bu instead | This document's core recommendation is superseded by the actual decision (order a second netcup box, not relocate the Hetzner one). Add a superseding note at the top: "Superseded 2026-09-14/15 — owner ordered netcup core-bu instead of relocating app1-bu; see app4-migration-plan.md and this matrix" | Without an explicit supersession note, a future reader (or scanner) will treat this doc's Hetzner-relocation recommendation as still-live guidance, wasting effort or causing a wrong action |
| 31 | docs/infrastructure/key-inventory.md §2 (Server Root Passwords) app1-bu row: "Warm standby (core-bu)" |
Parenthetical calls the Hetzner box "core-bu" — same naming collision as row 17, but specifically in the credentials table | Fix at P1 per row 17, but re-verify at P2 that the role note is accurate once core-bu (netcup) is actually live and taking over standby duties — app1-bu's row should then say "legacy warm standby, superseded by core-bu, retirement candidate" | Credential documentation errors compound: an admin reading this table during an incident needs to know which box is the ACTIVE standby right now |
| 32 | Gitea itpp-infrastructure repo (git.itpropartner.com, confirmed reachable, HTTP 200, API responsive) — same file set as this local clone |
Remote origin is https://ippadmin:***@git.itpropartner.com/ippadmin/itpp-infrastructure.git; per git remote -v this is the only remote, and it's the push target for every doc listed above |
No new remote needed; simply confirm every P1/P2 edit above gets committed and pushed to this repo (the audit found this repo push-current as of 2026-09-12 commit 199baad). The disaster-recovery Gitea repo (also confirmed reachable) should receive a parallel supersession note if it duplicates any of the standby/DR content |
Local-only edits that never reach Gitea are invisible to anyone cloning fresh, and to hermes-migration/server-recovery-bundle workflows that pull from Gitea, not from this specific Core filesystem |
| 33 | Ops Portal TRACCAR_HOST and similar single hardcoded backend-IP constants (server.py line ~1348) |
TRACCAR_HOST = "152.53.39.202" (app2, unaffected by this migration, listed here only because it lives in the same file as the app1-bu/app4 dicts flagged in rows 21-22) |
No change required — app2 is out of scope. Listed for completeness so a future editor doing a "replace all IPs" pass on server.py doesn't accidentally touch this unrelated constant |
Prevents scope creep / accidental breakage during the app4 edit pass |
P3 — Must change at app1-bu retirement (every place that would break or lie once the Hetzner box is gone)
| # | Artifact | Current text/value (evidence) | Required new text/value | Why it matters |
|---|---|---|---|---|
| 34 | /root/.hermes/references/decommissioned-hosts.json |
No entry for app1-bu / 5.161.225.131 (it is currently live, correctly absent) | Add an entry: {"hostname": "app1-bu", "ip": "5.161.225.131", "decommissioned": "<retirement date>", "services": ["hermes-gateway", "caddy", "docker"], "note": "Hetzner CPX21 warm standby, superseded by netcup core-bu (159.195.204.203). Retired <reason>. Provider order cancellation is a USER action per server-decommissioning skill."} |
This is THE canonical source stale-reference-verify.py and doc-live-verify.py both read — everything else in this phase depends on this entry existing first |
| 35 | /root/.hermes/scripts/health-master-watchdog.py REMOTE_SERVERS line ~94 |
("app1-bu", "5.161.225.131") |
Remove this tuple entirely (or comment it out with a decommission note, per the anita-mnz precedent at line ~92 of the same file: # hermes-gateway-anita.service moved to the dedicated box... It is checked remotely via REMOTE_USER_UNITS below, NOT locally here.) |
A dead host left in REMOTE_SERVERS alerts "unreachable" every single watchdog cycle forever — this exact failure mode is documented in the server-decommissioning skill as the recurring pitfall |
| 36 | /root/.hermes/scripts/vps-threshold-check.sh line 27 |
`"app1-bu | 5.161.225.131 | 0"` |
| 37 | /root/.hermes/scripts/security-compliance-check.sh line 6 |
...5.161.225.131:app1-bu" (trailing entry in the SERVERS string) |
Remove 5.161.225.131:app1-bu from the string |
Nightly compliance check SSHes to every listed host; a deleted server means permanent SSH-failure noise |
| 38 | hermes-standby-deployment skill (entire skill file) |
Skill is written entirely around app1-bu/Hetzner: "Hetzner CPX21 recommended", enable_rescue API calls, Hetzner-specific rescue-mode key injection, SERVER_ID referencing Hetzner's API, cost tables in EUR/Hetzner pricing |
This skill needs either (a) a full rewrite for netcup-based core-bu deployment (no Hetzner rescue mode, no Hetzner API, different provisioning flow per server-provisioning-standard), or (b) an explicit "ARCHIVED — app1-bu retired, see for core-bu" banner if a new skill is authored separately |
The entire deployment runbook (registering SSH keys via Hetzner API, rescue-mode key injection, CPX11→CPX21 upgrade math) is Hetzner-specific and becomes 100% inapplicable once core-bu (netcup, no rescue mode, standard provisioning) is the standby. Leaving this skill as the "how to deploy a standby" reference after retirement would actively mislead the next person who has to redeploy or troubleshoot |
| 39 | /root/.hermes/references/hermes-dr-plan-v2.md (entire document) |
Describes app1-bu exclusively: Core (netcup RS 2000 — 152.53.192.33) / app1-bu (Hetzner CPX21 — 5.161.114.8) (line 12, itself already stale — 5.161.114.8 was replaced by 5.161.225.131 back on 2026-07-24), SSH key itpp-infra-v2 deployed "to app1-bu via rescue", Hetzner-specific fencing script, CPX11→CPX21 upgrade tables |
Full rewrite required: replace every app1-bu/Hetzner/CPX21/rescue-mode reference with core-bu/netcup/RS-2000-twin/standard-provisioning equivalents, or mark the document ARCHIVED and write a v3 | This is the master DR plan referenced by the hermes-standby-deployment skill's own "Related Documents" table as "the target architecture" — if it still describes a deleted box as the target, DR execution during a real incident follows a runbook for infrastructure that no longer exists |
| 40 | /root/.hermes/references/3-Per-Server-Runbooks.md lines 16, 24, 35, 50-51, 54, 74 |
Hermes orchestration hub — highest priority, protected by warm standby (app1-bu); app1-bu (standby), S3 backup bucket...; If app1-bu is healthy, use failover...; Both Core and app1-bu active simultaneously — mitigated by...; app1-bu itself is restored like any Standby-tier host if it fails: provision replacement CPX11... |
Replace every app1-bu reference with core-bu across all 6+ occurrences; the CPX11 replacement-provisioning instruction (line 74) becomes netcup RS-line provisioning instead | Per-server runbooks are the document someone opens DURING an incident — six stale references to a deleted host in the runbook for CORE's own failover path is a P0-during-an-incident risk |
| 41 | /root/.hermes/references/itpp-recovery-manual.md (8 occurrences across lines 17, 48, 52, 72, 272-326, 676-679, 1031-1037, 1185-1327) |
Extensive: table of contents anchor #4-app1-bu-standby-hetzner--51611148 (itself referencing the OLD retired IP 5.161.114.8 in the anchor text — a pre-existing stale anchor found during this audit), full §4 "app1-bu Standby (Hetzner — 5.161.225.131)" section with SSH commands hardcoded to root@5.161.225.131, StrongSwan/L2TP fallback note "runs on both Core and app1-bu" |
Full section rewrite: new §4 "core-bu Standby (netcup — 159.195.204.203)" with corrected SSH commands, updated StrongSwan/L2TP note, and the TOC anchor fixed (it currently embeds a wrong IP even for app1-bu) | This is the master recovery manual — 39,541 bytes, 1351 lines, the document row 4 of the key-inventory's "Recovery Priority" list implicitly assumes is accurate. Every ssh root@5.161.225.131 command in it will fail post-retirement, and following it during a real outage wastes the exact minutes DR is supposed to save |
| 42 | /root/.hermes/references/restore-runbooks.md line 5 |
**Rollback plan:** Revert DNS to standby core-bu (5.161.225.131). Wipe server and restart runbook. |
This line ALREADY uses "core-bu" for the Hetzner IP (the naming collision from row 17, found here too) — at retirement, either this line's target no longer exists (if app1-bu/old-core-bu is deleted) or it needs redirecting to the NEW core-bu at 159.195.204.203 | A rollback plan pointing at a deleted IP is worse than no rollback plan — it will be trusted and fail silently mid-incident |
| 43 | /root/.hermes/references/restore-plan-2026-07-11.md (8+ occurrences, e.g. lines 11, 38, 151-210, 332-333) |
Entire "Recovery Path B: Failover to Warm Standby (app1-bu)" section with live ssh -i /root/.ssh/itpp-infra root@5.161.225.131 commands |
This is a dated incident-response document (July 11, 2026) describing a past incident — per the server-decommissioning skill's "Stale-IP cleanup in docs" guidance, dated historical incident reports are typically left as-is with an inline decommission annotation rather than rewritten, since they document what WAS done, not current procedure. Recommend annotating the top of the doc: "Historical — app1-bu (5.161.225.131) referenced below was retired ; current standby is core-bu (159.195.204.203)" rather than rewriting the incident narrative |
Rewriting historical incident reports to reflect infrastructure that didn't exist at the time falsifies the record; annotation preserves both truth and utility |
| 44 | docs/infrastructure/key-inventory.md lines 24 (SSH deployment scope), 42 (root password table), 221 (Unknown/Not Found §13 duplicate app1-bu row) |
All servers (Core, app1, app2, app3, app1-bu, home router) / app1-bu | 5.161.225.131 | Hetzner CPX21 | itpp-infra SSH key | Warm standby (core-bu) / duplicate row 221: **app1-bu** | 5.161.225.131 | Hetzner CPX21 — accessed via itpp-infra SSH key only. (note: this row is misplaced inside §13 "Unknown/Not Found", itself a pre-existing doc defect) |
Remove app1-bu from the SSH deployment scope line; remove or annotate-as-historical both the §2 row and the misplaced §13 duplicate row | Three separate mentions of app1-bu in one document, one of them filed under the wrong section header — all three need the retirement edit or two of three will be missed |
| 45 | docs/infrastructure/key-inventory.md §11 Tailscale table, app1-bu row |
app1-bu | 100.112.23.21 | Linux | ⚠️ Offline (7d) — confirmed still true live via tailscale status on 2026-09-15 (shows 100.112.23.21 app1-bu ... offline, last seen 61d ago, plus a SECOND stale node 100.95.212.28 app1-bu-1 currently idle) |
At retirement: remove the app1-bu (100.112.23.21) row entirely; also remove/rename the app1-bu-1 node once confirmed it's the same retired box (Tailscale auto-renamed it per the hermes-standby-deployment skill's Aug 7 pitfall note: "the node appears as app1-bu-1 (auto-renamed because the old app1-bu node was parked offline for 22+ days)") |
Live evidence found DURING this audit: there are currently TWO Tailscale nodes for the same physical box (app1-bu offline 61 days, app1-bu-1 idle/active). This is already a live inconsistency, not just a future one — flag now, clean up at retirement |
| 46 | README.md (multiple: lines 81-94, 176-177, 253, 277) |
**Hostname:** core-bu / **IP:** 5.161.225.131 / ~~wphost02-backup~~... **REMOVED — wphost02 decommissioned** (shows the precedent pattern to follow) / warm-standby-sync | Every 10 min | core-bu ← S3 | DR readiness / **Provider diversity:** core-bu stays at Hetzner specifically so a netcup outage can't kill both Core and standby simultaneously |
Follow the exact strikethrough-and-note pattern already used for wphost02-backup (line 176) for every app1-bu/old-core-bu line: strike it, note "REMOVED — app1-bu decommissioned , superseded by core-bu (netcup, 159.195.204.203)". The "provider diversity" claim (line 277) becomes FALSE once core-bu is also netcup — this line must be corrected to state the diversity argument no longer applies (or reframed around Nuremberg vs Manassas regional diversity within netcup, which is weaker than true provider diversity) | Row 277 is a substantive factual claim ("provider diversity... netcup outage can't kill both") that becomes false, not just outdated, once the standby is also netcup. This is the highest-risk single line in the whole sweep — see Top 5 below |
| 47 | docs/infrastructure/app4-migration-plan.md §7 line 205 |
Decide standby scope: app1-bu is a warm standby for Core, not for customer apps. app4 relies on S3 backups unless a customer-app standby is separately approved. |
Update to reference core-bu instead of app1-bu once it's the standby of record, and resolve the open question (does core-bu ALSO need to be app4's standby, or does app4 remain S3-only for DR) | This is an explicitly flagged open decision in the plan itself — retirement of app1-bu is the forcing function to finally resolve it |
| 48 | docs/architecture.md line 17 |
` | app1-bu | 5.161.225.131 | CPX21 (3 vCPU, 4 GB RAM, 80 GB) | Hetzner | Warm standby — provider diversity. Auto-restores from S3. | ` |
| 49 | docs/infrastructure/standby-host-replacement-2026-09-14.md §"Rebuild procedure" (final section) |
Reference the hermes-standby-deployment skill. Sequence: create CPX21 in fsn1 or nbg1, run the standby deploy script, restore from s3://hermes-vps-backups/live/, verify the failover cron... |
This entire rebuild procedure is Hetzner-specific and becomes fully inapplicable; either delete this section or replace with the netcup core-bu equivalent procedure once one exists | A "how to rebuild the standby" procedure that references a decommissioned provider's rescue-mode tooling is actively harmful if followed after retirement |
| 50 | /root/.hermes/references/network-diagram.md lines 26-27, 33-39, 54, 67 |
ASCII diagram box labeled [Standby Host] (app1-bu); [Legacy Net] (Hetzner) box listing decommissioned services plus implying app1-bu lives there too; **Standby Host (core-bu)**: 5.161.225.131 — Warm standby (CPX21), connected via Tailscale to Core. (again the naming collision); Tailscale Mesh description says "between Primary Host (Core) and Standby Host (core-bu)" |
Redraw the ASCII diagram: Standby Host box becomes core-bu at 159.195.204.203 (netcup box, not "Legacy Net"/Hetzner); remove the app1-bu box or move it to a "retired" annotation | ASCII diagrams are easy to skip during edits because they're not table rows — this file's diagram will visually lie about network topology post-retirement if not redrawn |
| 51 | /root/.hermes/references/master-apps-services.md (10 occurrences: lines 19, 30, 260-271, 307, 356-367, 386, 421-423) |
Extensive core-bu-as-Hetzner-name usage: **core-bu** | 5.161.225.131 | CPX11 (2C/2G/40G) | Hetzner | Warm standby for Core; ### core-bu (Hetzner CPX11 — 5.161.225.131) — Warm Standby; Tailscale | BSD | Core, core-bu; **Hermes Agent** | Proprietary | Core, core-bu; **DR Standby:** core-bu (Hetzner) boots and auto-restores from S3 if Core is down for 2+ minutes. |
Ten separate lines in one document all need the app1-bu/core-bu rename at P1 (existing collision) and then a further retirement edit at P3. This is the single most-referenced document for the naming collision found in this sweep | Ten independent edit points in one file is exactly the kind of surface a manual sweep misses one or two of — recommend a scripted find/replace pass on this file specifically, verified line-by-line afterward |
| 52 | Gitea disaster-recovery repo (confirmed reachable via API, git.itpropartner.com/ippadmin/disaster-recovery, last commit referenced in restore-test-log.md line 43: disaster-recovery.git | 5 | 3bc6d08 Fix: app1-bu CPX11 → CPX21 spec) |
Repo history shows this repo has previously been edited specifically for app1-bu spec corrections — implying it contains its own copy of DR content that will need the same app1-bu→core-bu treatment | Clone and sweep this repo's content directly (not done in this audit — only confirmed reachability and one commit-log reference via the local restore-test-log.md) before declaring the retirement documentation complete |
This repo was NOT directly inspected in this pass (Core's local checkout doesn't contain it) — flagged as an open item, not a completed row |
Artifacts referencing hosts that no longer exist (already stale, independent of app4/core-bu/app1-bu)
These were found while sweeping the same files and are stale right now, unrelated to the current provisioning work. Listed per the task's explicit request.
| Host / IP | Where still referenced (live, not historical-annotated) | Status per decommissioned-hosts.json |
|---|---|---|
wphost02 / 5.161.62.38 |
docs/monitoring/uptime-kuma-monitoring-plan.md §2.5 header still titled "wphost02 (5.161.62.38) — RunCloud (LEGACY — being migrated)" (present tense, "being migrated", not "migrated"); /root/.hermes/references/network-diagram.md line 59 "wphost02: 5.161.62.38 — DECOMMISSIONING" (present-progressive, not past); master-apps-services.md line 25/43/70/245/275/386/421 refer to wphost02 in present tense in several spots even though the file elsewhere (line correctly) notes migration |
Decommissioned 2026-08-28, confirmed in decommissioned-hosts.json and doc-live-verify.py's known_historical |
178.156.130.130 (old standalone Hudu) |
network-diagram.md does not list it directly but decommissioned-hosts.json carries it with "decommissioned": null (the field is present but not dated, unlike the other entries) |
Listed in graveyard but missing a decommission DATE — a data-quality gap in the source-of-truth file itself |
5.161.114.8 (old app1-bu IP) |
/root/.hermes/references/hermes-dr-plan-v2.md line 12: app1-bu (Hetzner CPX21 — 5.161.114.8); itpp-recovery-manual.md TOC anchor #4-app1-bu-standby-hetzner--51611148 still embeds this dead IP in the anchor slug even though the section body correctly uses 5.161.225.131 |
Decommissioned 2026-07-24, correctly in known_historical, but doc-live-verify.py's own DR audit (dr-issue-log.md, 2026-09-15 entry) already flags this exact hit as "benign" since the file is historical — confirms the scanner is working as designed for THIS one, but the two live-doc mentions above were not caught because doc-live-verify skips files with decommission-marker words nearby, not exact-line context |
178.156.167.181 (old admin-ai) |
hetzner-server-inventory.md (already self-marked "ARCHIVED — Historical Reference Only") and app-inventory.csv (also self-marked ARCHIVED) — both correctly annotated |
Decommissioned, correctly archived |
87.99.144.163 (old app1) |
decommissioned-hosts.json graveyard entry has "decommissioned": null — no date |
Same data-quality gap as 178.156.130.130 |
87.99.159.142 (Tony's old VPS) |
/root/.hermes/references/ip-dns-changes.md line 95: tony.iamgmb.com | 87.99.159.142 | Tony VPS | Tony's Hermes — this is a LIVE DNS record row in a document dated August 12, 2026, for a server the same file's own line 26 lists as "DELETED Jul 14" |
Contradiction WITHIN the same document: line 26 says deleted, line 95 lists it as a current DNS target |
Top 5 highest-risk stale references found
-
README.md line 277 — "provider diversity" claim becomes factually false, not just outdated. Once core-bu (netcup) replaces or joins app1-bu (Hetzner) as Core's standby, the stated rationale "core-bu stays at Hetzner specifically so a netcup outage can't kill both Core and standby simultaneously" is wrong the moment the standby is also netcup. This is a substantive risk claim in the primary README, not a cosmetic IP mismatch, and the same false claim is repeated in
docs/architecture.mdline 17 andmaster-apps-services.mdline 271. -
The "core-bu" naming collision already exists across 6 files before the new box is even ordered.
README.md,docs/infrastructure/key-inventory.md,network-diagram.md,master-apps-services.md,restore-runbooks.md, andip-dns-changes.mdall currently call the Hetzner box (5.161.225.131) "core-bu" — exactly the name the owner's plan assigns to the new netcup box (perserver-architecture-planskill's 2026-09-14 correction note). If this isn't fixed before core-bu goes live, every future reference to "core-bu" in these six files is ambiguous between two physically different servers with different providers, specs, and failure domains. -
itpp-recovery-manual.md— 8 live SSH commands hardcoded toroot@5.161.225.131inside the master recovery runbook. This is the document opened during an actual incident. Post-retirement, every one of these commands fails, and the failure mode during a live outage (typing a command that connects to nothing, or worse, to a re-leased IP with a different owner) is worse than having no runbook at all. -
hermes-standby-deploymentskill andhermes-dr-plan-v2.mdare entirely Hetzner-API-specific (rescue mode,enable_rescue, Hetzner SSH key registration, CPX11→CPX21 upgrade math) and become 100% inapplicable to a netcup-based core-bu. There is currently no equivalent skill or plan describing how to deploy/audit/troubleshoot a netcup-based warm standby — this is a capability gap, not just a stale reference, and it will be discovered mid-incident if not addressed before app1-bu is actually retired. -
Two live Tailscale nodes currently exist for one physical box (
app1-buat 100.112.23.21, offline 61 days;app1-bu-1at 100.95.212.28, idle/active) — found live during this audit, not from a doc. This is a pre-existing data-quality problem in the mesh itself (documented as a known Tailscale auto-rename behavior in thehermes-standby-deploymentskill's pitfalls, but never cleaned up) that will complicate identifying which Tailscale node to decommission when app1-bu is formally retired, and could cause the retirement checklist to miss one of the two.
Counts
- Total artifacts (files) touched by at least one required edit: 33 distinct files/scripts/configs identified with concrete row-level evidence, plus 1 remote Gitea repo flagged as unswept (
disaster-recovery), plus 1 external Uptime Kuma admin UI not inspected (config lives outside the filesystem, referenced only viauptime-kuma-monitoring-plan.md). - P1 (provisioning-time): 23 rows (rows 1-23)
- P2 (service cutover): 10 rows (rows 24-33)
- P3 (app1-bu retirement): 19 rows (rows 34-52)
- Already-stale references (pre-existing, independent of this migration): 6 items (wphost02 present-tense language, 2 undated graveyard entries, 1 dead-IP anchor slug, 1 contradiction within ip-dns-changes.md re: Tony's VPS)
Row counts were verified against the table bodies above (23 + 10 + 19 = 52 total numbered rows, matching the highest row number used).