From a6a32f6914b1a47240c2b33b0a35ea825655ad74 Mon Sep 17 00:00:00 2001 From: root Date: Tue, 15 Sep 2026 12:57:46 -0400 Subject: [PATCH] core-bu warm standby built + proven (6 defects fixed); standby package build receipt (2026-09-15) --- CHANGELOG.md | 25 ++++++- .../core-bu-standby-package-2026-09-15.md | 71 +++++++++++++++++++ 2 files changed, 95 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 0e678f1..a0d040a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,6 +1,29 @@ # itpp-infrastructure — CHANGELOG -## 2026-09-15 — app4 + core-bu Provisioned (Nuremberg) +## 2026-09-15 — core-bu Warm Standby BUILT and proven (DISARMED; arming is a posture decision) + +- **core-bu (`159.195.204.203`, netcup Nuremberg) is now a working Hermes warm standby.** Every step was executed on the box and verified by reading the result back. +- **Hermes parity, proven not assumed:** `/usr/local/lib/hermes-agent` mirrored from Core. core-bu reports `v0.21.0 (2026.8.31) · local 8dbf07e9 (+1 carried commit)`, Python `3.11.15`, method `git` — identical to Core. A real one-shot agent turn returned `STANDBY-SMOKE-OK`, so the box can serve, not just install. +- **State parity verified against the live box:** `state.db` **140,201,984 bytes on both**, `quick_check=ok`; `sessions` 151/151; cron `jobs.json` **89 jobs / 86 enabled on both**; 99 references; memories incl. `MEMORY.md`. The standby store sits one 11-minute checkpoint behind (14,222 vs 14,262 messages) by design — the failover path re-syncs before it serves. +- **Probe path proven:** core-bu -> Core over Tailscale (`100.71.155.7`) returns `core` / `active` non-interactively. core-bu joined the tailnet as `100.113.119.108` with `RunSSH=false`, because Tailscale SSH intercepts non-interactive key auth and would have broken the probe. +- **All three failover action paths exercised under DRYRUN (gate lifted for the test, restored after):** + - health-bad: `BAD 1 ` -> `BAD 2 ` -> `[dryrun] decision reached: would fence live box and take over (nothing done)`; + - host-down: `[dryrun] decision reached: host-down path would sync from S3 and start the standby (nothing done; reason: Live host not responding to pings)`; + - failback: `[dryrun] failback would notify and stop this box's gateway (nothing done)`, with a decoy process proving DRYRUN does not kill and the real kill command does. +- **Alert channels verified:** Telegram `getMe` -> `ok=True shonuff_is_a_bot`; SMTP login OK on port **2525**. +- **Six defects found and fixed during the build:** + 1. **No outbound `itpp-infra` private key on core-bu** — the health probe could not reach Core at all. The provisioning record conflates the *inbound* `authorized_keys` entry with an outbound key. Installed; fingerprint `SHA256:Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ` (matches Core). + 2. **`DRYRUN=1` was not a global no-op** (three separate action paths). The host-down path fell through to the S3 sync and the gateway start; the failback path notified and ran its kill. A "safe" dry-run against an unreachable Core would have started a **second poller on the same bot token**. All paths now guarded. + 3. **The failover's final step could not have worked.** `hermes gateway start` requires an installed unit and exits 1 without one. Unit installed with `--no-start-now --no-start-on-login` (present, `disabled`, `inactive`, 0 processes). + 4. **SMTP port was wrong for netcup.** Shipped `SMTP_PORT=587`; netcup blocks outbound 25/465/587 and 587 **times out** from core-bu, so email alerts would have failed silently forever. Now `2525`. app1-bu is on Hetzner and works on both, so it was left alone. + 5. **The failover sync would clobber the standby's own scripts.** Core's `~/.hermes/scripts/` is inside the synced tree, so `live/scripts/` holds **app1-bu's** copies (4,429-byte watchdog vs core-bu's host-specific 19,852-byte one). A wholesale failover sync would have overwritten `PROBE_HOST` / `PROBE_SSH_KEY` / the disarm gate at the worst moment. Added `--exclude "scripts/hermes-standby-*"`. + 6. **`app1-bu` — the previously armed standby — had no `/root/.config/systemd/user/` at all** while its watchdog also calls `hermes gateway start`. **Its failover could not complete its final step either.** Unit installed there too; its email path independently verified good on both ports. +- **Disarm gate verified three ways:** watchdog logs the notice once per day then exits 0; sync logs `DISARMED, sync skipped`; boot restore exits before starting anything. +- **Not armed, deliberately.** Arming core-bu requires disarming app1-bu first (the single-armed rule is enforced in code by `DISARM_FILE`), and it **trades provider diversity away**: Core and core-bu are both netcup, while app1-bu is the only non-netcup box. That is a posture decision for Germaine, not a build step. Stage C (live failover + failback drill, which fences Core's real gateway) also remains, and needs a scheduled window. +- **Bloat note:** the wholesale pull brought 2.9 GB of Core's `docker/` and 1.1 GB of `data/` that core-bu does not run; left in place (472 GB free) but the failover sync should probably exclude them. +- Full build receipt and evidence: `docs/infrastructure/core-bu-standby-package-2026-09-15.md` section 9. + +2026-09-15 — app4 + core-bu Provisioned (Nuremberg) - **Renamed by Germaine (recorded at rename time per the changelog mandate):** the two netcup Nuremberg boxes ordered 2026-09-14 are now **app4** and **core-bu**. - **core-bu** = `v2202609377162521279.megasrv.de`, `159.195.204.203/22`, IPv6 `2a0a:4cc0:c2:b3e0:9409:42ff:fe3b:5397` — netcup **RS 2000 G12** (8 vCPU / 15 GB / 503 GB), exact twin of Core. Role: **Core's warm standby**. diff --git a/docs/infrastructure/core-bu-standby-package-2026-09-15.md b/docs/infrastructure/core-bu-standby-package-2026-09-15.md index 5395c52..4110602 100644 --- a/docs/infrastructure/core-bu-standby-package-2026-09-15.md +++ b/docs/infrastructure/core-bu-standby-package-2026-09-15.md @@ -478,3 +478,74 @@ retired per the P3 checklist in `reference-update-matrix-2026-09-15.md` (Option Nothing was installed on core-bu, app1-bu, or Core. No systemd unit was created or modified anywhere. No production process was started, stopped, or restarted. + +--- + +## 9. Build receipt — executed 2026-09-15 (Sho'Nuff) + +**Status: BUILT + PROVEN + DISARMED (armed-ready).** Every step below was executed on +`core-bu` and verified by reading the result back, not by assuming the step worked. + +### 9.1 What was actually done + +| Step | Evidence | +|---|---| +| Tailscale joined | node `core-bu` = `100.113.119.108`; `RunSSH=false` (Tailscale SSH intercepts non-interactive key auth) | +| Outbound SSH key installed | `/root/.ssh/itpp-infra`, fingerprint `SHA256:Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ` (identical to Core) | +| Probe path proven | core-bu -> Core over tailnet returns `core` / `active` non-interactively | +| Hermes installed | `/usr/local/lib/hermes-agent` mirrored from Core; `v0.21.0 (2026.8.31) · local 8dbf07e9 (+1 carried commit)`, Python `3.11.15`, install method `git` | +| Functional smoke test | real one-shot turn returned `STANDBY-SMOKE-OK` (session `20260915_123810_6d5834`) | +| Config/state staged | `.env` + `config.yaml` present (`600`); Telegram token sha256 matches Core (`ba939adaf8354e72...`) | +| Scripts staged | watchdog / sync / restore at `/root/.hermes/scripts/`, mode `700`, hashes matched source | +| Gateway unit | `/root/.config/systemd/user/hermes gateway.service` — `enabled=disabled`, `active=inactive`, 0 processes | +| Crons | `*/5` watchdog + `*/10` sync; pre-existing `root-essentials-backup` preserved | +| Boot unit | `hermes-standby.service` enabled (inert while disarmed) | +| Linger | `Linger=yes` | +| Alert channels | Telegram `getMe` -> `ok=True shonuff_is_a_bot`; SMTP login OK on port **2525** | + +### 9.2 Proof of the state machine (dry-run, zero production risk) + +- Forced definitive-bad twice: `DEGRADED (1/2, definitive=1)` -> `DEGRADED (2/2, definitive=1)` + -> `[dryrun] decision reached: would fence live box and take over (nothing done)`. +- State file carried all three fields (`BAD `) — the "missing field + resets the counter forever" trap is absent. +- Disarm gate proven three ways: watchdog logs the notice once per day then exits 0; sync logs + `DISARMED, sync skipped`; boot restore exits before starting anything. + +### 9.3 Defects found during the build (all fixed and re-verified) + +1. **No outbound `itpp-infra` private key on core-bu.** The health probe could not reach Core + at all (`Permission denied (publickey)`). Section 1's provisioning claim covers the + *inbound* `authorized_keys` entry, not an outbound key — the two were conflated. +2. **`DRYRUN=1` was not a global no-op.** The health-bad path returned early, but the host-down + path fell through to the S3 sync and the gateway start. A "safe" dry-run test against an + unreachable Core would have started a **second poller on the same bot token**. Fixed with a + guard on the action block; proven by forcing the host-down path to its dry-run stop. +3. **The failover's final step could not have worked.** `hermes gateway start` requires an installed unit + (`_require_service_installed` exits 1 otherwise), and core-bu had none. +4. **SMTP port was wrong for netcup.** The watchdog shipped `SMTP_PORT=587`; netcup blocks + outbound 25/465/587 and 587 **times out** from core-bu, so every email alert would have + failed silently. Changed to `2525` (verified: `LOGIN OK (starttls=True)`). app1-bu is on + Hetzner and works on both ports, so it was left alone. +5. **The failover sync would clobber the standby's own scripts.** Core's `~/.hermes/scripts/` + sits inside the synced tree, so `live/scripts/` contains **app1-bu's** copies + (a 4,429-byte watchdog vs core-bu's host-specific 18 KB one). A wholesale failover sync + would have overwritten core-bu's `PROBE_HOST` / `PROBE_SSH_KEY` / disarm gate at the worst + possible moment. Added `--exclude "scripts/hermes-standby-*"` to the failover sync. + +### 9.4 Same defect class on the previously armed standby (FIXED 2026-09-15) + +`app1-bu` had **no** `/root/.config/systemd/user/` directory at all, while its watchdog calls +`hermes gateway start`. Its failover therefore could not complete its final step either. Unit installed there +with `--no-start-now --no-start-on-login`; now `enabled=disabled`, `active=inactive`, 0 +processes. Its email path was independently verified good (both 587 and 2525). + +### 9.5 Not done, deliberately + +- **Arming is not done.** Arming core-bu requires disarming app1-bu first (the single-armed + rule is enforced in code by `DISARM_FILE`), and it trades away **provider diversity**: + Core and core-bu are both netcup, while app1-bu is the only non-netcup box. That is a + posture decision, not a build step. +- **Stage C (live failover + failback drill) is not done.** It fences Core's real gateway, + so it needs a scheduled window and explicit approval. +- Vaultwarden entry and `key-inventory.md` update still pending.