# migration-plan-app4-core-bu-2026-09-15.md **Owner:** IT Pro Partner (Germaine Brown) **Created:** 2026-09-15 **Status:** ACTIVE (supersedes the draft `app4-migration-plan.md` of 2026-08-15; that file is kept for history) **Scope:** (a) move Core's customer-facing services onto the new `app4`; (b) stand up `core-bu` as Core's warm standby with a working failover **and failback**; (c) update every document, record and reference that names hosts, IPs or service locations. **Status as of 2026-09-15 (post-Phase-3-start):** This plan's original Section 4 phase lists undercounted scope — the verified inventory (`core-service-inventory-2026-09-15.md`) found roughly 30 customer-facing services plus 3 separate DocuSeal instances, not the ~5 implied below. See Sections 10-12 (appended) for the corrected scope, backup gaps, and defects found. **Actual migration progress to date: `microbin` has been staged on `app4` and verified working at the application layer (container up, served content matches Core); the DNS flip for microbin has NOT happened yet — it is still served live from Core.** No other service has moved. Treat Section 4's phase list as directional; Section 10 is the authoritative scope. --- ## 1. What is verified today (2026-09-15) Everything in this section was measured on the live boxes, not copied from a doc. ### 1.1 The two new boxes | | **app4** | **core-bu** | | --- | --- | --- | | Role | Core's customer-facing services | Core's warm standby | | Hostname | `v2202609377162521278.quicksrv.de` | `v2202609377162521279.megasrv.de` | | IPv4 | `159.195.205.80/22` | `159.195.204.203/22` | | IPv6 | `2a0a:4cc0:c2:bcbf:34b3:8cff:fea2:2892` | `2a0a:4cc0:c2:b3e0:9409:42ff:fe3b:5397` | | Model | netcup RS 4000 G12, 12 vCPU / 32 GB / 1007 GB | netcup RS 2000 G12, 8 vCPU / 16 GB / 503 GB | | Location | Nuremberg (NBG) | Nuremberg (NBG) | | RTT from Core | 100.5 ms | 100.5 ms | `core-bu` is the exact twin of Core (same 8 vCPU / 16 GB / 503 GB shape), which is what a standby should be. ### 1.2 Provisioning status: COMPLETE and verified Both boxes were provisioned to the ITPP standard on 2026-09-15 and each item was verified, not asserted: - Debian 13, hostname set, timezone `America/New_York`, 8 GB swap (9 GB on app4). - `ippadmin` user with NOPASSWD sudo; `itpp-infra` key installed for **both** `root` and `ippadmin`. - `ufw` **active** (22/80/443, plus 9100 only from Core and from the tailnet). - `fail2ban` active, `unattended-upgrades` active. - Docker CE **29.8.0** + Compose **v5.5.1** (upstream repo, matching app1's `docker-compose-plugin 5.3.1~trixie` family). - `node_exporter` listening on `:9100` - HTTP 200 from Core, **closed** from a third host. - `awscli` + Wasabi credentials in `/root/.aws/credentials` (mode 600). - sshd hardened to the fleet convention: `PermitRootLogin without-password`, `PasswordAuthentication no`, `AllowUsers ippadmin root`. Verified four ways per box (root key login, ippadmin key login + `sudo -n`, password auth refused, non-allowlisted user refused) with the self-reverting lockout guard armed. ### 1.3 Backups: enrolled and restore-tested | Host | Script | Schedule | Destination | Evidence | | --- | --- | --- | --- | --- | | app4 | `root-essentials-backup.sh` | 04:45 ET | `s3://hermes-vps-backups/root-backup/app4/` | first run: upload + in-script download/extract verify OK | | core-bu | `root-essentials-backup.sh` | 05:15 ET | `s3://hermes-vps-backups/root-backup/core-bu/` | first run: upload + in-script download/extract verify OK | Note the known gap this replicates: `root-essentials-backup.sh` **excludes `*.db` by design**. Any database these boxes end up hosting needs its own `sqlite3 .backup` / `pg_dump` job, exactly as `anita-mnz` needed `hermes-db-backup.sh`. This is an explicit Phase 4/5 acceptance item, not an optional nicety. ### 1.4 Monitoring: registered A `node_exporter` job was added to the **live** Prometheus config (`/root/docker/monitoring/prometheus/prometheus.yml`) - the job did not previously exist. `up=1` verified for `core:9100`, `app4:9100`, `core-bu:9100`. **Finding (pre-existing, not caused by this work):** `node_exporter` is **not running** on app1, app2, app3 or app1-bu; and the node_exporter target list in `/opt/prometheus/prometheus.yml` lives in a file Prometheus never loaded (it still names decommissioned `wphost02` and `178.156.131.57`). Host metrics for the existing fleet were therefore never collected. Tracked in Section 8. ### 1.5 Credentials Console/root credentials for both boxes are recorded in `/root/.hermes/references/new-servers-2026-09-15.md` (mode 600, root only). Password auth is disabled on both boxes, so those passwords are **console/rescue only**. Vaultwarden was **locked** at the time of writing, so the Vaultwarden entries are still owed (Section 9, Q3). --- ## 2. End state | Host | Role after migration | | --- | --- | | **Core** (152.53.192.33, RS 2000, Manassas) | Hermes + its direct dependencies (browserless, camofox-browser, SearXNG, Super Search MCP), Prometheus/Telegraf/Grafana, mikrotik-exporter, Caddy for Core-local routes. **No customer-facing apps.** | | **app4** (159.195.205.80, RS 4000 G12, Nuremberg) | All customer-facing apps, their databases, and all customer-facing Caddy routes + TLS. | | **core-bu** (159.195.204.203, RS 2000 G12, Nuremberg) | Warm standby for Core (Hermes, its state, its watchdog). Dormant until failover. | | **app1-bu** (5.161.225.131, Hetzner CPX21, Ashburn) | Retirement candidate once `core-bu` is proven. **Only remaining non-netcup box** (Section 3.2). | | **anita-mnz** (159.195.16.30, netcup, Manassas) | Unchanged. Anita's dedicated Hermes box. | | **app1 / app2 / app3** (Manassas) | Unchanged by this plan. | --- ## 3. Deviations and risks you must decide on ### 3.1 app4 is in Nuremberg, not Manassas (NEW, material) Measured: **100.5 ms RTT Core -> app4**, versus 0.5 ms Core -> app2 (Manassas) and 1.6 ms -> app1-bu (Ashburn). The old draft assumed Manassas. Consequences: - **Customer latency on app4-hosted sites.** Typical US East users add roughly 80-110 ms per round trip versus a Manassas host. For static sites this is mostly invisible; for interactive apps (DocuSeal signing flow, Ops Portal, TimeTrex) it is user-visible. - **Core <-> app4 chatter crosses the Atlantic.** Any Core->app4 API call, monitoring scrape, backup pull, or Caddy proxy hit pays ~100 ms. This is acceptable if app4 is self-contained, and painful if the two are chatty. Design rule for this migration: **app4 must not depend on Core at request time.** - **Benefit, and it is real:** Core (US) and its standby (EU) now fail independently. A Nuremberg outage does not touch Core, and a Manassas outage does not touch the standby. The old pair (Core + app1-bu Ashburn) were 1.6 ms apart and shared the US East corridor. **Options:** (A) accept Nuremberg and design app4 to be self-contained (recommended, zero cost, boxes are paid); (B) re-order app4 as a Manassas RS 4000 and keep the Nuremberg box as the standby. This is a money decision, so it is yours. ### 3.2 Provider diversity is now unmet Core, app1, app2, app3, app4, core-bu and anita-mnz are **all netcup**. `app1-bu` (Hetzner) is the only other provider, and this plan retires it. Mitigation options: keep app1-bu as the *provider-diverse* last-resort standby even after core-bu is primary (cheapest option, EUR 31.99/mo), or move the off-site backup/DR device to a non-netcup provider. **This must be decided before app1-bu is deleted**, and the DR principle that has governed the org so far ("a netcup outage must not kill both live and standby") is currently **satisfied by geography but not by provider**. ### 3.3 app4 has no standby of its own `app4` becomes the single host for every customer-facing service. If it dies, customer apps are down until S3 restore. `core-bu` is shaped for Core, not for the customer tier (16 GB, and it is meant to be dormant). Options: (A) accept S3-restore RTO for app4; (B) let core-bu carry a cold/secondary copy of app4's data; (C) budget a second app-tier box. Recommend a decision **now**, because it changes what core-bu should replicate. ### 3.4 DNS authority is split Confirmed by the old plan's own checklist and this project's history: `itpropartner.com` is on **SiteGround nameservers (manual panel, no API)**; `fleettracker360.com` and `voipsimplicity.com` are on **Cloudflare**; `iamgmb.com`, `intelsight.io`, `debtrecoveryexperts.com` need per-domain `dig NS` verification in Phase 0. Every cutover record must be changed in the correct panel or it is a silent no-op. Section 6 lists the records. --- ## 4. Migration phases Each phase has a gate: **the next phase does not start until the gate's evidence exists.** ### Phase 0 - Inventory and recon (Core, read-only) - IN PROGRESS Deliverable: `docs/infrastructure/core-service-inventory-2026-09-15.md` (+ `.json`) - every container, unit, port, volume, database, cron job, TLS cert and Caddy route on Core, with sizes and dependencies. Gate: inventory lists every Caddy site block and its upstream, and explicitly resolves two conflicts that existing docs disagree on: 1. **DocuSeal** is recorded on **Core** by the Aug 15 draft but on **App1** by `backup-plan.md` (4:00 AM job). 2. **SearXNG** likewise. Only the live `docker ps` / Caddyfile settles it. ### Phase 1 - Provision app4 + core-bu, monitoring first - **COMPLETE (2026-09-15)** See Section 1. Gate met: both boxes verified; backups restore-tested; both scraped by Prometheus; no customer app touched. ### Phase 2 - Access and naming (needs your input) - Enroll both boxes in Tailscale (needs a reusable auth key or your approval of the login URL - Section 9, Q1). - Decide DNS names: `app4.itpropartner.com` and `core-bu.itpropartner.com` A/AAAA records, added in the correct panel (SiteGround for `itpropartner.com`). Internal access and monitoring already work **by IP**, so this is not blocking, but the docs and the recovery manual read better with names. - Install Caddy on app4 with `default_bind 159.195.205.80` (avoids the Tailscale :443 conflict). Gate: `tailscale status` shows both nodes; name resolution works from Core. ### Phase 3 - Prove the pattern on low-risk apps - Move **microbin** (`127.0.0.1:8260`) first: single container, one volume, no database. - Move **Uptime Kuma** second: it is the monitoring tool, so it must be moved carefully and its own downtime window announced. - For each: stop on Core, rsync the volume, start on app4, verify side-by-side with `curl --resolve :443:159.195.205.80`, then flip DNS, then soak 24 h. - This phase validates the runbook (per-service steps, verification and rollback) before any customer app moves. Gate: microbin and Uptime Kuma both served from app4 with app4-issued TLS, verified externally, and their S3 backups land from app4 - not from Core - with a restore test on at least one. ### Phase 4 - Data foundation + first real app - Provision Postgres and Redis **fresh** on app4 (internal-only binds, least privilege, no public 5432/6379). - Move the **Ops Portal backend** (`:8090`), then **DocuSeal** (SQLite + attachments + its internal Redis), then **TimeTrex** (Postgres-backed). - Databases: `pg_dump -Fc` per database, restore on app4, then **compare row counts per major table**, not a spot check. SQLite: `sqlite3 .backup`, never `cp`. - **New backup jobs on app4 for every database it now hosts** (`*.db` is excluded from the essentials archive). Gate: row counts match; `curl --resolve` responses match Core; app4-backup + restore test for each moved DB. ### Phase 5 - Customer sites and the voice stack - rsync every static customer site root (`*.iamgmb.com`, `*.intelsight.io`, `*.fleettracker360.com`, `*.debtrecoveryexperts.com`) to app4; `caddy validate` the app4 config; pre-issue TLS. - Voice stack (`voice.*`, `voice-open.*`): enumerate Twilio webhooks and any external endpoints in Phase 0 and update them **before** the DNS flip, or calls break after cutover. Gate: every moving domain answers from app4 with a valid cert; voice end-to-end call tested. ### Phase 6 - DNS cutover, soak, decommission on Core - Lower TTL to 60-300 on every moving record **24 h before** the flip (correct panel per domain). - Flip one domain at a time, low traffic first, verifying each (`dig +short @1.1.1.1`, then `curl -sI`). - Keep Core's Caddy blocks as a 301 redirect to app4 during a 24-72 h soak; remove with targeted edits and the caddy-audit hook (never a whole-file rewrite). - Then: stop/remove the moved containers on Core, retain volumes + images **30 days** as rollback, decommission the customer schemas in Core's Postgres/Redis. Gate: 72 h soak with no rollback; Core runs zero customer-facing apps; rollback path still intact. ### Phase 7 - core-bu standby, failover AND failback proven (parallel with 3-6) - Install the standby package (sync + watchdog with a health-based decision branch, fence-before-takeover, and **automatic failback**, which the current app1-bu scripts do not have). - **Only one standby may be armed at a time.** Disarm app1-bu before arming core-bu, or a Core hiccup makes both answer as the same Telegram bot. - Prove it with a real, announced test: failover, then failback, then confirm the standby is dormant again. Gate: failover and failback both demonstrated with evidence, and app1-bu verifiably disarmed. ### Phase 8 - Documentation and reference sweep Deliverable: `docs/infrastructure/reference-update-matrix-2026-09-15.md` - every artifact that names a host. Includes at minimum: `key-inventory.md` (done), `backup-plan.md` (done), `CHANGELOG.md` (done), `app-inventory.csv`, `server-architecture-plan`, `server-dr-plans`, `dr-issue-log`, the recovery manual, `decommissioned-hosts.json` + `stale-reference-verify.py` (add `app1-bu` when retired), Prometheus config (live one, **and delete/refresh the dead `/opt/prometheus/prometheus.yml`**), Grafana dashboards, Uptime Kuma monitors, `health-master-watchdog.py`, Hermes cron `jobs.json` live-config fields, Hudu assets, the ops portal, client-facing runbooks, and any skill that hardcodes a host or IP. Gate: `stale-reference-verify.py` and `doc-live-verify.py` both clean; every doc cites the new IPs. --- ## 5. Acceptance criteria (whole project) 1. Core hosts no customer-facing app; every moved domain answers from app4 with a valid TLS cert. 2. Every moved service has: a data migration that was verified by counts/sizes, a health check, and a tested rollback. 3. Every database on app4 has its own backup job with a **performed restore test** (a green cron entry is not evidence). 4. `core-bu` failover **and** failback both demonstrated; exactly one standby armed at any time. 5. Documentation matrix closed out: no live surface names a decommissioned host or a stale IP. 6. `app1-bu` is either retired (with the provider-diversity decision recorded) or explicitly retained as the provider-diverse standby. --- ## 6. DNS and Caddy change checklist - [ ] `dig NS` every moving domain; record the authoritative panel in Phase 0. - [ ] Pre-write all moving site blocks into app4's Caddyfile; `caddy validate`; pre-issue certs. - [ ] `default_bind 159.195.205.80` in app4's Caddy global block. - [ ] TTL 60-300 at least 24 h before each flip. - [ ] Flip per domain; verify `dig +short @1.1.1.1` and `curl -sI https://`. - [ ] Keep Core blocks as 301s for the soak window; then targeted removal + caddy-audit hook. - [ ] Update Http->Https and any `CNAME`/`www` records in the **same** panel as the A record. --- ## 7. Rollback - Before each phase: snapshot DNS records, Core Caddyfile, Core `docker ps`/volume list. - Phases 3-5: stop on app4, flip DNS back to Core, restart the Core container. Core volumes are untouched. - Phase 6: with low TTL, the flip back propagates in minutes; Core blocks are retained during soak. - Data: Core volumes/images retained 30 days. After that, restore from app4's S3 backups (which is why the Phase 1/4 restore tests are mandatory). - `core-bu`: failback is part of the design, not an afterthought; the standby stands down on its own. --- ## 8. Follow-up findings raised by this work (not fixed here) | # | Finding | Impact | Owner | | --- | --- | --- | --- | | 1 | `node_exporter` not running on app1, app2, app3, app1-bu | No host metrics for the fleet | This project (Phase 8) | | 2 | `/opt/prometheus/prometheus.yml` is a dead file containing decommissioned hosts (wphost02, 178.156.131.57) and was never loaded | Misleading; wasted trust | Phase 8 | | 3 | netcup SCP/CCP API auth returns HTTP 500 / 404 (worked in July) | Provisioning automation via API is dead | Separate | | 4 | `hermes-standby-sync.sh` has a ping-based failback flaw and no failback logic | Standby reliability | Phase 7 | | 5 | `hermes-snapshot.sh:34` does a live `VACUUM INTO` | Store churn | Separate | --- ## 9. Open decisions (needed from Germaine) **Q1 - Tailscale:** add both boxes to the tailnet. Need either a reusable auth key, or approve the login URL from each box. *Update 2026-09-15:* still open, plus a new cleanup item — **two Tailscale nodes exist for one physical box** (`app1-bu`, offline 61 days, and `app1-bu-1`, idle). Decide whether to retire the stale `app1-bu` node now or wait until Q4 (app1-bu retirement) is decided; either way one of the two entries is dead weight today. See Section 12.4. **Q2 - Nuremberg vs Manassas for app4:** accept Nuremberg (design app4 self-contained) or re-order app4 in Manassas and repurpose the Nuremberg box? Section 3.1. *Update 2026-09-15:* still open and now more consequential — the corrected scope (Section 10) is ~30 services, not ~5, so the Atlantic-RTT cost of a wrong answer here is paid by far more customer-facing traffic than originally assessed. **Q3 - Vaultwarden:** the CLI is locked. Unlock it (or tell me when) and I will file the two new server items. *Update 2026-09-15:* still open, no change. **Q4 - app1-bu:** retire it, or keep it as the provider-diverse standby? Section 3.2 - this is the only remaining non-netcup box. *Update 2026-09-15:* still open. Note the duplicate-Tailscale-node finding in Q1/Section 12.4 is directly tied to this decision — retiring app1-bu should also retire its stale tailnet entry in the same action. **Q5 - app4 standby scope:** accept S3-restore RTO, or should core-bu carry a secondary copy? Section 3.3. *Update 2026-09-15:* still open, and now higher-stakes: with ~30 services and ~17 of them carrying zero backup coverage today (Section 11), an app4-only failure mode before those gaps are closed would be a real data-loss event, not just an availability gap. Recommend resolving Section 11's gaps before finalizing this answer. --- ## 10. Corrected scope and migration waves (added 2026-09-15, post-inventory) The counts above (Sections 1-9) predate `core-service-inventory-2026-09-15.md`. That inventory is now the source of truth for scope. It found: - **13 Docker containers** on Core, including **three separate DocuSeal instances** (`docuseal` :8091, `docuseal-dre` :8094, `docuseal-modelortho` :8092 — each its own bind-mounted `./data` dir and its own `.env`; none of this is the single "DocuSeal" the original plan assumed), plus TimeTrex, microbin, Uptime Kuma, Grafana, Prometheus, Telegraf, mikrotik-exporter, SearXNG, browserless, camofox-browser. - **~30 additional customer-facing systemd services** the original plan never named: `auth-api`, `diglocate-api`, `dre-mcp`, `dre-portal`, `ft360-mcp`, `hermes-voice`, `hotnow-api`, `intelsight-api`, `ops-portal`, `osint-api`, `outlook-upload`, `pipeline-api`, `pry`, `pta-registration`, `rally`, `seemytrip`, `shark-game`, `shopping-cart`, `survey-registration`, `transitpin`, `twilio-mcp`, `verdicttank-api`, `verdicttank-worker`, `voice-agent`, `voice-agent-stt` — plus two disabled-but-present units (`crawl4ai`, `hermes-control-deck`) that need an explicit keep/remove decision, not a silent carry-over. - Two independent Postgres engines (host 17.10 for `hotnow`; TimeTrex's own containerized 16), not one shared instance as Section 3/4 originally assumed. - Two stale/dead Caddy routes (`crm.debtrecoveryexperts.com`, `crm.intelsight.io` — both proxy to a `localhost:3003` nothing is listening on) that must be resolved (delete or repoint) before cutover, not moved as-is. **Total real scope: ~30 customer-facing services + 3 DocuSeal instances**, versus the ~5 the phase list in Section 4 was written around. Section 4's phases remain valid as a *pattern* (prove low-risk, then data foundation, then bulk) but its per-phase examples are illustrative, not the full list. This section supersedes that impression with actual waves: ### Wave A — no state, no database (lowest risk, validates the runbook) `microbin` (in progress — see Status line at top), `uptime-kuma`. Matches original Phase 3. ### Wave B — SQLite-only services with a single owner and no cross-service dependency `diglocate-api`, `dre-portal`, `dre-mcp`, `intelsight-api`, `pipeline-api`, `pry`, `pta-registration`, `survey-registration`, `seemytrip`, `shark-game`, `shopping-cart`, `transitpin`, `outlook-upload`, `ft360-mcp`. Rationale: each carries its own SQLite file and no shared Postgres/Redis dependency, so they can move independently and in any order without a shared-data-foundation blocker; doing this wave before Wave C also means the ~14 zero-backup services in this group (Section 11) get backup coverage on app4 sooner. ### Wave C — shared data foundation + its dependents Provision Postgres + Redis fresh on app4, then move `hotnow-api` (owns the `hotnow` Postgres DB + Redis DB1 consumer), `auth-api` (has an existing backup script — lowest-risk of this wave), `rally`, `verdicttank-api` + `verdicttank-worker`, `osint-api`. Rationale: these either own or depend on a database, so they wait for the fresh Postgres/Redis instance the original Phase 4 already calls for, and moving the database owner (`hotnow-api`) first de-risks the rest of the wave. ### Wave D — DocuSeal (all three instances) + TimeTrex `docuseal`, `docuseal-dre`, `docuseal-modelortho` (SQLite + attachments + internal Redis each — move and verify one at a time, they are not interchangeable despite sharing an image), then `timetrex` (its own containerized Postgres 16, needs a container-level `pg_dump`, not a host-level one). Rationale: highest customer-visibility (signing flow, time tracking) and the most complex data migrations — done last, once the runbook is proven on Waves A-C. ### Wave E — voice stack + remaining `voice-agent`, `voice-agent-stt`, `hermes-voice`, `twilio-mcp`. Rationale: unchanged from original Phase 5 — Twilio webhook endpoints must be updated before DNS flip or calls break. Kept last because it is externally integrated (Twilio-side config, not just DNS) and hardest to roll back quickly. ### Explicitly out of scope for this migration (stays on Core) `gitea-runner`, `hermes-assistant`, `hermes-browser`, `hermes-socat-8787`, `host-metrics-exporter`, `node_exporter`, `osint-person`, `snmp-metrics`, `super-search`, `wazuh-agent`, `hermes-gateway`, `caddy` (trimmed to Core-only routes post-cutover), and the two disabled units `crawl4ai` / `hermes-control-deck` pending a keep/remove decision. --- ## 11. Backup coverage gaps (added 2026-09-15, post-inventory) The inventory found **zero backup coverage** for the following services today. Each has its own SQLite DB (or, for HotNow, a Postgres DB + Redis DB1) with no S3 job protecting it. These must get a backup job on app4 as part of whichever wave moves them (Section 10) — do not carry the gap forward: **Update 2026-09-15 (later the same day) - most of this list is now closed.** Two backup jobs were built, run, and proven by download-and-restore after this audit was written: - `app-data-backup.sh` (Core 2:15 AM, app4 5:10 AM) - SQLite snapshots taken through the `.backup` API and `quick_check`ed, plus `pg_dump` of the `hotnow` Postgres DB; the Sunday 2:50 AM `--full` run adds upload/attachment blobs. Proven: 10/10 DBs `quick_check=ok`, valid 52-entry `pg_restore` TOC for `hotnow.dump`, 275 service code files byte-identical on restore. - `docker-stack-defs-backup.sh` (Core 1:40 AM, app4 5:10 AM) - compose files, `.env`, configs and service code for every `/root/docker` stack. Proven: 557 files, all 9 compose files byte-identical on restore. **Now covered:** HotNow (Postgres), IntelSight, Pipeline API, PTA registration, PTA survey, SeeMyTrip (DB + uploads), Shopping Cart (uploads), DRE Portal (DB + uploads), Rally, Ops Portal, VerdictTank `users.db`, DigLocate (uploads), and every `/root/docker` stack definition. **Still not covered:** `osint-api`, `osint-person`, `outlook-upload`, `voice-agent`, `voice-agent-stt`, `ft360-mcp`, `hermes-voice`, `twilio-mcp`, and the HotNow **Redis DB1**. Note `shark-game` *is* covered - it is in the `root-essentials-backup.sh` include list, so its entry below was an undercount. The audit list that follows is retained as the record of what was found at the time. 1. `diglocate-api` — no backup script 2. `hotnow-api` — Postgres `hotnow` DB + Redis DB1, no backup script 3. `intelsight-api` — no backup script (has `intelsight.db`) 4. `osint-api` — no backup script 5. `osint-person` (MCP) — no backup script 6. `outlook-upload` — no backup script 7. `pipeline-api` — no backup script 8. `pry` — no backup script 9. `pta-registration` — no backup script 10. `survey-registration` (pta-survey) — no backup script 11. `seemytrip` — no backup script 12. `shark-game` — no backup script 13. `shopping-cart` — no backup script 14. `dre-portal` — no backup script (`dre.db`) 15. `voice-agent` — no backup script 16. `voice-agent-stt` — no backup script 17. `verdicttank-api` / `verdicttank-worker` — `users.db` has collection cron jobs but no DB backup job (counted as one gap; the two units share the same unbacked database) Partial/unclear coverage worth flagging separately (not counted in the ~17 above, but not solid either): `ft360-mcp` (stats/export scripts only, no DB backup), `rally` (debug/dump scripts only, no scheduled backup), `ops-portal` (not explicitly separated from `core-services-backup.sh` — unconfirmed), `hermes-voice` and `twilio-mcp` (backup status unknown, not verified either way). --- ## 12. Known issues and defects found during Phase 3 (added 2026-09-15) ### 12.1 Fleet provisioning defect: empty Docker data root on both new boxes Both `app4` and `core-bu` shipped with `/var/lib/docker` **empty** despite Docker CE reporting installed and the service active. Every image pull failed with `GetImageBlob` / `overlay2: no such file or directory` (ENOENT) errors. Root cause not yet traced to a specific provisioning step; fixed on both boxes by stopping Docker, `rm -rf /var/lib/docker/*`, and restarting the daemon (which recreates the expected directory structure), then proven with a real `docker run hello-world` completing successfully on both. This is a **fleet-level provisioning defect**, not a one-off — it hit both new boxes identically, so any future box built from the same provisioning process should be checked for this before being marked "Docker ready." ### 12.2 Docs publish pipeline dead for ~1 month The MkDocs publish pipeline that puts these documents on the live site has been non-functional for approximately a month, for three compounding reasons: (a) a dead/expired Gitea access token used by the publish job, (b) a branch-resolution bug in the pipeline script that was silently masked by a trailing `|| true`, so failures never surfaced as failures, and (c) the deploy step pointing at the wrong path. **No Gitea Actions runner exists anywhere in the fleet**, so even with the above three items fixed, automated publish still cannot fire today — publishing currently requires the owner to run it manually. This explains why a plan this stale could stay "published" without anyone noticing the drift. ### 12.3 app4's essentials backup does not cover `/root/docker` `root-essentials-backup.sh` on app4 (Section 1.3) does not include `/root/docker` in its backup paths. Any compose stack migrated onto app4 under `/root/docker/...` (which is where Docker-based services conventionally live per Section 10's Wave A/D services) will **not** be captured by the essentials backup as currently configured. This must be fixed before or immediately after the first Docker-based service lands on app4 — otherwise migrated services are *less* protected on app4 than they were on Core. **Fixed 2026-09-15.** Rather than adding 5.8 GB of mostly-reproducible installs to the essentials tarball, a dedicated job was written: `docker-stack-defs-backup.sh` archives compose files, `.env`, configs and service code for every stack (240 files / 6.4 KB on Core), excluding `venv/`, `node_modules/`, caches and any file >= 1 MiB. It runs on Core at 1:40 AM and on app4 at 5:10 AM, is namespaced per host in S3 (`s3://hermes-vps-backups/docker-stack-defs/-.tar.gz`), and was proven by downloading the archive back and byte-comparing all 9 compose files against live. ### 12.4 Duplicate Tailscale node for one physical box Two Tailscale nodes exist representing the same physical `app1-bu` host: `app1-bu` (offline 61 days) and `app1-bu-1` (idle, presumably re-registered under a new name at some point without removing the old entry). This is a cleanup item tied to the app1-bu retirement decision (Section 9, Q4) — the stale `app1-bu` node should be removed from the tailnet regardless of which way Q4 is decided, since it is dead weight either way. --- ## Appendix A - Verified evidence log (2026-09-15) | Check | Result | | --- | --- | | RTT Core -> app4 / core-bu | 100.5 ms / 100.5 ms | | RTT Core -> app2 (Manassas) / app1-bu (Ashburn) | 0.5 ms / 1.6 ms | | `authorized_keys` (root + ippadmin, both boxes) | present, sha256 `102c80e5...`, 1 line each | | Key login from Core | root OK, ippadmin `sudo -n` -> root, both boxes | | Password auth / non-allowlisted user | refused on both | | `ufw status` | active, 22/80/443 + 9100 from Core/tailnet | | Docker / Compose | 29.8.0 / v5.5.1 both | | `node_exporter` | 200 from Core; refused from app2 (third host) | | Prometheus `up{job="node_exporter"}` | core, app4, core-bu = 1 | | First backup run | upload + download/extract verify OK on both | | Credentials file | `/root/.hermes/references/new-servers-2026-09-15.md`, mode 600 | ---