docs(infra): app4 + core-bu provisioning, migration plan, verified inventory

- migration-plan-app4-core-bu-2026-09-15.md: 8-phase plan (Nuremberg decision,
  provider-diversity gap, acceptance criteria, rollback, DNS/Caddy checklist)
- core-service-inventory-2026-09-15: verified Core inventory, ~30 customer-facing
  services (the Aug 15 plan listed 5), 3 DocuSeal instances, TimeTrex Postgres,
  dead Caddy routes
- reference-update-matrix-2026-09-15: 52 artifacts that name a host
- fix naming collision: 6 files called the Hetzner box core-bu, the name core-bu
  now claims; app1-bu = 5.161.225.131, core-bu = 159.195.204.203 (netcup Nuremberg)
- correct the false provider-diversity claim (the standby is now netcup too)
- supersede app4-migration-plan.md (wrong region reported, silent on core-bu)
This commit is contained in:
ShoNuff
2026-09-15 10:19:46 -04:00
parent 199baadedc
commit cac3cde372
11 changed files with 1204 additions and 8 deletions
@@ -0,0 +1,127 @@
# Standby Host Replacement: app1-bu Cost and Location Review
**Date:** 2026-09-14
**Status:** Recommendation, awaiting decision
**Author:** Sho'Nuff
**Source of truth:** Hetzner invoice 086001081978 (2026-08-27, period 07/2026) + live Hetzner Cloud API pull, 2026-09-14
## Verdict (BLUF)
The DR plan records app1-bu at roughly $14/mo. That is wrong. The real figure is **EUR 31.99/mo for the CPX21 alone, EUR 37.06/mo all in** (about USD 40.39). The fix is not a new provider. It is the **same box in a different location**: Hetzner Falkenstein or Nuremberg prices the identical CPX21 at **EUR 9.49/mo**.
**Recommendation: move the standby from Ashburn to fsn1 or nbg1. Saves EUR 22.50/mo (USD 294/yr) and closes a real diversity gap at the same time.**
## What the invoice actually shows
Invoice 086001081978, period 07/2026, net EUR 114.44, 0% VAT.
It is not one server. The July bill carried eleven servers mid-teardown plus supporting resources:
| Line | Item | Qty | Unit | EUR |
|---|---|---|---|---|
| 1 | Backup (20% of instance price) | 3 | | 7.23 |
| 2 | CPX11 | 1 hr | 0.0280 | 0.03 |
| 3 | CPX11 | 4 x 1,368 hr | 0.0096 | 13.13 |
| 4 | CPX21 | 3 x 1,069 hr | 0.0192 | 20.52 |
| 5 | CPX21 | 2 x 571 hr | 0.0513 | 29.29 |
| 6 | CPX21 | 1 month | 11.99 | 11.99 |
| 7 | CPX41 | 1 x 359 hr | 0.0625 | 22.44 |
| 8 | Primary IPv4 | 11 x 3,368 hr | 0.0008 | 2.69 |
| 9 | Primary IPv4 | 1 month | 0.5000 | 0.50 |
| 10 | Snapshot | 462.289 GB-mo | 0.0143 | 6.61 |
| 11-13 | Traffic | | | 0.00 |
The EUR 114.44 is a July artifact, not today's run rate. Most of those servers no longer exist.
Note line 5: two CPX21 instances at **EUR 0.0513/hr**. That is the Ashburn hourly rate, and it is the same rate the API reports today. Line 4 at EUR 0.0192/hr is the EU rate. The 3.5x US premium is visible inside a single invoice.
## What is actually live today
Live API pull, 2026-09-14. Exactly **one** server remains:
- **app1-bu.itpropartner.com**, id 151357515, CPX21, 3 vCPU / 4 GB / 80 GB, location **ash** (Ashburn, VA), status running, created 2026-07-15, public IPv4 5.161.225.131, backups **disabled**
- 2 primary IPs (IPv4 + IPv6), both assigned to that server
- 9 snapshots, 319.2 GB total
- 0 volumes, 0 firewalls
### Current monthly run rate
| Component | EUR/mo |
|---|---|
| CPX21 in Ashburn | 31.99 |
| Primary IPv4 | 0.50 |
| Snapshot storage (319.2 GB at 0.0143) | 4.57 |
| **Total** | **37.06** (~USD 40.39) |
### Same stack in an EU location
| Component | EUR/mo |
|---|---|
| CPX21 in fsn1 or nbg1 | 9.49 |
| Primary IPv4 | 0.50 |
| Snapshot storage | 4.57 |
| **Total** | **14.56** (~USD 15.87) |
**Delta: EUR 22.50/mo, EUR 270.00/yr, about USD 294/yr.**
## The diversity finding
The rule is that a netcup outage must not take out both Core and its standby. Today both sit in the same US East corridor: Core in Manassas, VA and app1-bu in Ashburn, VA. That is roughly 150 miles and one weather system, one grid region, one set of upstream transit providers.
Provider diversity is satisfied. Regional diversity is not.
Falkenstein or Nuremberg fixes it. The standby would move from the same corridor as production to a separate continent, which is what a warm standby is supposed to be.
## Options
**Option A: Move the standby to Hetzner fsn1 or nbg1 (RECOMMENDED)**
- Cost: EUR 14.56/mo all in, down from EUR 37.06
- Keeps the entire runbook: Hetzner API, rescue mode with key injection, snapshot tooling, S3 restore path
- Improves regional diversity from 150 miles to 4,000
- Latency Core to standby rises from roughly 10ms to roughly 95ms. The 10-minute failover poll does not care, and a 1.17 GB DB sync at that latency is still tens of seconds
- Reversibility: high. This is a rebuild from the standby deploy script plus an S3 restore. Nine snapshots and the S3 backup chain remain as the safety net throughout
- Risk: low. Same provider, same tooling, same image lineage
**Option B: Move to a US provider with monthly billing and a real API (DigitalOcean 2 GiB verified at USD 12.00/mo)**
- Keeps latency low and adds a genuinely separate provider on top of netcup
- Breaks the API-driven recovery runbook, which is Hetzner-specific. Rescue mode, key injection, and scripted power control would all need rewriting and re-testing
- Price per resource is worse than Hetzner EU: USD 12.00 for 2 GiB / 1 vCPU against EUR 9.49 for 4 GiB / 3 vCPU
- Vultr and Akamai/Linode price per region, so a specific number requires a per-region pull
**Option C: Stay in Ashburn at EUR 31.99**
- Rejected. Paying 3.4x for the same instance, in the same corridor as production
**Rejected: SSD Nodes prepaid**
- 14-day refund window then a 1 to 3 year commitment, an overselling reputation, and a thin API. Wrong profile for the asset that has to work when everything else is down
## Dead storage on the account
Snapshot review, same API pull:
- 3 wphost02 snapshots, **151.6 GB, EUR 2.17/mo**, for a host decommissioned 2026-08-28
- 1 legacy snapshot "Back-online" (2026-06-05), **112.1 GB, EUR 1.60/mo**, superseded by the current weekly chain
- 1 legacy snapshot "migration-complete" (2025-11-23), 7.6 GB, EUR 0.11/mo
- 4 current app1-bu weeklies, 48.0 GB, EUR 0.69/mo, KEEP
Removing the decommissioned and superseded snapshots saves **EUR 3.88/mo, EUR 46.55/yr**. Not proposed for silent deletion. Confirm first, per standing rule.
## Combined opportunity
| Action | EUR/mo | USD/yr |
|---|---|---|
| Relocate standby to EU | 22.50 | 294 |
| Retire dead snapshots | 3.88 | 51 |
| **Total** | **26.38** | **345** |
## Record correction required
The DR plan records the app1-bu line as a CPX21 at roughly $14/mo. Both halves are wrong:
1. **Price:** the Ashburn CPX21 is EUR 31.99/mo, not about $14
2. **Spec:** CPX21 is 3 vCPU / 4 GB / 80 GB, not 4C/8G
Corrections should land in the DR plan v2 and any derived runbook once the location decision is made, so the standby budget is not understated again.
## Rebuild procedure
Reference the `hermes-standby-deployment` skill. Sequence: create CPX21 in fsn1 or nbg1, run the standby deploy script, restore from s3://hermes-vps-backups/live/, verify the failover cron and the 10-minute health check, then destroy the Ashburn instance. Keep the Ashburn box alive until the EU standby passes a restore test. Backed up is not finished; a confirmed restore test is.