docs(infra): app4 + core-bu provisioning, migration plan, verified inventory
- migration-plan-app4-core-bu-2026-09-15.md: 8-phase plan (Nuremberg decision, provider-diversity gap, acceptance criteria, rollback, DNS/Caddy checklist) - core-service-inventory-2026-09-15: verified Core inventory, ~30 customer-facing services (the Aug 15 plan listed 5), 3 DocuSeal instances, TimeTrex Postgres, dead Caddy routes - reference-update-matrix-2026-09-15: 52 artifacts that name a host - fix naming collision: 6 files called the Hetzner box core-bu, the name core-bu now claims; app1-bu = 5.161.225.131, core-bu = 159.195.204.203 (netcup Nuremberg) - correct the false provider-diversity claim (the standby is now netcup too) - supersede app4-migration-plan.md (wrong region reported, silent on core-bu)
This commit is contained in:
@@ -14,7 +14,14 @@
|
||||
| **app1** | 152.53.36.131 | RS 4000 G9.5 (12 vCPU EPYC, 32 GB RAM, 1 TB SSD) | netcup | Service hub — AI gateway, CRM, signing, TTS, auth, automation |
|
||||
| **app2** | 152.53.39.202 | RS 4000 G9.5 (12 vCPU EPYC, 32 GB RAM, 1 TB SSD) | netcup | Infrastructure — Gitea, Hudu, Ubiquiti controllers, Traccar, DNS, SIEM |
|
||||
| **app3** | 152.53.241.111 | RS 4000 G9.5 (12 vCPU EPYC, 32 GB RAM, 1 TB SSD) | netcup | Web hosting — CloudPanel CE (WordPress/static/PHP), client sites |
|
||||
| **app1-bu** | 5.161.225.131 | CPX21 (3 vCPU, 4 GB RAM, 80 GB) | Hetzner | Warm standby — provider diversity. Auto-restores from S3. |
|
||||
| **app1-bu** | 5.161.225.131 | CPX21 (3 vCPU, 4 GB RAM, 80 GB) | Hetzner | Warm standby for Core — **the only non-netcup box**. Auto-restores from S3. |
|
||||
| **app4** | 159.195.205.80 | RS 4000 G12 (12 vCPU EPYC, 32 GB RAM, 1 TB NVMe) | netcup | **Customer-facing services** (migration target off Core). Nuremberg. Provisioned 2026-09-15. |
|
||||
| **core-bu** | 159.195.204.203 | RS 2000 G12 (8 vCPU EPYC, 16 GB RAM, 503 GB) | netcup | **Core's warm standby** (new role, Nuremberg). Provisioned 2026-09-15. |
|
||||
|
||||
> **Provider diversity (2026-09-15):** the standby pair is now Core (netcup, Manassas) + **core-bu (netcup, Nuremberg)**. That gives
|
||||
> **regional** separation but **not provider** separation: a netcup-wide incident affects both. `app1-bu` (Hetzner) is the only
|
||||
> non-netcup box; provider diversity therefore depends on retaining it. Open decision — see
|
||||
> `docs/infrastructure/migration-plan-app4-core-bu-2026-09-15.md` Q4.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,3 +1,10 @@
|
||||
> **SUPERSEDED 2026-09-15 — do not use as the live plan.** The boxes are now delivered and provisioned:
|
||||
> **app4 = 159.195.205.80** and **core-bu = 159.195.204.203**, both netcup **Nuremberg**, not Manassas as assumed below.
|
||||
> This file's §3 recommendation ("RS 4000 G12 ... Manassas VA"), its Phase 0-4 ordering, and its silence on `core-bu`
|
||||
> are all out of date. Live plan: `migration-plan-app4-core-bu-2026-09-15.md`. Verified inventory:
|
||||
> `core-service-inventory-2026-09-15.md` (which found ~30 customer-facing services this document never listed).
|
||||
> Kept for history only.
|
||||
|
||||
# app4 Scoping and Migration Plan
|
||||
|
||||
**Owner:** IT Pro Partner (Germaine Brown)
|
||||
|
||||
@@ -0,0 +1,288 @@
|
||||
{
|
||||
"meta": {
|
||||
"host": "Core (152.53.192.33)",
|
||||
"spec": "netcup RS 2000, 8 vCPU / 15GB RAM / 503GB disk, Debian 13, Manassas VA",
|
||||
"scan_date": "2026-09-15",
|
||||
"method": "read-only commands executed directly on Core; no service started/stopped/restarted; hermes-gateway left untouched",
|
||||
"counts": {
|
||||
"docker_containers_total": 13,
|
||||
"docker_containers_running": 13,
|
||||
"systemd_services_running": 57,
|
||||
"caddy_site_blocks": 49,
|
||||
"postgres_databases_excl_templates": 1,
|
||||
"hermes_cron_jobs": 87,
|
||||
"root_system_crontab_entries": 12,
|
||||
"listening_tcp_sockets": 66,
|
||||
"listening_udp_sockets": 14,
|
||||
"tls_certs_acme": 47,
|
||||
"tls_certs_self_signed_internal": 7
|
||||
}
|
||||
},
|
||||
"services": [
|
||||
{"name": "auth-api", "type": "systemd", "purpose": "ITPP centralized Auth API / SSO", "ports": ["127.0.0.1:8500"], "data_paths": [{"path": "/root/projects/auth/auth.db", "size": "32M (dir)"}], "config_paths": ["/root/projects/auth/.env"], "hostnames": ["auth.itpropartner.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "yes - auth-api-backup.sh (3:15/4:35 AM)", "migration_classification": "needs-decision"},
|
||||
{"name": "diglocate-api", "type": "systemd", "purpose": "811 locate ticket management API", "ports": ["127.0.0.1:8000"], "data_paths": [{"path": "/root/projects/diglocate", "size": "118M"}], "config_paths": ["none found"], "hostnames": ["dig.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "dre-mcp", "type": "systemd", "purpose": "DRE MCP server (Hermes tool for DRE data)", "ports": [], "data_paths": [{"path": "/root/docker/dre-mcp", "size": "171M"}], "config_paths": ["/root/.hermes/.env"], "hostnames": [], "hermes_related": true, "customer_facing": true, "backup_coverage": "no", "migration_classification": "needs-decision"},
|
||||
{"name": "dre-portal", "type": "systemd", "purpose": "DRE customer portal API (FastAPI/uvicorn)", "ports": ["127.0.0.1:8093"], "data_paths": [{"path": "/opt/dre-portal/data/dre.db", "size": "110M (dir)"}], "config_paths": ["/opt/dre-portal/.env"], "hostnames": ["internal.debtrecoveryexperts.com", "my.debtrecoveryexperts.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "ft360-mcp", "type": "systemd", "purpose": "FleetTracker360 MCP server", "ports": [], "data_paths": [{"path": "/root/docker/ft360-mcp", "size": "60K"}], "config_paths": ["none found"], "hostnames": [], "hermes_related": true, "customer_facing": true, "backup_coverage": "partial - stats/export scripts only, no DB backup found", "migration_classification": "needs-decision"},
|
||||
{"name": "gitea-runner", "type": "systemd", "purpose": "Gitea Actions Runner (core)", "ports": ["0.0.0.0:35849"], "data_paths": [{"path": "/var/lib/gitea-runner", "size": "not measured"}], "config_paths": ["none found"], "hostnames": [], "hermes_related": false, "customer_facing": false, "backup_coverage": "no", "migration_classification": "stays-on-Core"},
|
||||
{"name": "hermes-assistant", "type": "systemd", "purpose": "Hermes Assistant PWA backend", "ports": ["127.0.0.1:8082"], "data_paths": [{"path": "/root/hermes-assistant", "size": "169M"}], "config_paths": ["none found"], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "unknown", "migration_classification": "stays-on-Core"},
|
||||
{"name": "hermes-browser", "type": "systemd", "purpose": "Headless Chromium (CDP browser) for Hermes", "ports": ["127.0.0.1:9222"], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "n/a - stateless", "migration_classification": "stays-on-Core"},
|
||||
{"name": "hermes-socat-8787", "type": "systemd", "purpose": "Hermes API port forward (8787 -> 8642) for HermesX mobile", "ports": ["0.0.0.0:8787"], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "n/a", "migration_classification": "stays-on-Core"},
|
||||
{"name": "hermes-voice", "type": "systemd", "purpose": "Hermes Voice (SvelteKit frontend)", "ports": ["127.0.0.1:4331"], "data_paths": [{"path": "/opt/hermes-voice", "size": "123M"}], "config_paths": ["/opt/hermes-voice/.env"], "hostnames": ["voice.itpropartner.com"], "hermes_related": true, "customer_facing": true, "backup_coverage": "unknown", "migration_classification": "needs-decision"},
|
||||
{"name": "host-metrics-exporter", "type": "systemd", "purpose": "systemd/Docker/disk/memory metrics exporter", "ports": [], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": false, "customer_facing": false, "backup_coverage": "n/a", "migration_classification": "stays-on-Core"},
|
||||
{"name": "hotnow-api", "type": "systemd", "purpose": "HotNow API backend", "ports": ["127.0.0.1:8001"], "data_paths": [{"path": "/root/hotnow-api", "size": "66M"}, {"path": "postgresql db hotnow", "size": "16MB"}], "config_paths": ["/root/hotnow-api/.env"], "hostnames": ["api.hotnow.io"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "intelsight-api", "type": "systemd", "purpose": "IntelSight API service", "ports": ["127.0.0.1:8099"], "data_paths": [{"path": "/root/intelsight-api/intelsight.db", "size": "64M (dir)"}], "config_paths": ["none found"], "hostnames": ["my.intelsight.io"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "node_exporter", "type": "systemd", "purpose": "Prometheus Node Exporter", "ports": ["0.0.0.0:9100"], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": false, "customer_facing": false, "backup_coverage": "n/a", "migration_classification": "stays-on-Core"},
|
||||
{"name": "ops-portal", "type": "systemd", "purpose": "ITPP Ops Portal backend", "ports": ["127.0.0.1:8090"], "data_paths": [{"path": "/opt/ops-portal/ops.db", "size": "148M (dir)"}], "config_paths": ["/root/.hermes/.env"], "hostnames": ["ops.itpropartner.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "unknown - not explicitly listed in backup-plan.md as separate from core-services-backup.sh", "migration_classification": "move-to-app4"},
|
||||
{"name": "osint-api", "type": "systemd", "purpose": "OSINT Tool API", "ports": ["127.0.0.1:8100"], "data_paths": [{"path": "/opt/osint-api", "size": "323M"}], "config_paths": ["/root/.hermes/.env"], "hostnames": ["ops.itpropartner.com (/api/search/*)"], "hermes_related": true, "customer_facing": true, "backup_coverage": "no", "migration_classification": "needs-decision"},
|
||||
{"name": "osint-person", "type": "systemd", "purpose": "OSINT Person MCP server", "ports": [], "data_paths": [{"path": "/root/docker/osint-person-mcp", "size": "314M"}], "config_paths": ["/root/.hermes/.env"], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "no", "migration_classification": "stays-on-Core"},
|
||||
{"name": "outlook-upload", "type": "systemd", "purpose": "Outlook folder upload receiver (chunked/resumable)", "ports": ["127.0.0.1:8240"], "data_paths": [{"path": "/root/upload-staging", "size": "not measured"}], "config_paths": ["none"], "hostnames": ["core.itpropartner.com (/o-.../*)"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "needs-decision"},
|
||||
{"name": "pipeline-api", "type": "systemd", "purpose": "Project Pipeline API - ITPP customer portal backend", "ports": ["127.0.0.1:8200"], "data_paths": [{"path": "/root/projects/pipeline/pipeline.db", "size": "30M (dir)"}], "config_paths": ["none found"], "hostnames": ["my.itpropartner.com (/api/pipeline/*)"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "pry", "type": "systemd", "purpose": "PRY unified OSINT search backend", "ports": ["127.0.0.1:8905"], "data_paths": [{"path": "/root/docker/pry", "size": "66M"}], "config_paths": ["/root/.hermes/.env"], "hostnames": ["pry.iamgmb.com"], "hermes_related": true, "customer_facing": true, "backup_coverage": "no", "migration_classification": "needs-decision"},
|
||||
{"name": "pta-registration", "type": "systemd", "purpose": "TIMAPTA Membership Registration", "ports": ["127.0.0.1:8114"], "data_paths": [{"path": "/opt/pta-registration/pta.db", "size": "53M (dir)"}], "config_paths": ["none"], "hostnames": ["register.timapta.org", "pta.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "rally", "type": "systemd", "purpose": "Rally Family Calendar", "ports": ["127.0.0.1:8105"], "data_paths": [{"path": "/opt/rally/data/rally.db", "size": "within 294M dir"}], "config_paths": ["none found"], "hostnames": ["rally.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "partial - debug/sync scripts exist, no scheduled DB backup confirmed", "migration_classification": "move-to-app4"},
|
||||
{"name": "seemytrip", "type": "systemd", "purpose": "SeeMyTrip collaborative trip media pipeline", "ports": ["127.0.0.1:8113"], "data_paths": [{"path": "/opt/seemytrip/data/seemytrip.db", "size": "within 212M dir"}], "config_paths": ["none"], "hostnames": ["seemytrip.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "shark-game", "type": "systemd", "purpose": "Shark Attack Fantasy Game backend", "ports": ["0.0.0.0:8083"], "data_paths": [{"path": "/root/shark-game/backend/game.db", "size": "within 127M dir"}], "config_paths": ["none"], "hostnames": ["shark.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "shopping-cart", "type": "systemd", "purpose": "Shopping Cart Builder", "ports": ["127.0.0.1:8101"], "data_paths": [{"path": "/opt/shopping-cart", "size": "147M"}], "config_paths": ["none"], "hostnames": ["shopping.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "snmp-metrics", "type": "systemd", "purpose": "SNMP Metrics HTTP Server", "ports": ["0.0.0.0:8105"], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": false, "customer_facing": false, "backup_coverage": "n/a", "migration_classification": "stays-on-Core"},
|
||||
{"name": "super-search", "type": "systemd", "purpose": "Super Search MCP Server (web_search/web_extract for Hermes)", "ports": ["0.0.0.0:8899"], "data_paths": [{"path": "/root/docker/super-search", "size": "1.1G"}], "config_paths": ["/root/.hermes/.env"], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "unknown", "migration_classification": "stays-on-Core"},
|
||||
{"name": "survey-registration", "type": "systemd", "purpose": "TIMA Location Survey", "ports": ["127.0.0.1:8115"], "data_paths": [{"path": "/opt/pta-survey/survey.db", "size": "50M (dir)"}], "config_paths": ["none"], "hostnames": ["survey.iamgmb.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "transitpin", "type": "systemd", "purpose": "TransitPin WebSocket Relay", "ports": ["127.0.0.1:8210"], "data_paths": [{"path": "/opt/transitpin", "size": "496K"}], "config_paths": ["none"], "hostnames": ["status.itpropartner.com (/api/msg*)", "shopping.iamgmb.com (/api/msg*)"], "hermes_related": false, "customer_facing": true, "backup_coverage": "yes - transitpin-backup.sh (app3-associated, verify covers Core path)", "migration_classification": "move-to-app4"},
|
||||
{"name": "twilio-mcp", "type": "systemd", "purpose": "Twilio MCP Server", "ports": [], "data_paths": [{"path": "/root/docker/twilio-mcp", "size": "171M"}], "config_paths": ["/root/.hermes/.env"], "hostnames": [], "hermes_related": true, "customer_facing": true, "backup_coverage": "unknown", "migration_classification": "stays-on-Core"},
|
||||
{"name": "verdicttank-api", "type": "systemd", "purpose": "VerdictTank API - form handler and PDF generation", "ports": ["127.0.0.1:8201"], "data_paths": [{"path": "/opt/verdicttank/users.db", "size": "within 15M dir"}], "config_paths": ["/etc/verdicttank.env"], "hostnames": ["verdicttank.com", "www.verdicttank.com", "ops.verdicttank.com", "core.itpropartner.com (/api/verdicttank/*)"], "hermes_related": false, "customer_facing": true, "backup_coverage": "partial - collect scripts exist, no DB backup confirmed", "migration_classification": "move-to-app4"},
|
||||
{"name": "verdicttank-worker", "type": "systemd", "purpose": "VerdictTank review worker - polls queue, runs multi-model panel", "ports": [], "data_paths": [{"path": "/opt/verdicttank", "size": "15M"}], "config_paths": ["/etc/verdicttank.env"], "hostnames": [], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "voice-agent-stt", "type": "systemd", "purpose": "Voice Agent STT (faster-whisper)", "ports": ["127.0.0.1:9000"], "data_paths": [{"path": "/opt/voice-agent", "size": "470M"}], "config_paths": ["none"], "hostnames": [], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "needs-decision"},
|
||||
{"name": "voice-agent", "type": "systemd", "purpose": "Voice Agent (open-source stack)", "ports": ["127.0.0.1:9101"], "data_paths": [{"path": "/opt/voice-agent", "size": "470M"}], "config_paths": ["/opt/voice-agent/.env"], "hostnames": ["voice-open.itpropartner.com"], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "move-to-app4"},
|
||||
{"name": "wazuh-agent", "type": "systemd", "purpose": "Wazuh SIEM/XDR agent", "ports": [], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": false, "customer_facing": false, "backup_coverage": "n/a", "migration_classification": "stays-on-Core"},
|
||||
{"name": "caddy", "type": "systemd", "purpose": "Reverse proxy for all Core-hosted domains", "ports": ["152.53.192.33:80", "152.53.192.33:443"], "data_paths": [{"path": "/var/lib/caddy", "size": "not separately measured"}], "config_paths": ["/etc/caddy/Caddyfile"], "hostnames": ["all 49 site blocks - see caddy_sites"], "hermes_related": false, "customer_facing": true, "backup_coverage": "unknown - Caddyfile not confirmed in root-essentials-backup.sh path list", "migration_classification": "stays-on-Core (trim customer routes post-cutover per plan)"},
|
||||
{"name": "postgresql@17-main", "type": "systemd", "purpose": "Host PostgreSQL cluster", "ports": ["127.0.0.1:5432", "[::1]:5432"], "data_paths": [{"path": "/var/lib/postgresql", "size": "71M"}], "config_paths": [], "hostnames": [], "hermes_related": false, "customer_facing": true, "backup_coverage": "no dedicated pg_dump script found for hotnow db", "migration_classification": "needs-decision (only db is hotnow, which is moving)"},
|
||||
{"name": "redis-server", "type": "systemd", "purpose": "Redis cache/keystore", "ports": ["127.0.0.1:6379", "[::1]:6379"], "data_paths": [{"path": "/var/lib/redis", "size": "8.0K"}], "config_paths": [], "hostnames": [], "hermes_related": false, "customer_facing": true, "backup_coverage": "no", "migration_classification": "needs-decision (HotNow is confirmed consumer via DB1)"},
|
||||
{"name": "crawl4ai", "type": "systemd", "purpose": "Crawl4AI extraction microservice", "ports": ["8910 (configured, service disabled/inactive)"], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "n/a - inactive", "migration_classification": "needs-decision (disabled, unclear if still needed)"},
|
||||
{"name": "hermes-control-deck", "type": "systemd", "purpose": "Hermes Control Deck Backend API Server", "ports": ["127.0.0.1:8200 (configured, service disabled/inactive)"], "data_paths": [], "config_paths": [], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "n/a - inactive", "migration_classification": "needs-decision (port 8200 conflicts with active pipeline-api)"},
|
||||
{"name": "hermes-gateway", "type": "systemd (user unit)", "purpose": "Hermes Agent Gateway - messaging platform integration", "ports": ["0.0.0.0:8642"], "data_paths": [], "config_paths": ["/root/.hermes/config.yaml"], "hostnames": [], "hermes_related": true, "customer_facing": false, "backup_coverage": "yes - hermes-backup.sh + hermes-live-sync", "migration_classification": "stays-on-Core (untouched during this audit)"}
|
||||
],
|
||||
"containers": [
|
||||
{"name": "docuseal", "image": "docuseal/docuseal:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:8091->3000/tcp"], "restart_policy": "always", "mounts": [{"type": "bind", "source": "/root/docker/docuseal/data", "dest": "/data"}], "compose_file": "/root/docker/docuseal/docker-compose.yml", "hostnames": ["sign.itpropartner.com"], "migration_classification": "move-to-app4"},
|
||||
{"name": "docuseal-dre", "image": "docuseal/docuseal:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:8094->3000/tcp"], "restart_policy": "always", "mounts": [{"type": "bind", "source": "/root/docker/docuseal-dre/data", "dest": "/data"}], "compose_file": "/root/docker/docuseal-dre/docker-compose.yml", "hostnames": ["sign.debtrecoveryexperts.com"], "migration_classification": "move-to-app4"},
|
||||
{"name": "docuseal-modelortho", "image": "docuseal/docuseal:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:8092->3000/tcp"], "restart_policy": "always", "mounts": [{"type": "bind", "source": "/root/docker/docuseal-modelortho/data", "dest": "/data"}], "compose_file": "/root/docker/docuseal-modelortho/docker-compose.yml", "hostnames": ["sign.modelortho.com"], "migration_classification": "move-to-app4"},
|
||||
{"name": "timetrex", "image": "skewll/timetrex:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:8085->80/tcp"], "restart_policy": "unless-stopped", "mounts": [{"type": "bind", "source": "/root/docker/timetrex/storage", "dest": "/storage"}, {"type": "bind", "source": "/root/docker/timetrex/logs", "dest": "/logs"}, {"type": "bind", "source": "/root/docker/timetrex/database", "dest": "/database"}, {"type": "bind", "source": "/root/docker/timetrex/timetrex.ini.php", "dest": "/var/www/html/timetrex/timetrex.ini.php"}], "compose_file": "/root/docker/timetrex/docker-compose.yml", "hostnames": ["timetrex.iamgmb.com"], "migration_classification": "move-to-app4", "note": "runs its own internal PostgreSQL 16 instance in /database, separate from host PG17"},
|
||||
{"name": "microbin", "image": "danielszabo99/microbin:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:8260->8080/tcp"], "restart_policy": "unless-stopped", "mounts": [{"type": "bind", "source": "/opt/microbin/data", "dest": "/app/pasta_data"}], "compose_file": "/opt/microbin/docker-compose.yml", "hostnames": ["share.itpropartner.com"], "migration_classification": "move-to-app4"},
|
||||
{"name": "uptime-kuma", "image": "louislam/uptime-kuma:latest", "status": "Up 2 hours (healthy)", "ports": ["0.0.0.0:3001->3001/tcp"], "restart_policy": "unless-stopped", "mounts": [{"type": "bind", "source": "/root/docker/uptime-kuma/data", "dest": "/app/data"}], "compose_file": "/root/docker/uptime-kuma/docker-compose.yml", "hostnames": ["uptimekuma.itpropartner.com"], "migration_classification": "move-to-app4"},
|
||||
{"name": "grafana", "image": "grafana/grafana:11.4.0", "status": "Up 2 hours", "ports": ["3002 (container-internal, exposed on all host interfaces)"], "restart_policy": "unless-stopped", "mounts": [{"type": "volume", "source": "grafana_data_final", "dest": "/var/lib/grafana"}], "compose_file": "not found", "hostnames": [], "migration_classification": "stays-on-Core"},
|
||||
{"name": "prometheus", "image": "prom/prometheus:latest", "status": "Up 2 hours", "ports": ["9090 (container)"], "restart_policy": "unless-stopped", "mounts": [{"type": "bind", "source": "/root/docker/monitoring/prometheus/prometheus.yml", "dest": "/etc/prometheus/prometheus.yml"}, {"type": "volume", "source": "prometheus_data", "dest": "/prometheus"}, {"type": "bind", "source": "/var/lib/prometheus/textfile", "dest": "/var/lib/prometheus/textfile"}], "compose_file": "not directly located", "hostnames": [], "migration_classification": "stays-on-Core"},
|
||||
{"name": "telegraf", "image": "telegraf:latest", "status": "Up 2 hours", "ports": [], "restart_policy": "unless-stopped", "mounts": [{"type": "bind", "source": "/root/docker/monitoring/telegraf/telegraf.conf", "dest": "/etc/telegraf/telegraf.conf"}], "compose_file": "not directly located", "hostnames": [], "migration_classification": "stays-on-Core"},
|
||||
{"name": "mikrotik-exporter", "image": "swoga/mikrotik-exporter:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:9436->9436/tcp"], "restart_policy": "unless-stopped", "mounts": [{"type": "bind", "source": "/root/docker/monitoring/mikrotik-exporter/config.yml", "dest": "/etc/mikrotik-exporter/config.yml"}], "compose_file": "not directly located", "hostnames": [], "migration_classification": "stays-on-Core"},
|
||||
{"name": "searxng", "image": "searxng/searxng:latest", "status": "Up 2 hours", "ports": ["127.0.0.1:8888->8080/tcp"], "restart_policy": "always", "mounts": [{"type": "bind", "source": "/root/docker/searxng/searxng-data", "dest": "/etc/searxng"}, {"type": "bind", "source": "/root/docker/searxng/searxng-themes", "dest": "/usr/local/searxng/searx/static/themes"}], "compose_file": "/root/docker/searxng/docker-compose.yml", "hostnames": [], "migration_classification": "stays-on-Core", "note": "backup-plan.md incorrectly lists this as removed/stale; it is running"},
|
||||
{"name": "browserless", "image": "browserless/chrome:latest", "status": "Up 2 hours", "ports": ["0.0.0.0:3000->3000/tcp"], "restart_policy": "always", "mounts": [], "compose_file": "not found", "hostnames": [], "migration_classification": "stays-on-Core"},
|
||||
{"name": "camofox-browser", "image": "camofox-browser:dataimpulse", "status": "Up 2 hours", "ports": ["0.0.0.0:9377->9377/tcp"], "restart_policy": "unless-stopped", "mounts": [], "compose_file": "not found", "hostnames": [], "migration_classification": "stays-on-Core"}
|
||||
],
|
||||
"volumes": [
|
||||
{"name": "grafana_data_final", "size": "14.65MB", "used_by": "grafana", "status": "active"},
|
||||
{"name": "grafana_data", "size": "50.13MB", "used_by": "none (orphaned, older grafana volume)", "status": "orphaned"},
|
||||
{"name": "grafana_data_v3", "size": "14.59MB", "used_by": "none (orphaned)", "status": "orphaned"},
|
||||
{"name": "prometheus_data", "size": "118.9MB", "used_by": "prometheus", "status": "active"},
|
||||
{"name": "twenty_db-data", "size": "71.46MB", "used_by": "none (Twenty CRM migrated to App1, orphaned on Core)", "status": "orphaned"},
|
||||
{"name": "twenty_server-local-data", "size": "0B", "used_by": "none", "status": "orphaned"},
|
||||
{"name": "act-toolcache", "size": "0B", "used_by": "gitea-runner (possibly)", "status": "unclear"}
|
||||
],
|
||||
"ports": [
|
||||
{"port": 443, "bind": "152.53.192.33", "process": "caddy", "protocol": "tcp"},
|
||||
{"port": 80, "bind": "152.53.192.33", "process": "caddy", "protocol": "tcp"},
|
||||
{"port": 443, "bind": "100.71.155.7 (tailscale)", "process": "tailscaled", "protocol": "tcp"},
|
||||
{"port": 22, "bind": "0.0.0.0 / [::]", "process": "sshd", "protocol": "tcp"},
|
||||
{"port": 5432, "bind": "127.0.0.1 / [::1]", "process": "postgres", "protocol": "tcp"},
|
||||
{"port": 6379, "bind": "127.0.0.1 / [::1]", "process": "redis-server", "protocol": "tcp"},
|
||||
{"port": 3000, "bind": "0.0.0.0 / [::]", "process": "docker-proxy (browserless)", "protocol": "tcp"},
|
||||
{"port": 3001, "bind": "0.0.0.0 / [::]", "process": "docker-proxy (uptime-kuma)", "protocol": "tcp"},
|
||||
{"port": 3002, "bind": "* (all interfaces)", "process": "grafana", "protocol": "tcp"},
|
||||
{"port": 9377, "bind": "0.0.0.0 / [::]", "process": "docker-proxy (camofox)", "protocol": "tcp"},
|
||||
{"port": 8000, "bind": "127.0.0.1", "process": "uvicorn (diglocate-api)", "protocol": "tcp"},
|
||||
{"port": 8001, "bind": "127.0.0.1", "process": "uvicorn (hotnow-api)", "protocol": "tcp"},
|
||||
{"port": 8082, "bind": "127.0.0.1", "process": "python3 (hermes-assistant)", "protocol": "tcp"},
|
||||
{"port": 8083, "bind": "0.0.0.0", "process": "python3 (shark-game)", "protocol": "tcp"},
|
||||
{"port": 8085, "bind": "127.0.0.1", "process": "docker-proxy (timetrex)", "protocol": "tcp"},
|
||||
{"port": 8090, "bind": "127.0.0.1", "process": "uvicorn (ops-portal)", "protocol": "tcp"},
|
||||
{"port": 8091, "bind": "127.0.0.1", "process": "docker-proxy (docuseal)", "protocol": "tcp"},
|
||||
{"port": 8092, "bind": "127.0.0.1", "process": "docker-proxy (docuseal-modelortho)", "protocol": "tcp"},
|
||||
{"port": 8093, "bind": "127.0.0.1", "process": "uvicorn (dre-portal)", "protocol": "tcp"},
|
||||
{"port": 8094, "bind": "127.0.0.1", "process": "docker-proxy (docuseal-dre)", "protocol": "tcp"},
|
||||
{"port": 8099, "bind": "0.0.0.0", "process": "python3 (intelsight-api)", "protocol": "tcp"},
|
||||
{"port": 8100, "bind": "127.0.0.1", "process": "uvicorn (osint-api)", "protocol": "tcp"},
|
||||
{"port": 8101, "bind": "127.0.0.1", "process": "uvicorn (shopping-cart)", "protocol": "tcp"},
|
||||
{"port": 8105, "bind": "0.0.0.0", "process": "python (rally) / snmp-metrics collision note", "protocol": "tcp"},
|
||||
{"port": 8113, "bind": "127.0.0.1", "process": "uvicorn (seemytrip)", "protocol": "tcp"},
|
||||
{"port": 8114, "bind": "127.0.0.1", "process": "uvicorn (pta-registration)", "protocol": "tcp"},
|
||||
{"port": 8115, "bind": "127.0.0.1", "process": "uvicorn (survey-registration)", "protocol": "tcp"},
|
||||
{"port": 8200, "bind": "127.0.0.1", "process": "python (pipeline-api)", "protocol": "tcp"},
|
||||
{"port": 8201, "bind": "127.0.0.1", "process": "python3 (verdicttank-api)", "protocol": "tcp"},
|
||||
{"port": 8210, "bind": "127.0.0.1", "process": "node (transitpin)", "protocol": "tcp"},
|
||||
{"port": 8240, "bind": "127.0.0.1", "process": "python3 (outlook-upload)", "protocol": "tcp"},
|
||||
{"port": 8260, "bind": "127.0.0.1", "process": "docker-proxy (microbin)", "protocol": "tcp"},
|
||||
{"port": 8500, "bind": "127.0.0.1", "process": "uvicorn (auth-api)", "protocol": "tcp"},
|
||||
{"port": 8787, "bind": "0.0.0.0", "process": "socat (hermes forward)", "protocol": "tcp"},
|
||||
{"port": 8642, "bind": "0.0.0.0", "process": "hermes (gateway internal)", "protocol": "tcp"},
|
||||
{"port": 8888, "bind": "127.0.0.1", "process": "docker-proxy (searxng)", "protocol": "tcp"},
|
||||
{"port": 8899, "bind": "0.0.0.0", "process": "python3 (super-search MCP)", "protocol": "tcp"},
|
||||
{"port": 8900, "bind": "127.0.0.1", "process": "python3", "protocol": "tcp"},
|
||||
{"port": 8901, "bind": "127.0.0.1", "process": "python3", "protocol": "tcp"},
|
||||
{"port": 8902, "bind": "127.0.0.1", "process": "python3", "protocol": "tcp"},
|
||||
{"port": 8903, "bind": "127.0.0.1", "process": "python3", "protocol": "tcp"},
|
||||
{"port": 8905, "bind": "127.0.0.1", "process": "python (pry)", "protocol": "tcp"},
|
||||
{"port": 9000, "bind": "127.0.0.1", "process": "uvicorn (voice-agent-stt)", "protocol": "tcp"},
|
||||
{"port": 9090, "bind": "*", "process": "prometheus", "protocol": "tcp"},
|
||||
{"port": 9100, "bind": "*", "process": "node_exporter", "protocol": "tcp"},
|
||||
{"port": 9101, "bind": "127.0.0.1", "process": "uvicorn (voice-agent)", "protocol": "tcp"},
|
||||
{"port": 9222, "bind": "127.0.0.1", "process": "chrome (hermes-browser CDP)", "protocol": "tcp"},
|
||||
{"port": 9273, "bind": "*", "process": "telegraf", "protocol": "tcp"},
|
||||
{"port": 9274, "bind": "127.0.0.1", "process": "python3", "protocol": "tcp"},
|
||||
{"port": 9275, "bind": "127.0.0.1", "process": "python3", "protocol": "tcp"},
|
||||
{"port": 9436, "bind": "127.0.0.1", "process": "docker-proxy (mikrotik-exporter)", "protocol": "tcp"},
|
||||
{"port": 25, "bind": "127.0.0.1 / [::1]", "process": "exim4", "protocol": "tcp"},
|
||||
{"port": 500, "bind": "0.0.0.0 / [::]", "process": "charon (strongswan)", "protocol": "udp"},
|
||||
{"port": 4500, "bind": "0.0.0.0 / [::]", "process": "charon (strongswan)", "protocol": "udp"},
|
||||
{"port": 1701, "bind": "0.0.0.0", "process": "xl2tpd", "protocol": "udp"},
|
||||
{"port": 51821, "bind": "0.0.0.0 / [::]", "process": "unowned (WireGuard per UFW comment)", "protocol": "udp"},
|
||||
{"port": 41641, "bind": "0.0.0.0 / [::]", "process": "tailscaled", "protocol": "udp"},
|
||||
{"port": 5353, "bind": "0.0.0.0 / [::]", "process": "avahi-daemon", "protocol": "udp"}
|
||||
],
|
||||
"caddy_sites": [
|
||||
{"hostname": "core.itpropartner.com", "backend": ["127.0.0.1:8240", "127.0.0.1:8201", "static /var/www"], "tls": "acme, expires 2026-12-05"},
|
||||
{"hostname": "sign.itpropartner.com", "backend": ["127.0.0.1:8091"], "tls": "acme, expires 2026-12-06"},
|
||||
{"hostname": "sign.modelortho.com", "backend": ["127.0.0.1:8092"], "tls": "acme, expires 2026-11-20"},
|
||||
{"hostname": "sign.debtrecoveryexperts.com", "backend": ["127.0.0.1:8094"], "tls": "acme, expires 2026-11-21"},
|
||||
{"hostname": "ops.itpropartner.com", "backend": ["127.0.0.1:8090", "127.0.0.1:8100", "static"], "tls": "acme, expires 2026-12-06"},
|
||||
{"hostname": "shark.iamgmb.com", "backend": ["127.0.0.1:8083"], "tls": "acme, expires 2026-12-06"},
|
||||
{"hostname": "internal.debtrecoveryexperts.com", "backend": ["127.0.0.1:8093", "static", "basic_auth protected"], "tls": "acme, expires 2026-12-06"},
|
||||
{"hostname": "portal.debtrecoveryexperts.com", "backend": ["redirect to my.debtrecoveryexperts.com/start"], "tls": "acme, expires 2026-12-06"},
|
||||
{"hostname": "pay.debtrecoveryexperts.com", "backend": ["static"], "tls": "acme, expires 2026-10-24"},
|
||||
{"hostname": "my.debtrecoveryexperts.com", "backend": ["127.0.0.1:8093", "static"], "tls": "acme, expires 2026-11-20"},
|
||||
{"hostname": "crm.debtrecoveryexperts.com", "backend": ["localhost:3003 - NOTHING LISTENING"], "tls": "acme, expires 2026-12-06", "note": "dead backend"},
|
||||
{"hostname": "dig.iamgmb.com", "backend": ["127.0.0.1:8000", "static"], "tls": "acme, expires 2026-10-26"},
|
||||
{"hostname": "uptimekuma.itpropartner.com", "backend": ["localhost:3001"], "tls": "acme, expires 2026-12-11"},
|
||||
{"hostname": "gps.fleettracker360.com", "backend": ["152.53.39.202:8082 (App2, remote)"], "tls": "acme, expires 2026-12-13"},
|
||||
{"hostname": "my.itpropartner.com", "backend": ["152.53.241.111:8090 (App3, remote)", "127.0.0.1:8200", "static"], "tls": "acme, expires 2026-10-15"},
|
||||
{"hostname": "status.itpropartner.com", "backend": ["127.0.0.1:8210", "127.0.0.1:3001", "static"], "tls": "acme, expires 2026-10-16"},
|
||||
{"hostname": "track.fleettracker360.com", "backend": ["152.53.39.202:5055 (App2, remote)"], "tls": "http only, no TLS block"},
|
||||
{"hostname": "hear.fleettracker360.com", "backend": ["static"], "tls": "acme, expires 2026-10-17"},
|
||||
{"hostname": "voice.itpropartner.com", "backend": ["127.0.0.1:4331"], "tls": "acme, expires 2026-10-24"},
|
||||
{"hostname": "voice-open.itpropartner.com", "backend": ["127.0.0.1:9101"], "tls": "acme, expires 2026-10-24"},
|
||||
{"hostname": "auth.itpropartner.com", "backend": ["127.0.0.1:8500", "static"], "tls": "acme, expires 2026-10-30"},
|
||||
{"hostname": "my.intelsight.io", "backend": ["127.0.0.1:8099", "static"], "tls": "acme, expires 2026-10-24"},
|
||||
{"hostname": "intelsight.io", "backend": ["static"], "tls": "acme, expires 2026-10-28"},
|
||||
{"hostname": "intelsight.iamgmb.com", "backend": ["static"], "tls": "acme, expires 2026-10-26"},
|
||||
{"hostname": "schedule.iamgmb.com", "backend": ["static"], "tls": "acme (not individually captured)"},
|
||||
{"hostname": "seemytrip.iamgmb.com", "backend": ["127.0.0.1:8113", "static"], "tls": "acme, expires 2026-11-03"},
|
||||
{"hostname": "rally.iamgmb.com", "backend": ["127.0.0.1:8105", "static"], "tls": "acme, expires 2026-10-28"},
|
||||
{"hostname": "shopping.iamgmb.com", "backend": ["127.0.0.1:8101", "127.0.0.1:8210", "static"], "tls": "acme, expires 2026-10-25"},
|
||||
{"hostname": "pry.iamgmb.com", "backend": ["127.0.0.1:8905", "static"], "tls": "acme, expires 2026-10-26 (http listed as http:// in Caddyfile but cert present)"},
|
||||
{"hostname": "crm.intelsight.io", "backend": ["localhost:3003 - NOTHING LISTENING"], "tls": "acme, expires 2026-10-28", "note": "dead backend"},
|
||||
{"hostname": "www.hotnow.io", "backend": ["redirect to hotnow.io"], "tls": "acme, expires 2026-10-31"},
|
||||
{"hostname": "hotnow.io", "backend": ["static"], "tls": "acme, expires 2026-10-31"},
|
||||
{"hostname": "app.hotnow.io", "backend": ["static"], "tls": "acme, expires 2026-10-31"},
|
||||
{"hostname": "api.hotnow.io", "backend": ["127.0.0.1:8001"], "tls": "acme (ZeroSSL CA), expires 2026-10-31"},
|
||||
{"hostname": "admin.hotnow.io", "backend": ["static"], "tls": "acme, expires 2026-10-31"},
|
||||
{"hostname": "timetrex.iamgmb.com", "backend": ["127.0.0.1:8085"], "tls": "acme, expires 2026-11-02"},
|
||||
{"hostname": "webmail.timapta.org", "backend": ["redirect to heracles.mxrouting.net"], "tls": "tls internal (self-signed), expires 2026-09-15"},
|
||||
{"hostname": "webmail.transitpin.com", "backend": ["redirect to heracles.mxrouting.net"], "tls": "tls internal (self-signed), expires 2026-09-15"},
|
||||
{"hostname": "webmail.rfptank.com", "backend": ["redirect to heracles.mxrouting.net"], "tls": "tls internal (self-signed), expires 2026-09-16"},
|
||||
{"hostname": "webmail.radartank.com", "backend": ["redirect to heracles.mxrouting.net"], "tls": "tls internal (self-signed), expires 2026-09-16"},
|
||||
{"hostname": "webmail.verdicttank.com", "backend": ["redirect to heracles.mxrouting.net"], "tls": "tls internal (self-signed), expires 2026-09-16"},
|
||||
{"hostname": "share.itpropartner.com", "backend": ["127.0.0.1:8260"], "tls": "acme, expires 2026-11-02"},
|
||||
{"hostname": "verdicttank.com, www.verdicttank.com", "backend": ["127.0.0.1:8201", "static"], "tls": "acme, expires 2026-11-06"},
|
||||
{"hostname": "voipsimplicity.itpropartner.com", "backend": ["static"], "tls": "tls internal (self-signed), expires 2026-09-15"},
|
||||
{"hostname": "forefront.itpropartner.com", "backend": ["static"], "tls": "tls internal (self-signed), expires 2026-09-15"},
|
||||
{"hostname": "ops.verdicttank.com", "backend": ["127.0.0.1:8201", "static"], "tls": "acme, expires 2026-11-17"},
|
||||
{"hostname": "register.timapta.org", "backend": ["127.0.0.1:8114"], "tls": "acme, expires 2026-12-08"},
|
||||
{"hostname": "pta.iamgmb.com", "backend": ["127.0.0.1:8114"], "tls": "acme, expires 2026-12-09"},
|
||||
{"hostname": "survey.iamgmb.com", "backend": ["127.0.0.1:8115"], "tls": "acme, expires 2026-12-09"}
|
||||
],
|
||||
"databases": [
|
||||
{"engine": "PostgreSQL 17.10", "instance": "host (postgresql@17-main.service)", "port": 5432, "databases": [
|
||||
{"name": "hotnow", "owner": "hotnow_app", "size": "16 MB"},
|
||||
{"name": "postgres", "owner": "postgres", "size": "7510 kB"},
|
||||
{"name": "template0", "owner": "postgres", "size": "7353 kB"},
|
||||
{"name": "template1", "owner": "postgres", "size": "7582 kB"}
|
||||
]},
|
||||
{"engine": "PostgreSQL 16", "instance": "containerized inside timetrex container", "port": "not published to host", "databases": [{"name": "timetrex", "owner": "timetrex", "size": "not separately queryable from host"}]},
|
||||
{"engine": "Redis 8.0.2", "instance": "host (redis-server.service)", "port": 6379, "used_memory": "758688 bytes (741K)", "appendonly": false, "save_policy": "3600 1 300 100 60 10000", "keyspace_db0": 0, "keyspace_db1": 0, "known_consumer": "hotnow-api (redis://localhost:6379/1 in .env)"}
|
||||
],
|
||||
"cron": [
|
||||
{"source": "root system crontab", "entries": 12, "detail": "boys-mail-monitor (hourly+daily), shark-game scraper (daily), shark-draft-reminder (15min), hermes-backup.sh (1AM), backup-audit-check.sh (2AM), root-essentials-backup.sh (3AM), system-config-sync.sh (4AM), snmp-collect.sh (1min), core-services-backup.sh (1:30AM), watchdog-wg-tunnel.sh (10min), ops-report-collect/send (23:30)"},
|
||||
{"source": "/etc/crontab", "entries": 4, "detail": "stock Debian run-parts hourly/daily/weekly/monthly"},
|
||||
{"source": "/etc/cron.d/", "entries": 5, "detail": "e2scrub_all, kernel(fstrim), php(sessionclean), sysstat - all stock; reap-chrome (custom, 30min, added 2026-09-09 DR fix)"},
|
||||
{"source": "/root/.hermes/cron/jobs.json", "entries": 87, "detail": "full list in tool log; 82 last_status=ok, 5 last_status=error (Doc-Live Verify, Security Compliance Check, OSINT Tool Daily Discovery, Super Search Daily Discovery, Nous LLM Pricing Weekly Report - all Hermes-internal automation, not customer migration scope)"}
|
||||
],
|
||||
"certs": {
|
||||
"acme_count": 47,
|
||||
"self_signed_internal_count": 7,
|
||||
"ca_providers": ["Let's Encrypt (acme-v02.api.letsencrypt.org)", "ZeroSSL (api.hotnow.io only)"],
|
||||
"self_signed_hosts": ["webmail.verdicttank.com", "webmail.rfptank.com", "webmail.radartank.com", "webmail.timapta.org", "webmail.transitpin.com", "voipsimplicity.itpropartner.com", "forefront.itpropartner.com"],
|
||||
"nearest_expiry": "2026-09-15/16 (self-signed webmail/internal stubs, low stakes)",
|
||||
"furthest_expiry": "2026-12-13 (gps.fleettracker360.com)",
|
||||
"note": "app4 must pre-issue its own certs for every moving hostname before DNS cutover; Core certs are not portable without importing Caddy's TLS storage"
|
||||
},
|
||||
"data_paths": [
|
||||
{"path": "/root/docker/uptime-kuma/data", "size": "487M", "owner": "uptime-kuma"},
|
||||
{"path": "/root/docker/super-search", "size": "1.1G", "owner": "super-search MCP"},
|
||||
{"path": "/opt/voice-agent", "size": "470M", "owner": "voice-agent + voice-agent-stt"},
|
||||
{"path": "/opt/rally", "size": "294M", "owner": "rally"},
|
||||
{"path": "/opt/osint-api", "size": "323M", "owner": "osint-api"},
|
||||
{"path": "/root/docker/osint-person-mcp", "size": "314M", "owner": "osint-person MCP"},
|
||||
{"path": "/opt/seemytrip", "size": "212M", "owner": "seemytrip"},
|
||||
{"path": "/opt/ops-portal", "size": "148M", "owner": "ops-portal"},
|
||||
{"path": "/opt/shopping-cart", "size": "147M", "owner": "shopping-cart"},
|
||||
{"path": "/root/docker/dre-mcp", "size": "171M", "owner": "dre-mcp"},
|
||||
{"path": "/root/docker/twilio-mcp", "size": "171M", "owner": "twilio-mcp"},
|
||||
{"path": "/root/shark-game", "size": "127M", "owner": "shark-game"},
|
||||
{"path": "/root/hermes-assistant", "size": "169M", "owner": "hermes-assistant"},
|
||||
{"path": "/opt/hermes-voice", "size": "123M", "owner": "hermes-voice"},
|
||||
{"path": "/root/projects/diglocate", "size": "118M", "owner": "diglocate-api"},
|
||||
{"path": "/opt/dre-portal", "size": "110M", "owner": "dre-portal"},
|
||||
{"path": "/var/lib/docker/volumes/prometheus_data/_data", "size": "114.9M", "owner": "prometheus"},
|
||||
{"path": "/opt/pta-registration", "size": "53M", "owner": "pta-registration"},
|
||||
{"path": "/opt/pta-survey", "size": "50M", "owner": "survey-registration"},
|
||||
{"path": "/root/docker/timetrex", "size": "50M", "owner": "timetrex"},
|
||||
{"path": "/root/hotnow-api", "size": "66M", "owner": "hotnow-api"},
|
||||
{"path": "/root/intelsight-api", "size": "64M", "owner": "intelsight-api"},
|
||||
{"path": "/root/docker/pry", "size": "66M", "owner": "pry"},
|
||||
{"path": "/root/projects/pipeline", "size": "30M", "owner": "pipeline-api"},
|
||||
{"path": "/root/projects/auth", "size": "32M", "owner": "auth-api"},
|
||||
{"path": "/var/lib/docker/volumes/grafana_data_final/_data", "size": "15M", "owner": "grafana"},
|
||||
{"path": "/opt/verdicttank", "size": "15M", "owner": "verdicttank-api + worker"},
|
||||
{"path": "/var/lib/postgresql", "size": "71M", "owner": "host PostgreSQL"},
|
||||
{"path": "/root/docker/docuseal-dre/data", "size": "2.5M", "owner": "docuseal-dre"},
|
||||
{"path": "/root/docker/docuseal-modelortho/data", "size": "1.4M", "owner": "docuseal-modelortho"},
|
||||
{"path": "/root/docker/docuseal/data", "size": "688K", "owner": "docuseal"},
|
||||
{"path": "/opt/transitpin", "size": "496K", "owner": "transitpin"},
|
||||
{"path": "/opt/microbin/data", "size": "4.0K", "owner": "microbin"},
|
||||
{"path": "/var/lib/redis", "size": "8.0K", "owner": "redis-server"}
|
||||
],
|
||||
"disk_summary": {
|
||||
"filesystem": "/dev/vda4",
|
||||
"total": "503G",
|
||||
"used": "157G (33%)",
|
||||
"available": "326G",
|
||||
"var_lib_docker": "31G",
|
||||
"root_docker": "5.8G",
|
||||
"opt": "6.4G"
|
||||
},
|
||||
"memory_cpu_summary": {
|
||||
"ram_total": "15Gi",
|
||||
"ram_used": "4.9Gi",
|
||||
"ram_buff_cache": "11Gi",
|
||||
"swap_total": "8.0Gi",
|
||||
"swap_used": "1.9Gi",
|
||||
"vcpus": 8
|
||||
},
|
||||
"ambiguities": [
|
||||
"grafana, prometheus, telegraf, mikrotik-exporter containers: no discoverable docker-compose.yml at the expected paths for grafana specifically; likely started via manual docker run or an unfound compose file",
|
||||
"crawl4ai.service and hermes-control-deck.service are disabled/inactive; hermes-control-deck's configured port 8200 collides with the actively-running pipeline-api.service - would fail to start if ever enabled",
|
||||
"IPsec (charon/strongSwan, UDP 500/4500) and L2TP (xl2tpd, UDP 1701) are running with no documented owner found in the reviewed plan/backup docs",
|
||||
"UFW rule '8890/tcp temp map server' has nothing currently listening on that port; purpose and lifecycle unclear",
|
||||
"crm.debtrecoveryexperts.com and crm.intelsight.io Caddy blocks proxy to localhost:3003 but nothing listens there on Core (Twenty CRM migrated to App1); unclear if these routes are dead code or missing a proxy hop",
|
||||
"gitea-runner.service binds 0.0.0.0:35849 - purpose of this specific port from Core's perspective alone is not fully confirmed",
|
||||
"Grafana is bound to *:3002 on all interfaces, not evidently proxied through Caddy and not restricted to localhost/Tailscale - worth confirming this exposure is intentional",
|
||||
"TimeTrex's internal PostgreSQL 16 instance could not be queried from host psql (container-internal, not published); migration must use docker exec pg_dump as timetrex-backup.sh already does",
|
||||
"diglocate-api, ft360-mcp, pipeline-api, pry, rally, seemytrip, shark-game, shopping-cart, transitpin have no discoverable .env or explicit external DB config beyond their local SQLite files found in a shallow scan - deeper source review would be needed to fully rule out hidden Postgres/Redis dependencies"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,287 @@
|
||||
# Core Service Inventory — 2026-09-15
|
||||
|
||||
Host: Core (152.53.192.33, netcup RS 2000, 8 vCPU / 15 GB / 503 GB, Debian 13, Manassas VA)
|
||||
Purpose: exhaustive, evidence-based inventory to plan the app4 migration (see `app4-migration-plan.md`).
|
||||
Method: read-only commands only, run directly on Core. No service was started, stopped, or restarted. `hermes-maintenance` was never invoked. The Hermes gateway (`hermes-gateway.service`, user unit) was left untouched and observed only via `systemctl --user status`.
|
||||
|
||||
All commands quoted below were actually executed on Core on 2026-09-15. Numbers are copy-pasted from real output, not estimated.
|
||||
|
||||
## Headline counts (verified)
|
||||
|
||||
| Metric | Count | Command |
|
||||
|---|---|---|
|
||||
| Docker containers (all, all running) | 13 | `docker ps -a --format ...` |
|
||||
| systemd services in `running` state | 57 | `systemctl list-units --type=service --state=running` |
|
||||
| Caddy site blocks in `/etc/caddy/Caddyfile` | 49 | `grep -cE '^[a-zA-Z0-9].*\{$\|^http://.*\{$' /etc/caddy/Caddyfile` |
|
||||
| PostgreSQL databases (excl. template0/1/postgres) | 1 (`hotnow`) | `sudo -u postgres psql -c '\l+'` |
|
||||
| Hermes cron jobs (`/root/.hermes/cron/jobs.json`) | 87 | Python count of `jobs[]` |
|
||||
| Listening TCP sockets | 66 | `ss -tlnp` |
|
||||
| Listening UDP sockets | 14 | `ss -ulnp` |
|
||||
| TLS certs, real ACME (Let's Encrypt / ZeroSSL) | 47 | `find .../certificates -mindepth 2 -maxdepth 2 ! -path '*/local/*'` |
|
||||
| TLS certs, Caddy `tls internal` self-signed | 7 | `find .../certificates/local -mindepth 1 -maxdepth 1` |
|
||||
| System crontab entries (root) | 12 | `crontab -l` |
|
||||
| `/etc/cron.d/*` files | 6 (all stock Debian/PHP/sysstat except `reap-chrome`) | `ls /etc/cron.d` |
|
||||
|
||||
## Hostile surprises found (read this section first)
|
||||
|
||||
1. **The app4 migration plan's "stays on Core" and "moves to app4" lists are both incomplete.** Core is running roughly 30 additional customer-facing FastAPI/Node/systemd services that the plan document never mentions: DigLocate API, DRE MCP + DRE Portal, FT360 MCP, HotNow API (with its own Postgres DB `hotnow` and its own Redis DB 1), IntelSight API, OSINT API, OSINT Person MCP, Outlook upload receiver, Pipeline API, PRY (OSINT search backend), PTA registration + survey, Rally family calendar, SeeMyTrip, Shark Attack Fantasy Game, Shopping Cart Builder, SNMP metrics server, Transitpin WebSocket relay, Twilio MCP, VerdictTank API + worker, Voice Agent + STT, Auth API, plus two disabled-but-present services (`crawl4ai`, `hermes-control-deck`). See `ambiguities[]` and the systemd table below. Every one of these has a live Caddy site (see caddy_sites) and needs an explicit stays/moves decision.
|
||||
2. **DocuSeal did NOT move to App1 as the backup-plan.md's "Migration History" section claims.** Three DocuSeal containers (`docuseal`, `docuseal-dre`, `docuseal-modelortho`) are running on Core right now on ports 8091/8094/8092, backed by bind-mounted `./data` dirs, each with its own `.env`. `backup-plan.md` still lists DocuSeal backups running against App1 (`app1/docuseal/`). The app4-migration-plan.md's "Inventory caveat" already flagged this: **confirmed here, DocuSeal is on Core, not migrated.** Same finding for SearXNG: it is running on Core (`searxng` container, port 127.0.0.1:8888, container up 2 hours) despite backup-plan.md marking `core/searxng/` as a stale/removed path.
|
||||
3. **PostgreSQL on Core has essentially nothing "shared."** Only one real customer database exists: `hotnow` (16 MB, owned by role `hotnow_app`). TimeTrex runs its own **separate PostgreSQL 16 instance inside its own Docker container** (bind-mounted `/root/docker/timetrex/database`), not the host Postgres 17 instance. The migration plan's assumption of "a shared Postgres/Redis instance... provisioned fresh on app4" understates this: there are two independent Postgres engines on Core (host 17.10 for HotNow, containerized 16 for TimeTrex), and they do not share data.
|
||||
4. **Redis on Core is nearly idle** (`used_memory` 741 KB, 0 keys in DB0, `DBSIZE 0`), but HotNow's `.env` explicitly points at `redis://localhost:6379/1`, meaning host Redis has at least one active customer consumer despite the migration plan describing Redis only as "shared cache" with "consumers TBD in Phase 0." HotNow is that consumer, but note DB1 was also empty at scan time (idle app or nothing warm).
|
||||
5. **`crm.debtrecoveryexperts.com` and `crm.intelsight.io` are dead Caddy routes.** Both proxy to `localhost:3003` (Twenty CRM), but nothing is listening on port 3003 — `ss -tlnp | grep 3003` returned nothing. Twenty CRM migrated to App1 per backup-plan.md, but the Caddy site blocks pointing at it were never removed from Core's Caddyfile. These will 502 today; they are not "moving" apps, they are stale entries to clean up (or confirm App1 handles them via a different route/CNAME not visible from Core).
|
||||
6. **UFW has two apparently-orphaned rules**: `8080/tcp on tailscale0` labeled "Vaultwarden via Tailscale" (Vaultwarden migrated off Core, port not listening) and `8890/tcp` labeled "temp map server" (also not listening). Neither is a running risk today but both are stale firewall state that should not be copied to app4.
|
||||
7. **13 `.env`/config files carry live secrets** (SMTP creds, LetterStream API keys, Docuseal API tokens, Hexclave keys, Twilio SIDs, provider API keys for a dozen LLM vendors, RingLogix creds, netcup creds) all in `/root/.hermes/.env`, `/opt/dre-portal/.env`, `/etc/verdicttank.env`, `/root/hotnow-api/.env`, `/opt/voice-agent/.env`, `/opt/hermes-voice/.env`. None of these are backed up in plaintext by policy (`auth-api-backup.sh` explicitly redacts values), but they must be manually and securely transferred to app4, not rsynced as part of a generic root-essentials backup.
|
||||
8. **Backup coverage gaps confirmed for the newly-discovered services.** No backup script exists for: DigLocate, HotNow, IntelSight, Pipeline, PRY, PTA registration, PTA survey, SeeMyTrip, Shark Game, Shopping Cart, Survey (pta-survey duplicate), Voice Agent, DRE Portal. Each of these has its own SQLite DB (see `data_paths[]`) with zero S3 backup today. `backup-plan.md`'s "Unbacked Services: none" claim (dated 2026-08-08) is now stale; it predates all of these services appearing in `systemctl list-units`.
|
||||
9. **Postgres cluster mismatch:** host runs PostgreSQL **17.10** (`postgresql@17-main.service`); TimeTrex's containerized instance is PostgreSQL **16**. A wholesale `pg_dumpall` migration approach from the plan's section 5 will not work across both without per-instance handling.
|
||||
|
||||
## containers[] — Docker (13 total, all running)
|
||||
|
||||
| name | image | ports | restart policy | bind mounts / volumes | compose file | classification |
|
||||
|---|---|---|---|---|---|---|
|
||||
| docuseal | docuseal/docuseal:latest | 127.0.0.1:8091->3000 | always | bind `/root/docker/docuseal/data` -> `/data` | `/root/docker/docuseal/docker-compose.yml` | move-to-app4 |
|
||||
| docuseal-dre | docuseal/docuseal:latest | 127.0.0.1:8094->3000 | always | bind `/root/docker/docuseal-dre/data` -> `/data` | `/root/docker/docuseal-dre/docker-compose.yml` | move-to-app4 |
|
||||
| docuseal-modelortho | docuseal/docuseal:latest | 127.0.0.1:8092->3000 | always | bind `/root/docker/docuseal-modelortho/data` -> `/data` | `/root/docker/docuseal-modelortho/docker-compose.yml` | move-to-app4 |
|
||||
| timetrex | skewll/timetrex:latest | 127.0.0.1:8085->80 | unless-stopped | bind storage/logs/database/ini.php | `/root/docker/timetrex/docker-compose.yml` | move-to-app4 |
|
||||
| microbin | danielszabo99/microbin:latest | 127.0.0.1:8260->8080 | unless-stopped | bind `/opt/microbin/data` -> `/app/pasta_data` | `/opt/microbin/docker-compose.yml` | move-to-app4 |
|
||||
| uptime-kuma | louislam/uptime-kuma:latest | 0.0.0.0:3001->3001 | unless-stopped | bind `/root/docker/uptime-kuma/data` -> `/app/data` | `/root/docker/uptime-kuma/docker-compose.yml` | move-to-app4 |
|
||||
| grafana | grafana/grafana:11.4.0 | 0.0.0.0:3002 (host network, exposed by container itself) | unless-stopped | named volume `grafana_data_final` -> `/var/lib/grafana` | none found (manual `docker run`?) | stays-on-Core |
|
||||
| prometheus | prom/prometheus:latest | none published (host network via container) 9090 exposed by container | unless-stopped | bind prometheus.yml, named volume `prometheus_data`, bind textfile dir | `/root/docker/monitoring/...` (compose not directly inspected) | stays-on-Core |
|
||||
| telegraf | telegraf:latest | 9273 (container) | unless-stopped | bind `/root/docker/monitoring/telegraf/telegraf.conf` | monitoring compose | stays-on-Core |
|
||||
| mikrotik-exporter | swoga/mikrotik-exporter:latest | 127.0.0.1:9436->9436 | unless-stopped | bind config.yml | monitoring compose | stays-on-Core |
|
||||
| searxng | searxng/searxng:latest | 127.0.0.1:8888->8080 | always | bind searxng-data, searxng-themes; named cache volume | `/root/docker/searxng/docker-compose.yml` | stays-on-Core (Super Search backend; NOTE: backup-plan.md incorrectly marks this removed) |
|
||||
| browserless | browserless/chrome:latest | 0.0.0.0:3000->3000 | always | none | none found | stays-on-Core (Hermes browser dependency) |
|
||||
| camofox-browser | camofox-browser:dataimpulse | 0.0.0.0:9377->9377 | unless-stopped | none | none found | stays-on-Core (Hermes stealth browser dependency) |
|
||||
|
||||
Sizes (`docker system df -v`): total image space several GB (largest: browserless 3.06GB, camofox 2.27GB, timetrex 1.7GB, twentycrm 1.12GB image present but unused/0 containers, kokoro-fastapi 3.63GB image present but unused). Named volumes of note: `prometheus_data` 118.9MB, `grafana_data_final` 14.65MB, `grafana_data` 50.13MB (older/orphaned), `grafana_data_v3` 14.59MB (orphaned), `twenty_db-data` 71.46MB (orphaned, Twenty CRM no longer runs on Core).
|
||||
|
||||
## systemd_units[] — running, non-stock (57 running total; below are the added/customer-relevant ones)
|
||||
|
||||
| unit | purpose | exec | port | env file | data path | classification |
|
||||
|---|---|---|---|---|---|---|
|
||||
| auth-api.service | ITPP Auth API (SSO) | uvicorn server:app :8500 | 127.0.0.1:8500 | `/root/projects/auth/.env` | `/root/projects/auth/auth.db` (32M dir) | needs-decision (auth.itpropartner.com is customer SSO; plan doesn't mention it) |
|
||||
| diglocate-api.service | 811 locate ticket mgmt API | uvicorn main:app :8000 | 127.0.0.1:8000 | none found | `/root/projects/diglocate` (118M) | move-to-app4 |
|
||||
| dre-mcp.service | DRE MCP server (Hermes tool) | python server.py | n/a (stdio/mcp) | `/root/.hermes/.env` | `/root/docker/dre-mcp` (171M) | needs-decision (Hermes MCP but customer=DRE data) |
|
||||
| dre-portal.service | DRE customer portal API | uvicorn app.main:app :8093 | 127.0.0.1:8093 | `/opt/dre-portal/.env` | `/opt/dre-portal/data/dre.db` (110M dir) | move-to-app4 |
|
||||
| ft360-mcp.service | FT360 (FleetTracker360) MCP server | python server.py | n/a | none found | `/root/docker/ft360-mcp` (60K) | needs-decision |
|
||||
| gitea-runner.service | Gitea Actions Runner (core) | act_runner daemon | :35849 | none found | `/var/lib/gitea-runner` | stays-on-Core (CI runner tied to Core, not customer-facing) |
|
||||
| hermes-assistant.service | Hermes Assistant PWA backend | python server.py | 127.0.0.1:8082 (python3 pid 1098) | none found | `/root/hermes-assistant` (169M) | stays-on-Core (Hermes-related) |
|
||||
| hermes-browser.service | Headless Chromium (CDP) | chrome --remote-debugging-port=9222 | 127.0.0.1:9222 | n/a | n/a | stays-on-Core (Hermes dependency) |
|
||||
| hermes-socat-8787.service | port forward for HermesX mobile | socat 8787->8642 | 0.0.0.0:8787 | n/a | n/a | stays-on-Core |
|
||||
| hermes-voice.service | Hermes Voice (SvelteKit) | node build/index.js | 127.0.0.1:4331 | `/opt/hermes-voice/.env` | `/opt/hermes-voice` (123M) | needs-decision (named "Hermes" but proxied at voice.itpropartner.com, a customer-facing plan item) |
|
||||
| host-metrics-exporter.service | systemd/docker/disk/mem exporter | python | n/a (writes textfile) | n/a | n/a | stays-on-Core |
|
||||
| hotnow-api.service | HotNow API backend | uvicorn main:app :8001 | 127.0.0.1:8001 | `/root/hotnow-api/.env` | `/root/hotnow-api` (66M) + Postgres db `hotnow` (16MB) + Redis DB1 | move-to-app4 |
|
||||
| intelsight-api.service | IntelSight API | python server.py :8099 | 127.0.0.1:8099 | none found | `/root/intelsight-api/intelsight.db` (64M dir) | move-to-app4 |
|
||||
| node_exporter.service | Prometheus node exporter | node_exporter | 0.0.0.0:9100 | n/a | n/a | stays-on-Core |
|
||||
| ops-portal.service | ITPP Ops Portal backend | uvicorn server:app :8090 | 127.0.0.1:8090 | `/root/.hermes/.env` | `/opt/ops-portal/ops.db` (148M dir) | move-to-app4 (per plan) |
|
||||
| osint-api.service | OSINT Tool Daily Discovery API | uvicorn api:app :8100 | 127.0.0.1:8100 | `/root/.hermes/.env` | `/opt/osint-api` (323M) | needs-decision |
|
||||
| osint-person.service | OSINT Person MCP server | python server.py | n/a | `/root/.hermes/.env` | `/root/docker/osint-person-mcp` (314M) | needs-decision (Hermes MCP tool) |
|
||||
| outlook-upload.service | Outlook folder upload receiver | python upload_server.py | 127.0.0.1:8240 | none | `/root/upload-staging` | needs-decision |
|
||||
| pipeline-api.service | Project Pipeline API (customer portal backend) | python server.py :8200 | 127.0.0.1:8200 | none found | `/root/projects/pipeline/pipeline.db` (30M dir) | move-to-app4 |
|
||||
| pry.service | PRY unified OSINT search backend | python server.py :8905 | 127.0.0.1:8905 | `/root/.hermes/.env` | `/root/docker/pry` (66M) | needs-decision |
|
||||
| pta-registration.service | TIMAPTA Membership Registration | uvicorn server:app :8114 | 127.0.0.1:8114 | none | `/opt/pta-registration/pta.db` (53M dir) | move-to-app4 |
|
||||
| rally.service | Rally Family Calendar | python run.py :8105 | 127.0.0.1:8105 | none found | `/opt/rally/data/rally.db` (294M dir) | move-to-app4 |
|
||||
| seemytrip.service | SeeMyTrip media pipeline | uvicorn server:app :8113 | 127.0.0.1:8113 | none | `/opt/seemytrip/data/seemytrip.db` (212M dir) | move-to-app4 |
|
||||
| shark-game.service | Shark Attack Fantasy Game backend | python server.py :8083 | 0.0.0.0:8083 | none | `/root/shark-game/backend/game.db` (127M dir) | move-to-app4 |
|
||||
| shopping-cart.service | Shopping Cart Builder | uvicorn app:app :8101 | 127.0.0.1:8101 | none | `/opt/shopping-cart` (147M) | move-to-app4 |
|
||||
| snmp-metrics.service | SNMP metrics HTTP server | python snmp-http-server.py :8105-adjacent (0.0.0.0:8105) | 0.0.0.0:8105 | n/a | n/a | stays-on-Core |
|
||||
| super-search.service | Super Search MCP server | python server.py :8899 | 0.0.0.0:8899 | `/root/.hermes/.env` | `/root/docker/super-search` (1.1G) | stays-on-Core (Hermes MCP) |
|
||||
| survey-registration.service | TIMA Location Survey | uvicorn server:app :8115 | 127.0.0.1:8115 | none | `/opt/pta-survey/survey.db` (50M dir) | move-to-app4 |
|
||||
| transitpin.service | TransitPin WebSocket relay | node server.js | 127.0.0.1:8210 | none | `/opt/transitpin` (496K) | move-to-app4 |
|
||||
| twilio-mcp.service | Twilio MCP server | python server.py | n/a | `/root/.hermes/.env` | `/root/docker/twilio-mcp` (171M) | stays-on-Core (Hermes MCP; also feeds voice stack) |
|
||||
| verdicttank-api.service | VerdictTank form handler + PDF gen | python api.py :8201 | 127.0.0.1:8201 | `/etc/verdicttank.env` | `/opt/verdicttank/users.db` (15M dir) | move-to-app4 |
|
||||
| verdicttank-worker.service | VerdictTank review worker (multi-model panel) | python worker.py | n/a | `/etc/verdicttank.env` | `/opt/verdicttank` | move-to-app4 |
|
||||
| voice-agent-stt.service | Voice Agent STT (faster-whisper) | uvicorn stt_server:app :9000 | 127.0.0.1:9000 | none | `/opt/voice-agent` (470M) | needs-decision (plan lists voice stack for app4 but not this STT sub-service explicitly) |
|
||||
| voice-agent.service | Voice Agent (open-source stack) | uvicorn agent_server:app :9101 | 127.0.0.1:9101 | `/opt/voice-agent/.env` | `/opt/voice-agent` | move-to-app4 (voice.itpropartner.com / voice-open.itpropartner.com) |
|
||||
| wazuh-agent.service | Wazuh SIEM agent | (Wazuh binary) | n/a | n/a | n/a | stays-on-Core (security agent, host-level) |
|
||||
| caddy.service | reverse proxy | caddy run | 152.53.192.33:80/443 | `/etc/caddy/Caddyfile` | `/var/lib/caddy` (TLS store) | stays-on-Core per plan (routes to be trimmed after cutover) |
|
||||
| postgresql@17-main.service | Postgres host instance | postgres | 127.0.0.1:5432, [::1]:5432 | n/a | `/var/lib/postgresql` (71M) | needs-decision (only DB is `hotnow`, which is moving) |
|
||||
| redis-server.service | Redis | redis-server | 127.0.0.1:6379, [::1]:6379 | n/a | `/var/lib/redis` (8K, no persistence file written; `appendonly no`, RDB save schedule set) | needs-decision (HotNow is the only confirmed consumer found) |
|
||||
|
||||
Disabled-but-present units (not running, found via unit files, no `.service` entry in the running list): `crawl4ai.service` (Crawl4AI extraction microservice, port 8910, disabled/inactive) and `hermes-control-deck.service` (backend API, port 8200 conflicts with pipeline-api's port 8200 if ever enabled — flagged in ambiguities). Both should be accounted for even though inactive.
|
||||
|
||||
Hermes-related systemd units confirmed via `systemctl --user list-units` (separate user-level manager, NOT touched or restarted): `hermes-gateway.service` (active running, PID 100998, 5.3G RAM) and `ssh-agent.service`. These stay on Core by definition of the migration and were only observed, never controlled.
|
||||
|
||||
## ports[] — Listening sockets (66 TCP, 14 UDP)
|
||||
|
||||
Representative table (full list captured via `ss -tlnp`/`ss -ulnp`, see raw command output in this task's tool log for the complete 80-row set):
|
||||
|
||||
| port | bind | process | service |
|
||||
|---|---|---|---|
|
||||
| 443 | 152.53.192.33 | caddy | Caddy HTTPS (customer + core sites) |
|
||||
| 80 | 152.53.192.33 | caddy | Caddy HTTP redirect |
|
||||
| 443 | 100.71.155.7 (Tailscale) | tailscaled | Tailscale-only HTTPS |
|
||||
| 22 | 0.0.0.0 / [::] | sshd | SSH |
|
||||
| 5432 | 127.0.0.1 / [::1] | postgres | Host PostgreSQL 17 |
|
||||
| 6379 | 127.0.0.1 / [::1] | redis-server | Redis |
|
||||
| 3000 | 0.0.0.0 / [::] | docker-proxy | browserless |
|
||||
| 3001 | 0.0.0.0 / [::] | docker-proxy | uptime-kuma |
|
||||
| 3002 | * (all interfaces) | grafana | Grafana (not port-mapped through Caddy's default_bind; exposed directly) |
|
||||
| 9377 | 0.0.0.0 / [::] | docker-proxy | camofox-browser |
|
||||
| 8090-8115, 8200-8210, 8500, 8899-8905 range | mostly 127.0.0.1 | uvicorn/python | see systemd_units table above |
|
||||
| 8787 | 0.0.0.0 | socat | Hermes API forward for mobile |
|
||||
| 8642 | 0.0.0.0 | hermes | Hermes gateway internal API (do not touch) |
|
||||
| 9090 | * | prometheus | Prometheus |
|
||||
| 9100 | * | node_exporter | Node exporter |
|
||||
| 9273 | * | telegraf | Telegraf |
|
||||
| 9436 | 127.0.0.1 | docker-proxy | mikrotik-exporter |
|
||||
| 9222 | 127.0.0.1 | chrome | Hermes headless browser CDP |
|
||||
| 25 | 127.0.0.1 / [::1] | exim4 | local mail relay |
|
||||
| 500, 4500 (UDP) | 0.0.0.0 / [::] | charon (strongSwan) | IPsec (unclear purpose — see ambiguities) |
|
||||
| 1701 (UDP) | 0.0.0.0 | xl2tpd | L2TP (unclear purpose — see ambiguities) |
|
||||
| 51821 (UDP) | 0.0.0.0 / [::] | (no owning process shown) | WireGuard, per UFW comment |
|
||||
| 41641 (UDP) | 0.0.0.0 / [::] | tailscaled | Tailscale |
|
||||
| 5353 (UDP) | 0.0.0.0 / [::] | avahi-daemon | mDNS |
|
||||
|
||||
## caddy_sites[] — 49 site blocks
|
||||
|
||||
Extracted from `/etc/caddy/Caddyfile` (513 lines; global block sets `default_bind 152.53.192.33`, `email info@itpropartner.com`). Full list of hostnames and their backend targets:
|
||||
|
||||
| hostname(s) | backend | notes |
|
||||
|---|---|---|
|
||||
| core.itpropartner.com | 127.0.0.1:8240, 127.0.0.1:8201, static /var/www | multi-path handler; Ops Portal / VerdictTank API paths mixed in |
|
||||
| sign.itpropartner.com | 127.0.0.1:8091 | DocuSeal |
|
||||
| sign.modelortho.com | 127.0.0.1:8092 | DocuSeal (modelortho) |
|
||||
| sign.debtrecoveryexperts.com | 127.0.0.1:8094 | DocuSeal (DRE) |
|
||||
| ops.itpropartner.com | 127.0.0.1:8090, 127.0.0.1:8100, static | Ops Portal + OSINT API |
|
||||
| shark.iamgmb.com | 127.0.0.1:8083 | Shark Game |
|
||||
| internal.debtrecoveryexperts.com | 127.0.0.1:8093, static, basic_auth | DRE Portal internal, password-protected |
|
||||
| portal.debtrecoveryexperts.com | redirect to my.debtrecoveryexperts.com/start | |
|
||||
| pay.debtrecoveryexperts.com | static | |
|
||||
| my.debtrecoveryexperts.com | 127.0.0.1:8093, static | DRE Portal |
|
||||
| crm.debtrecoveryexperts.com | localhost:3003 | **DEAD — nothing listening on 3003** |
|
||||
| dig.iamgmb.com | 127.0.0.1:8000, static | DigLocate |
|
||||
| uptimekuma.itpropartner.com | localhost:3001 | Uptime Kuma |
|
||||
| gps.fleettracker360.com | 152.53.39.202:8082 (App2, remote proxy) | Traccar on App2, not Core |
|
||||
| my.itpropartner.com | 152.53.241.111:8090 (App3, remote), 127.0.0.1:8200 (pipeline), static | Mixed remote+local |
|
||||
| status.itpropartner.com | 127.0.0.1:8210, 127.0.0.1:3001, static | Transitpin relay + Uptime Kuma API |
|
||||
| track.fleettracker360.com (http only) | 152.53.39.202:5055 (App2, remote) | |
|
||||
| hear.fleettracker360.com | static | |
|
||||
| voice.itpropartner.com | 127.0.0.1:4331 | Hermes Voice / SvelteKit |
|
||||
| voice-open.itpropartner.com | 127.0.0.1:9101 | Voice Agent |
|
||||
| auth.itpropartner.com | 127.0.0.1:8500, static | Auth API |
|
||||
| my.intelsight.io | 127.0.0.1:8099, static | IntelSight |
|
||||
| intelsight.io | static | landing page |
|
||||
| intelsight.iamgmb.com | static | landing page |
|
||||
| schedule.iamgmb.com | static | |
|
||||
| seemytrip.iamgmb.com | 127.0.0.1:8113, static | SeeMyTrip |
|
||||
| rally.iamgmb.com | 127.0.0.1:8105, static | Rally |
|
||||
| shopping.iamgmb.com | 127.0.0.1:8101, 127.0.0.1:8210, static | Shopping Cart |
|
||||
| pry.iamgmb.com (http only) | 127.0.0.1:8905, static | PRY |
|
||||
| crm.intelsight.io | localhost:3003 | **DEAD — same as above** |
|
||||
| www.hotnow.io | redirect to hotnow.io | |
|
||||
| hotnow.io | static | |
|
||||
| app.hotnow.io | static | |
|
||||
| api.hotnow.io | 127.0.0.1:8001 | HotNow API |
|
||||
| admin.hotnow.io | static | |
|
||||
| timetrex.iamgmb.com | 127.0.0.1:8085 | TimeTrex |
|
||||
| webmail.timapta.org, webmail.transitpin.com, webmail.rfptank.com, webmail.radartank.com, webmail.verdicttank.com | redirect to heracles.mxrouting.net, `tls internal` | 5 sites, self-signed certs |
|
||||
| share.itpropartner.com | 127.0.0.1:8260 | Microbin |
|
||||
| verdicttank.com, www.verdicttank.com | 127.0.0.1:8201, static | VerdictTank |
|
||||
| voipsimplicity.itpropartner.com | static, `tls internal` | |
|
||||
| forefront.itpropartner.com | static, `tls internal` | |
|
||||
| ops.verdicttank.com | 127.0.0.1:8201, static | |
|
||||
| register.timapta.org | 127.0.0.1:8114 | PTA registration |
|
||||
| pta.iamgmb.com | 127.0.0.1:8114 | PTA registration (dup route) |
|
||||
| survey.iamgmb.com | 127.0.0.1:8115 | PTA survey |
|
||||
|
||||
`caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile` returned **Valid configuration** (with a formatting-only warning, no functional errors).
|
||||
|
||||
## databases[] — PostgreSQL + Redis
|
||||
|
||||
PostgreSQL 17.10 (host, `postgresql@17-main.service`, port 5432):
|
||||
|
||||
| database | owner | size |
|
||||
|---|---|---|
|
||||
| hotnow | hotnow_app | 16 MB |
|
||||
| postgres | postgres | 7510 kB |
|
||||
| template0 | postgres | 7353 kB |
|
||||
| template1 | postgres | 7582 kB |
|
||||
|
||||
Roles: `hotnow_app` (no special attrs), `postgres` (superuser).
|
||||
|
||||
TimeTrex has its own **containerized PostgreSQL 16** instance (visible only via the bind-mounted `/root/docker/timetrex/database` PGDATA directory, `type=postgres, host=localhost, user=timetrex` in `timetrex.ini.php`). Not queryable from the host `psql` since it lives inside the container's own network namespace / port, and the container does not publish 5432 to the host.
|
||||
|
||||
Redis 8.0.2 (`redis-server.service`, port 6379): `used_memory` 741 KB, `used_memory_rss` 17 MB, `save 3600 1 300 100 60 10000`, `appendonly no`. `DBSIZE` on DB0 and DB1 both returned 0 at scan time. HotNow's `.env` references `redis://localhost:6379/1` confirming at least one real customer consumer even though it was empty when sampled.
|
||||
|
||||
## cron[] — system cron + Hermes cron
|
||||
|
||||
Root system crontab (`crontab -l`, 12 real entries beyond blank/env lines): boys-mail-monitor (hourly + daily), shark-game scraper (daily), shark-draft-reminder (15 min), hermes-backup.sh (1 AM), backup-audit-check.sh (2 AM), root-essentials-backup.sh (3 AM), system-config-sync.sh (4 AM), snmp-collect.sh (every minute), core-services-backup.sh (1:30 AM), watchdog-wg-tunnel.sh (10 min), ops-report-collect/send (23:30).
|
||||
|
||||
`/etc/crontab`: only stock Debian `run-parts` hourly/daily/weekly/monthly entries.
|
||||
|
||||
`/etc/cron.d/`: `e2scrub_all`, `kernel` (fstrim), `php` (session cleanup), `sysstat` — all stock Debian. One custom entry: `reap-chrome` (every 30 min, reaps orphaned Playwright Chrome processes, added per DR note 2026-09-09).
|
||||
|
||||
No other Linux user has a crontab (checked every user in `/etc/passwd`).
|
||||
|
||||
Hermes cron (`/root/.hermes/cron/jobs.json`): **87 jobs total**, `last_status: ok` for the great majority; 5 jobs currently show `last_status: error` (`Doc-Live Verify`, `Security Compliance Check`, `OSINT Tool Daily Discovery`, `Super Search Daily Discovery`, `Nous LLM Pricing Weekly Report`) — these are Hermes-internal automation, not customer-facing infra, and are flagged for the Hermes team separately, not part of this migration's scope. Backup-relevant Hermes cron jobs (docuseal-backup, docuseal-dre-backup, docuseal-modelortho-backup, timetrex-backup, auth-api-backup, and dozens of others for App1/2/3 services) confirm the backup-plan.md schedule is implemented as documented, with the caveat in surprise #8 above (several newly-found Core services have no corresponding job at all).
|
||||
|
||||
## certs[] — TLS
|
||||
|
||||
47 real ACME-issued certificates under `/var/lib/caddy/.local/share/caddy/certificates/acme-v02.api.letsencrypt.org-directory/` and one under the ZeroSSL CA path (`api.hotnow.io`). All checked with `openssl x509 -noout -enddate`; none are expired, nearest expiry is `status.itpropartner.com` and `voipsimplicity`-adjacent (self-signed, see below) in mid-October 2026, furthest is `pta.iamgmb.com` (Dec 9 2026). Full list of hostname -> expiry captured in the JSON companion file.
|
||||
|
||||
7 self-signed (`tls internal`) certs for internal/webmail redirect stubs: `webmail.verdicttank.com`, `webmail.rfptank.com`, `webmail.radartank.com`, `webmail.timapta.org`, `webmail.transitpin.com`, `voipsimplicity.itpropartner.com`, `forefront.itpropartner.com` — these expire within 1 year of issuance and are low-stakes redirect-only stubs.
|
||||
|
||||
Migration note: app4 must pre-issue its own certs for every moving hostname before DNS cutover (per the plan's checklist); the existing Core certs cannot be copied over and reused as-is without importing Caddy's storage, which the plan does not currently call for.
|
||||
|
||||
## data_paths[] — sizes (`du -sh`, all real measurements)
|
||||
|
||||
| path | size | owner service |
|
||||
|---|---|---|
|
||||
| /root/docker/uptime-kuma/data | 487M | uptime-kuma |
|
||||
| /root/docker/super-search/ | 1.1G | super-search MCP (venv-heavy) |
|
||||
| /opt/voice-agent | 470M | voice-agent/voice-agent-stt |
|
||||
| /opt/rally/data/rally.db (dir 294M) | 294M | rally |
|
||||
| /opt/osint-api | 323M | osint-api |
|
||||
| /root/docker/osint-person-mcp | 314M | osint-person MCP |
|
||||
| /opt/seemytrip/data | within 212M dir | seemytrip |
|
||||
| /opt/ops-portal | 148M | ops-portal |
|
||||
| /opt/shopping-cart | 147M | shopping-cart |
|
||||
| /root/docker/dre-mcp | 171M | dre-mcp |
|
||||
| /root/docker/twilio-mcp | 171M | twilio-mcp |
|
||||
| /root/shark-game | 127M | shark-game |
|
||||
| /root/hermes-assistant | 169M | hermes-assistant |
|
||||
| /opt/hermes-voice | 123M | hermes-voice |
|
||||
| /root/projects/diglocate | 118M | diglocate-api |
|
||||
| /opt/dre-portal | 110M | dre-portal |
|
||||
| /var/lib/docker/volumes/prometheus_data/_data | 114.9M | prometheus |
|
||||
| /opt/pta-registration | 53M | pta-registration |
|
||||
| /opt/pta-survey | 50M | survey-registration |
|
||||
| /root/docker/timetrex (container data) | 50M | timetrex |
|
||||
| /root/hotnow-api | 66M | hotnow-api |
|
||||
| /root/intelsight-api | 64M | intelsight-api |
|
||||
| /root/docker/pry | 66M | pry |
|
||||
| /root/projects/pipeline | 30M | pipeline-api |
|
||||
| /root/projects/auth | 32M | auth-api |
|
||||
| /var/lib/docker/volumes/grafana_data_final/_data | 15M | grafana |
|
||||
| /opt/verdicttank | 15M | verdicttank-api/worker |
|
||||
| /var/lib/postgresql | 71M | host PostgreSQL |
|
||||
| /root/docker/docuseal-dre/data | 2.5M | docuseal-dre |
|
||||
| /root/docker/docuseal-modelortho/data | 1.4M | docuseal-modelortho |
|
||||
| /root/docker/docuseal/data | 688K | docuseal |
|
||||
| /opt/transitpin | 496K | transitpin |
|
||||
| /opt/microbin/data | 4.0K | microbin |
|
||||
| /var/lib/redis | 8.0K | redis-server (no RDB dump on disk at scan time) |
|
||||
|
||||
Overall disk: `/dev/vda4` 503G total, 157G used (33%), 326G available. `/var/lib/docker` 31G, `/root/docker` 5.8G, `/opt` 6.4G.
|
||||
|
||||
Memory/CPU snapshot at scan time: 15Gi RAM total, 4.9Gi used, 11Gi buff/cache, 8Gi swap configured with 1.9Gi in use; 8 vCPUs.
|
||||
|
||||
## ambiguities[]
|
||||
|
||||
1. **`grafana`, `prometheus`, `telegraf`, `mikrotik-exporter` containers have no discoverable `docker-compose.yml`** in the paths their bind mounts imply beyond `/root/docker/monitoring/...` for the latter three; Grafana's compose file was not found at all (likely started via a one-off `docker run` or a compose file elsewhere not searched). Confirm before any monitoring-stack changes.
|
||||
2. **`crawl4ai.service` and `hermes-control-deck.service`** are defined but disabled/inactive. `hermes-control-deck` binds port 8200, the same port `pipeline-api.service` is actively using — if `hermes-control-deck` is ever enabled it will fail to bind. Needs a decision on whether either is still needed; neither showed up in the migration plan.
|
||||
3. **IPsec (`charon`/strongSwan) on UDP 500/4500 and L2TP (`xl2tpd`) on UDP 1701** are running with no obvious owner in the docs reviewed. Unclear if this is a legacy VPN endpoint for a customer or leftover config. Flag for the infra team, do not assume it is decommission-safe.
|
||||
4. **UFW rule `8890/tcp "temp map server"`** — nothing is listening on this port today. Unclear what service this was for or whether it is scheduled to return.
|
||||
5. **`crm.debtrecoveryexperts.com` and `crm.intelsight.io`** Caddy blocks point to `localhost:3003` with nothing listening. Either Twenty CRM on App1 is reached by a different mechanism not visible from Core (e.g. these routes are actually dead code) or there's a missing local proxy. Needs confirmation before deciding whether these Caddy blocks move, get deleted, or get repointed to App1 directly.
|
||||
6. **`gitea-runner.service`** listens on `*:35849` — purpose of this port (act_runner's own control port) unconfirmed from Core alone; likely not customer-facing, but flagged since it's a nonstandard high port bound to all interfaces.
|
||||
7. **Grafana bound to `*:3002`** on all interfaces (not proxied through Caddy's default_bind and not restricted to localhost or Tailscale) — worth checking if this is intentionally public; it wasn't found behind any Caddy site block in the scanned Caddyfile.
|
||||
8. **TimeTrex's internal PostgreSQL 16** could not be queried from the host `psql` (different engine version/instance inside the container, not published to host network). A migration plan needs a container-internal `pg_dump` step (documented in `timetrex-backup.sh`, which already does `docker exec timetrex ... pg_dump`), not a host-level one.
|
||||
9. Several running services have **no `.env` and no discoverable database config found in a shallow scan** (`diglocate-api`, `ft360-mcp`, `intelsight-api` — wait, intelsight-api does have `intelsight.db`, `pipeline-api`, `pry`, `rally`, `seemytrip`, `shark-game`, `shopping-cart`, `transitpin`) beyond the SQLite files already listed; deeper source inspection would be needed to confirm each has no hidden Postgres/Redis dependency the shallow `.env` grep missed.
|
||||
|
||||
## Backup coverage summary (cross-checked against backup-plan.md)
|
||||
|
||||
Covered by an existing script (verified script exists, not verified last-run freshness beyond what backup-plan.md states): Hermes (full + live sync), /root essentials, Grafana, Uptime Kuma, Docker volumes, Prometheus, Auth API, DocuSeal (x1 base + x2 named backups for -dre and -modelortho), TimeTrex.
|
||||
|
||||
**No backup coverage found** for (all newly-discovered in this audit): DigLocate API, FT360 MCP (only stats/export scripts found, not a DB backup), HotNow API + its Postgres DB + Redis DB1, IntelSight API, OSINT API, OSINT Person MCP, Pipeline API, PRY, PTA Registration, PTA Survey, Rally (only debug/dump scripts found, not a scheduled backup job matching this DB path), SeeMyTrip, Shark Game backend DB, Shopping Cart, DRE Portal (`dre.db`), Voice Agent, Voice Agent STT, VerdictTank's `users.db` (VerdictTank does have `hello-*-collect.py` cron jobs but not a DB backup job).
|
||||
|
||||
These gaps should be closed on app4 as part of Phase 1/3 of the migration, not carried over as-is.
|
||||
@@ -11,10 +11,10 @@
|
||||
|
||||
| Key Name | File | Type | Fingerprint (SHA256) | Purpose | Deployed To |
|
||||
|----------|------|------|-----------------------|---------|-------------|
|
||||
| **itpp-infra** | `/root/.ssh/itpp-infra` | ED25519 | `Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ` | Universal server admin key | All servers (Core, app1, app2, app3, app1-bu, home router). wphost02 DECOMMISSIONED (2026-08-28), removed from scope. |
|
||||
| **itpp-infra** | `/root/.ssh/itpp-infra` | ED25519 | `Jxh0bbT9dUV3q1DYYB3hHyhy/1TDj7Q8U4xrVmB38uQ` | Universal server admin key | All servers (Core, app1, app2, app3, app4, core-bu, app1-bu, home router). wphost02 DECOMMISSIONED (2026-08-28), removed from scope. |
|
||||
| **wisp_rsa** | `/root/.ssh/wisp_rsa` | ED25519 | `MxQw1oh90NibSgN2mDbKP+07/jE4FEUEBbFAzuk5DcI` | WISP MikroTik CCR router SSH | Home CCR router (10.77.0.2 via WireGuard) |
|
||||
| **germaine-personal** | `/root/.ssh/germaine-personal` | ED25519 | `dDbLH+bdPFcGU0mm1DpGa43ec0nUZ88YnpCi4p63y3I` | Germaine's personal key (from his machines) | Germaine's devices → Core |
|
||||
| **homelab** | `/root/.ssh/homelab` | ED25519 | `c1nts4wR9EU06/O/k895Pb2tGZublgnGWG6NoQrK/qs` | Homelab Proxmox/QNAP access | vm-host-01, vm-host-02, QNAP NAS |
|
||||
| **homelab** | `/root/.ssh/homelab` | ED25519 | `c1nts4wR9EU06/O/k895Pb2tGZublgnGWG6NoQrK/qs` | Homelab Proxmox/QNAP access | vm-host-01, QNAP NAS (vm-host-02 retired 2026-08-30) |
|
||||
| **siteground.key** | `/root/.ssh/siteground.key` | RSA (encrypted) | N/A (RSA, encrypted) | SiteGround SFTP backup (port 18765) | SiteGround shared hosting |
|
||||
| **authorized_keys** | `/root/.ssh/authorized_keys` | — | — | Who can SSH into Core | Core (this server) |
|
||||
|
||||
@@ -39,7 +39,9 @@ homelab.pub: ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIHT+727Cti4cZ2x6CiYDeDKZ9
|
||||
| **app1** | 152.53.36.131 | netcup RS 4000 | Via root or ippadmin+sudo | Root password in Vaultwarden |
|
||||
| **app2** | 152.53.39.202 | netcup RS 4000 | Via root or ippadmin+sudo | Root password in Vaultwarden |
|
||||
| **app3** | 152.53.241.111 | netcup RS 4000 | Via root or ippadmin+sudo | Root password in Vaultwarden |
|
||||
| **app1-bu** | 5.161.225.131 | Hetzner CPX21 | itpp-infra SSH key | Warm standby (core-bu) |
|
||||
| **app1-bu** | 5.161.225.131 | Hetzner CPX21 | itpp-infra SSH key | Warm standby for Core, **Hetzner = the only non-netcup box**. Legacy/superseded once `core-bu` (below) is proven. |
|
||||
| **core-bu** | 159.195.204.203 | netcup RS 2000 G12 | Via root or ippadmin+sudo (key only) | **Core's warm standby** (Nuremberg). Console root pw: `/root/.hermes/references/new-servers-2026-09-15.md` (chmod 600); Vaultwarden entry pending. Provisioned 2026-09-15. |
|
||||
| **app4** | 159.195.205.80 | netcup RS 4000 G12 | Via root or ippadmin+sudo (key only) | **Core's customer-facing services** (Nuremberg). Console root pw: `/root/.hermes/references/new-servers-2026-09-15.md` (chmod 600); Vaultwarden entry pending. Provisioned 2026-09-15. |
|
||||
|
||||
### Admin Account (all servers)
|
||||
|
||||
|
||||
@@ -0,0 +1,311 @@
|
||||
# migration-plan-app4-core-bu-2026-09-15.md
|
||||
|
||||
**Owner:** IT Pro Partner (Germaine Brown)
|
||||
**Created:** 2026-09-15
|
||||
**Status:** ACTIVE (supersedes the draft `app4-migration-plan.md` of 2026-08-15; that file is kept for history)
|
||||
**Scope:** (a) move Core's customer-facing services onto the new `app4`; (b) stand up `core-bu` as Core's
|
||||
warm standby with a working failover **and failback**; (c) update every document, record and reference that
|
||||
names hosts, IPs or service locations.
|
||||
|
||||
---
|
||||
|
||||
## 1. What is verified today (2026-09-15)
|
||||
|
||||
Everything in this section was measured on the live boxes, not copied from a doc.
|
||||
|
||||
### 1.1 The two new boxes
|
||||
|
||||
| | **app4** | **core-bu** |
|
||||
| --- | --- | --- |
|
||||
| Role | Core's customer-facing services | Core's warm standby |
|
||||
| Hostname | `v2202609377162521278.quicksrv.de` | `v2202609377162521279.megasrv.de` |
|
||||
| IPv4 | `159.195.205.80/22` | `159.195.204.203/22` |
|
||||
| IPv6 | `2a0a:4cc0:c2:bcbf:34b3:8cff:fea2:2892` | `2a0a:4cc0:c2:b3e0:9409:42ff:fe3b:5397` |
|
||||
| Model | netcup RS 4000 G12, 12 vCPU / 32 GB / 1007 GB | netcup RS 2000 G12, 8 vCPU / 16 GB / 503 GB |
|
||||
| Location | Nuremberg (NBG) | Nuremberg (NBG) |
|
||||
| RTT from Core | 100.5 ms | 100.5 ms |
|
||||
|
||||
`core-bu` is the exact twin of Core (same 8 vCPU / 16 GB / 503 GB shape), which is what a standby should be.
|
||||
|
||||
### 1.2 Provisioning status: COMPLETE and verified
|
||||
|
||||
Both boxes were provisioned to the ITPP standard on 2026-09-15 and each item was verified, not asserted:
|
||||
|
||||
- Debian 13, hostname set, timezone `America/New_York`, 8 GB swap (9 GB on app4).
|
||||
- `ippadmin` user with NOPASSWD sudo; `itpp-infra` key installed for **both** `root` and `ippadmin`.
|
||||
- `ufw` **active** (22/80/443, plus 9100 only from Core and from the tailnet).
|
||||
- `fail2ban` active, `unattended-upgrades` active.
|
||||
- Docker CE **29.8.0** + Compose **v5.5.1** (upstream repo, matching app1's `docker-compose-plugin 5.3.1~trixie` family).
|
||||
- `node_exporter` listening on `:9100` - HTTP 200 from Core, **closed** from a third host.
|
||||
- `awscli` + Wasabi credentials in `/root/.aws/credentials` (mode 600).
|
||||
- sshd hardened to the fleet convention: `PermitRootLogin without-password`, `PasswordAuthentication no`,
|
||||
`AllowUsers ippadmin root`. Verified four ways per box (root key login, ippadmin key login + `sudo -n`,
|
||||
password auth refused, non-allowlisted user refused) with the self-reverting lockout guard armed.
|
||||
|
||||
### 1.3 Backups: enrolled and restore-tested
|
||||
|
||||
| Host | Script | Schedule | Destination | Evidence |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| app4 | `root-essentials-backup.sh` | 04:45 ET | `s3://hermes-vps-backups/root-backup/app4/` | first run: upload + in-script download/extract verify OK |
|
||||
| core-bu | `root-essentials-backup.sh` | 05:15 ET | `s3://hermes-vps-backups/root-backup/core-bu/` | first run: upload + in-script download/extract verify OK |
|
||||
|
||||
Note the known gap this replicates: `root-essentials-backup.sh` **excludes `*.db` by design**. Any database
|
||||
these boxes end up hosting needs its own `sqlite3 .backup` / `pg_dump` job, exactly as `anita-mnz` needed
|
||||
`hermes-db-backup.sh`. This is an explicit Phase 4/5 acceptance item, not an optional nicety.
|
||||
|
||||
### 1.4 Monitoring: registered
|
||||
|
||||
A `node_exporter` job was added to the **live** Prometheus config
|
||||
(`/root/docker/monitoring/prometheus/prometheus.yml`) - the job did not previously exist. `up=1` verified for
|
||||
`core:9100`, `app4:9100`, `core-bu:9100`.
|
||||
|
||||
**Finding (pre-existing, not caused by this work):** `node_exporter` is **not running** on app1, app2, app3 or
|
||||
app1-bu; and the node_exporter target list in `/opt/prometheus/prometheus.yml` lives in a file Prometheus never
|
||||
loaded (it still names decommissioned `wphost02` and `178.156.131.57`). Host metrics for the existing fleet were
|
||||
therefore never collected. Tracked in Section 8.
|
||||
|
||||
### 1.5 Credentials
|
||||
|
||||
Console/root credentials for both boxes are recorded in `/root/.hermes/references/new-servers-2026-09-15.md`
|
||||
(mode 600, root only). Password auth is disabled on both boxes, so those passwords are **console/rescue only**.
|
||||
Vaultwarden was **locked** at the time of writing, so the Vaultwarden entries are still owed (Section 9, Q3).
|
||||
|
||||
---
|
||||
|
||||
## 2. End state
|
||||
|
||||
| Host | Role after migration |
|
||||
| --- | --- |
|
||||
| **Core** (152.53.192.33, RS 2000, Manassas) | Hermes + its direct dependencies (browserless, camofox-browser, SearXNG, Super Search MCP), Prometheus/Telegraf/Grafana, mikrotik-exporter, Caddy for Core-local routes. **No customer-facing apps.** |
|
||||
| **app4** (159.195.205.80, RS 4000 G12, Nuremberg) | All customer-facing apps, their databases, and all customer-facing Caddy routes + TLS. |
|
||||
| **core-bu** (159.195.204.203, RS 2000 G12, Nuremberg) | Warm standby for Core (Hermes, its state, its watchdog). Dormant until failover. |
|
||||
| **app1-bu** (5.161.225.131, Hetzner CPX21, Ashburn) | Retirement candidate once `core-bu` is proven. **Only remaining non-netcup box** (Section 3.2). |
|
||||
| **anita-mnz** (159.195.16.30, netcup, Manassas) | Unchanged. Anita's dedicated Hermes box. |
|
||||
| **app1 / app2 / app3** (Manassas) | Unchanged by this plan. |
|
||||
|
||||
---
|
||||
|
||||
## 3. Deviations and risks you must decide on
|
||||
|
||||
### 3.1 app4 is in Nuremberg, not Manassas (NEW, material)
|
||||
|
||||
Measured: **100.5 ms RTT Core -> app4**, versus 0.5 ms Core -> app2 (Manassas) and 1.6 ms -> app1-bu (Ashburn).
|
||||
The old draft assumed Manassas. Consequences:
|
||||
|
||||
- **Customer latency on app4-hosted sites.** Typical US East users add roughly 80-110 ms per round trip versus
|
||||
a Manassas host. For static sites this is mostly invisible; for interactive apps (DocuSeal signing flow,
|
||||
Ops Portal, TimeTrex) it is user-visible.
|
||||
- **Core <-> app4 chatter crosses the Atlantic.** Any Core->app4 API call, monitoring scrape, backup pull, or
|
||||
Caddy proxy hit pays ~100 ms. This is acceptable if app4 is self-contained, and painful if the two are chatty.
|
||||
Design rule for this migration: **app4 must not depend on Core at request time.**
|
||||
- **Benefit, and it is real:** Core (US) and its standby (EU) now fail independently. A Nuremberg outage does
|
||||
not touch Core, and a Manassas outage does not touch the standby. The old pair (Core + app1-bu Ashburn) were
|
||||
1.6 ms apart and shared the US East corridor.
|
||||
|
||||
**Options:** (A) accept Nuremberg and design app4 to be self-contained (recommended, zero cost, boxes are paid);
|
||||
(B) re-order app4 as a Manassas RS 4000 and keep the Nuremberg box as the standby. This is a money decision,
|
||||
so it is yours.
|
||||
|
||||
### 3.2 Provider diversity is now unmet
|
||||
|
||||
Core, app1, app2, app3, app4, core-bu and anita-mnz are **all netcup**. `app1-bu` (Hetzner) is the only other
|
||||
provider, and this plan retires it. Mitigation options: keep app1-bu as the *provider-diverse* last-resort
|
||||
standby even after core-bu is primary (cheapest option, EUR 31.99/mo), or move the off-site backup/DR device to a
|
||||
non-netcup provider. **This must be decided before app1-bu is deleted**, and the DR principle that has governed
|
||||
the org so far ("a netcup outage must not kill both live and standby") is currently **satisfied by geography but
|
||||
not by provider**.
|
||||
|
||||
### 3.3 app4 has no standby of its own
|
||||
|
||||
`app4` becomes the single host for every customer-facing service. If it dies, customer apps are down until
|
||||
S3 restore. `core-bu` is shaped for Core, not for the customer tier (16 GB, and it is meant to be dormant).
|
||||
Options: (A) accept S3-restore RTO for app4; (B) let core-bu carry a cold/secondary copy of app4's data;
|
||||
(C) budget a second app-tier box. Recommend a decision **now**, because it changes what core-bu should
|
||||
replicate.
|
||||
|
||||
### 3.4 DNS authority is split
|
||||
|
||||
Confirmed by the old plan's own checklist and this project's history: `itpropartner.com` is on **SiteGround
|
||||
nameservers (manual panel, no API)**; `fleettracker360.com` and `voipsimplicity.com` are on **Cloudflare**;
|
||||
`iamgmb.com`, `intelsight.io`, `debtrecoveryexperts.com` need per-domain `dig NS` verification in Phase 0.
|
||||
Every cutover record must be changed in the correct panel or it is a silent no-op. Section 6 lists the records.
|
||||
|
||||
---
|
||||
|
||||
## 4. Migration phases
|
||||
|
||||
Each phase has a gate: **the next phase does not start until the gate's evidence exists.**
|
||||
|
||||
### Phase 0 - Inventory and recon (Core, read-only) - IN PROGRESS
|
||||
|
||||
Deliverable: `docs/infrastructure/core-service-inventory-2026-09-15.md` (+ `.json`) - every container, unit,
|
||||
port, volume, database, cron job, TLS cert and Caddy route on Core, with sizes and dependencies.
|
||||
|
||||
Gate: inventory lists every Caddy site block and its upstream, and explicitly resolves two conflicts that
|
||||
existing docs disagree on:
|
||||
1. **DocuSeal** is recorded on **Core** by the Aug 15 draft but on **App1** by `backup-plan.md` (4:00 AM job).
|
||||
2. **SearXNG** likewise. Only the live `docker ps` / Caddyfile settles it.
|
||||
|
||||
### Phase 1 - Provision app4 + core-bu, monitoring first - **COMPLETE (2026-09-15)**
|
||||
|
||||
See Section 1. Gate met: both boxes verified; backups restore-tested; both scraped by Prometheus; no customer
|
||||
app touched.
|
||||
|
||||
### Phase 2 - Access and naming (needs your input)
|
||||
|
||||
- Enroll both boxes in Tailscale (needs a reusable auth key or your approval of the login URL - Section 9, Q1).
|
||||
- Decide DNS names: `app4.itpropartner.com` and `core-bu.itpropartner.com` A/AAAA records, added in the correct
|
||||
panel (SiteGround for `itpropartner.com`). Internal access and monitoring already work **by IP**, so this is
|
||||
not blocking, but the docs and the recovery manual read better with names.
|
||||
- Install Caddy on app4 with `default_bind 159.195.205.80` (avoids the Tailscale :443 conflict).
|
||||
|
||||
Gate: `tailscale status` shows both nodes; name resolution works from Core.
|
||||
|
||||
### Phase 3 - Prove the pattern on low-risk apps
|
||||
|
||||
- Move **microbin** (`127.0.0.1:8260`) first: single container, one volume, no database.
|
||||
- Move **Uptime Kuma** second: it is the monitoring tool, so it must be moved carefully and its own downtime
|
||||
window announced.
|
||||
- For each: stop on Core, rsync the volume, start on app4, verify side-by-side with
|
||||
`curl --resolve <domain>:443:159.195.205.80`, then flip DNS, then soak 24 h.
|
||||
- This phase validates the runbook (per-service steps, verification and rollback) before any customer app moves.
|
||||
|
||||
Gate: microbin and Uptime Kuma both served from app4 with app4-issued TLS, verified externally, and their S3
|
||||
backups land from app4 - not from Core - with a restore test on at least one.
|
||||
|
||||
### Phase 4 - Data foundation + first real app
|
||||
|
||||
- Provision Postgres and Redis **fresh** on app4 (internal-only binds, least privilege, no public 5432/6379).
|
||||
- Move the **Ops Portal backend** (`:8090`), then **DocuSeal** (SQLite + attachments + its internal Redis),
|
||||
then **TimeTrex** (Postgres-backed).
|
||||
- Databases: `pg_dump -Fc` per database, restore on app4, then **compare row counts per major table**, not a
|
||||
spot check. SQLite: `sqlite3 .backup`, never `cp`.
|
||||
- **New backup jobs on app4 for every database it now hosts** (`*.db` is excluded from the essentials archive).
|
||||
|
||||
Gate: row counts match; `curl --resolve` responses match Core; app4-backup + restore test for each moved DB.
|
||||
|
||||
### Phase 5 - Customer sites and the voice stack
|
||||
|
||||
- rsync every static customer site root (`*.iamgmb.com`, `*.intelsight.io`, `*.fleettracker360.com`,
|
||||
`*.debtrecoveryexperts.com`) to app4; `caddy validate` the app4 config; pre-issue TLS.
|
||||
- Voice stack (`voice.*`, `voice-open.*`): enumerate Twilio webhooks and any external endpoints in Phase 0 and
|
||||
update them **before** the DNS flip, or calls break after cutover.
|
||||
|
||||
Gate: every moving domain answers from app4 with a valid cert; voice end-to-end call tested.
|
||||
|
||||
### Phase 6 - DNS cutover, soak, decommission on Core
|
||||
|
||||
- Lower TTL to 60-300 on every moving record **24 h before** the flip (correct panel per domain).
|
||||
- Flip one domain at a time, low traffic first, verifying each (`dig +short @1.1.1.1`, then `curl -sI`).
|
||||
- Keep Core's Caddy blocks as a 301 redirect to app4 during a 24-72 h soak; remove with targeted edits and the
|
||||
caddy-audit hook (never a whole-file rewrite).
|
||||
- Then: stop/remove the moved containers on Core, retain volumes + images **30 days** as rollback, decommission
|
||||
the customer schemas in Core's Postgres/Redis.
|
||||
|
||||
Gate: 72 h soak with no rollback; Core runs zero customer-facing apps; rollback path still intact.
|
||||
|
||||
### Phase 7 - core-bu standby, failover AND failback proven (parallel with 3-6)
|
||||
|
||||
- Install the standby package (sync + watchdog with a health-based decision branch, fence-before-takeover, and
|
||||
**automatic failback**, which the current app1-bu scripts do not have).
|
||||
- **Only one standby may be armed at a time.** Disarm app1-bu before arming core-bu, or a Core hiccup makes both
|
||||
answer as the same Telegram bot.
|
||||
- Prove it with a real, announced test: failover, then failback, then confirm the standby is dormant again.
|
||||
|
||||
Gate: failover and failback both demonstrated with evidence, and app1-bu verifiably disarmed.
|
||||
|
||||
### Phase 8 - Documentation and reference sweep
|
||||
|
||||
Deliverable: `docs/infrastructure/reference-update-matrix-2026-09-15.md` - every artifact that names a host.
|
||||
Includes at minimum: `key-inventory.md` (done), `backup-plan.md` (done), `CHANGELOG.md` (done),
|
||||
`app-inventory.csv`, `server-architecture-plan`, `server-dr-plans`, `dr-issue-log`, the recovery manual,
|
||||
`decommissioned-hosts.json` + `stale-reference-verify.py` (add `app1-bu` when retired), Prometheus config
|
||||
(live one, **and delete/refresh the dead `/opt/prometheus/prometheus.yml`**), Grafana dashboards, Uptime Kuma
|
||||
monitors, `health-master-watchdog.py`, Hermes cron `jobs.json` live-config fields, Hudu assets, the ops portal,
|
||||
client-facing runbooks, and any skill that hardcodes a host or IP.
|
||||
|
||||
Gate: `stale-reference-verify.py` and `doc-live-verify.py` both clean; every doc cites the new IPs.
|
||||
|
||||
---
|
||||
|
||||
## 5. Acceptance criteria (whole project)
|
||||
|
||||
1. Core hosts no customer-facing app; every moved domain answers from app4 with a valid TLS cert.
|
||||
2. Every moved service has: a data migration that was verified by counts/sizes, a health check, and a tested
|
||||
rollback.
|
||||
3. Every database on app4 has its own backup job with a **performed restore test** (a green cron entry is not
|
||||
evidence).
|
||||
4. `core-bu` failover **and** failback both demonstrated; exactly one standby armed at any time.
|
||||
5. Documentation matrix closed out: no live surface names a decommissioned host or a stale IP.
|
||||
6. `app1-bu` is either retired (with the provider-diversity decision recorded) or explicitly retained as the
|
||||
provider-diverse standby.
|
||||
|
||||
---
|
||||
|
||||
## 6. DNS and Caddy change checklist
|
||||
|
||||
- [ ] `dig NS` every moving domain; record the authoritative panel in Phase 0.
|
||||
- [ ] Pre-write all moving site blocks into app4's Caddyfile; `caddy validate`; pre-issue certs.
|
||||
- [ ] `default_bind 159.195.205.80` in app4's Caddy global block.
|
||||
- [ ] TTL 60-300 at least 24 h before each flip.
|
||||
- [ ] Flip per domain; verify `dig +short @1.1.1.1` and `curl -sI https://<domain>`.
|
||||
- [ ] Keep Core blocks as 301s for the soak window; then targeted removal + caddy-audit hook.
|
||||
- [ ] Update Http->Https and any `CNAME`/`www` records in the **same** panel as the A record.
|
||||
|
||||
---
|
||||
|
||||
## 7. Rollback
|
||||
|
||||
- Before each phase: snapshot DNS records, Core Caddyfile, Core `docker ps`/volume list.
|
||||
- Phases 3-5: stop on app4, flip DNS back to Core, restart the Core container. Core volumes are untouched.
|
||||
- Phase 6: with low TTL, the flip back propagates in minutes; Core blocks are retained during soak.
|
||||
- Data: Core volumes/images retained 30 days. After that, restore from app4's S3 backups (which is why the
|
||||
Phase 1/4 restore tests are mandatory).
|
||||
- `core-bu`: failback is part of the design, not an afterthought; the standby stands down on its own.
|
||||
|
||||
---
|
||||
|
||||
## 8. Follow-up findings raised by this work (not fixed here)
|
||||
|
||||
| # | Finding | Impact | Owner |
|
||||
| --- | --- | --- | --- |
|
||||
| 1 | `node_exporter` not running on app1, app2, app3, app1-bu | No host metrics for the fleet | This project (Phase 8) |
|
||||
| 2 | `/opt/prometheus/prometheus.yml` is a dead file containing decommissioned hosts (wphost02, 178.156.131.57) and was never loaded | Misleading; wasted trust | Phase 8 |
|
||||
| 3 | netcup SCP/CCP API auth returns HTTP 500 / 404 (worked in July) | Provisioning automation via API is dead | Separate |
|
||||
| 4 | `hermes-standby-sync.sh` has a ping-based failback flaw and no failback logic | Standby reliability | Phase 7 |
|
||||
| 5 | `hermes-snapshot.sh:34` does a live `VACUUM INTO` | Store churn | Separate |
|
||||
|
||||
---
|
||||
|
||||
## 9. Open decisions (needed from Germaine)
|
||||
|
||||
**Q1 - Tailscale:** add both boxes to the tailnet. Need either a reusable auth key, or approve the login URL
|
||||
from each box.
|
||||
|
||||
**Q2 - Nuremberg vs Manassas for app4:** accept Nuremberg (design app4 self-contained) or re-order app4 in
|
||||
Manassas and repurpose the Nuremberg box? Section 3.1.
|
||||
|
||||
**Q3 - Vaultwarden:** the CLI is locked. Unlock it (or tell me when) and I will file the two new server items.
|
||||
|
||||
**Q4 - app1-bu:** retire it, or keep it as the provider-diverse standby? Section 3.2 - this is the only
|
||||
remaining non-netcup box.
|
||||
|
||||
**Q5 - app4 standby scope:** accept S3-restore RTO, or should core-bu carry a secondary copy? Section 3.3.
|
||||
|
||||
---
|
||||
|
||||
## Appendix A - Verified evidence log (2026-09-15)
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| RTT Core -> app4 / core-bu | 100.5 ms / 100.5 ms |
|
||||
| RTT Core -> app2 (Manassas) / app1-bu (Ashburn) | 0.5 ms / 1.6 ms |
|
||||
| `authorized_keys` (root + ippadmin, both boxes) | present, sha256 `102c80e5...`, 1 line each |
|
||||
| Key login from Core | root OK, ippadmin `sudo -n` -> root, both boxes |
|
||||
| Password auth / non-allowlisted user | refused on both |
|
||||
| `ufw status` | active, 22/80/443 + 9100 from Core/tailnet |
|
||||
| Docker / Compose | 29.8.0 / v5.5.1 both |
|
||||
| `node_exporter` | 200 from Core; refused from app2 (third host) |
|
||||
| Prometheus `up{job="node_exporter"}` | core, app4, core-bu = 1 |
|
||||
| First backup run | upload + download/extract verify OK on both |
|
||||
| Credentials file | `/root/.hermes/references/new-servers-2026-09-15.md`, mode 600 |
|
||||
@@ -0,0 +1,148 @@
|
||||
# Reference Update Matrix — app4 / core-bu Go-Live and app1-bu Retirement
|
||||
|
||||
**Date:** 2026-09-15
|
||||
**Author:** Sho'Nuff (read-only audit subagent, ran on Core)
|
||||
**Scope:** Every document, script, config, monitoring check, and record found on Core
|
||||
that names a host/service and must change when app4 (159.195.205.80) and the new
|
||||
netcup core-bu (159.195.204.203) go live, and when the Hetzner standby app1-bu
|
||||
(5.161.225.131) is eventually retired.
|
||||
**Method:** Live grep/read of `/root/projects/itpp-infrastructure`, `/root/.hermes/scripts`,
|
||||
`/root/.hermes/references`, `/root/.hermes/cron/jobs.json`, `/root/.hermes/config.yaml`,
|
||||
`/etc/caddy/Caddyfile`, `/etc/cron.d`, root crontab, systemd units, `/opt/ops-portal`,
|
||||
and read-only Gitea API calls to `git.itpropartner.com`. No file was edited. No service
|
||||
was restarted. Secrets are masked below where shown.
|
||||
**Read-only constraint honored:** all evidence below is quoted from files actually read
|
||||
on 2026-09-15. Line numbers refer to the file state at read time.
|
||||
|
||||
**IMPORTANT correction found during this audit:** `/root/.hermes/scripts/doc-live-verify.py`
|
||||
currently hardcodes `app3.itpropartner.com` specs as "RS 4000 (8 vCPU, 16 GB RAM, 320 GB SSD)"
|
||||
— this is stale even before app4/core-bu (app3 is actually 12 vCPU/32GB/1TB per the live
|
||||
`server-architecture-plan` skill and `architecture.md`). Flagged here since it lives in the
|
||||
same `SERVER_INVENTORY` structure that needs the P1/P2 edits below.
|
||||
|
||||
---
|
||||
|
||||
## How to read this matrix
|
||||
|
||||
- **Phase (P1/P2/P3)** — P1: edit at provisioning time (DNS/rDNS/inventory/monitoring/backup
|
||||
targets/decommissioned-hosts/SSH keys), before any service moves. P2: edit at the moment
|
||||
services actually cut over to app4 / core-bu takes over Core's standby role. P3: edit only
|
||||
once app1-bu is deleted from the Hetzner account.
|
||||
- **Current** is the exact text/value read from the file (quoted or line-numbered).
|
||||
- **Required new value** is what the artifact must say once the transition in that phase lands.
|
||||
- Rows are grouped by artifact class: docs, scripts, cron/systemd, Caddy, Gitea/repo, ops-portal,
|
||||
monitoring/decommission tooling.
|
||||
|
||||
---
|
||||
|
||||
## P1 — Must change at provisioning time (DNS/rDNS, inventory, monitoring watch lists, backup targets, decommissioned-hosts, SSH key distribution)
|
||||
|
||||
| # | Artifact | Current text/value (evidence) | Required new text/value | Why it matters |
|
||||
|---|---|---|---|---|
|
||||
| 1 | `docs/infrastructure/key-inventory.md` line 14 | `Universal server admin key... All servers (Core, app1, app2, app3, app1-bu, home router).` | Add app4 (159.195.205.80) and core-bu (159.195.204.203) to the deployed-to list for `itpp-infra` the moment the SSH key is injected at provisioning | Key inventory is the audit trail for "who can SSH where" — a missing entry means the new boxes are invisible to a credential audit |
|
||||
| 2 | `docs/infrastructure/key-inventory.md` §2 table (Server Root Passwords) | Rows for Core/app1/app2/app3/app1-bu only | Add app4 and core-bu rows: IP, netcup, access method, Vaultwarden pointer | Same table used for DR "how do I get in" — missing rows = no documented root access path |
|
||||
| 3 | `/root/.hermes/references/decommissioned-hosts.json` | No entries for app4/core-bu (they don't exist yet); app1-bu is still listed as live nowhere in this file (correct — it's not decommissioned) | No P1 change needed to this file for app4/core-bu (they are new, not decommissioned). Flagged here only to confirm scope: this file is a P3 artifact for app1-bu, not P1 | Prevents accidentally graveyard-listing a box that is being born, not retired |
|
||||
| 4 | `/root/.hermes/scripts/stale-reference-verify.py` (docstring + `SCAN_ROOTS`) | Scans `/etc/systemd/system/`, `/root/.hermes/scripts/`, `/etc/cron.d/`, `/var/spool/cron/crontabs/` for graveyard hits | No functional change required at P1 (app4/core-bu aren't graveyarded). Verify after P1 that no stray reference to the old "core-bu = Hetzner" naming (see row 20) creates a false-clean scan | The script's correctness depends on `decommissioned-hosts.json` staying accurate; a naming collision (two things called "core-bu") could mask a real stale reference later |
|
||||
| 5 | `/root/.hermes/scripts/doc-live-verify.py` lines 50-57 (`SERVER_INVENTORY`) | Dict has entries for `core.itpropartner.com`, `app1`, `app2`, `app3.itpropartner.com`, `app1-bu.itpropartner.com` only — no app4/core-bu entries exist | Add `app4.itpropartner.com` (159.195.205.80, netcup, RS 4000 G12, "customer-facing services") and a core-bu entry (159.195.204.203, netcup, RS 2000 G12 twin, "warm standby for Core") to `SERVER_INVENTORY` and `domain_map` (lines ~90-108) | Without this, the scanner has no baseline for the new IPs and can't tell a correct new reference from a typo'd one |
|
||||
| 6 | `/root/.hermes/scripts/health-master-watchdog.py` `REMOTE_SERVERS` (lines ~91-97) | `[("app1", "152.53.36.131"), ("app2", "152.53.39.202"), ("app3", "152.53.241.111"), ("app1-bu", "5.161.225.131"), ("anita-mnz", "159.195.16.30")]` | Add `("app4", "159.195.205.80")` and `("core-bu", "159.195.204.203")` at provisioning so ping/reachability + backup-freshness checks cover the new boxes from day one, per the precedent set for anita-mnz on 2026-09-11 | This exact list is what the skill's own pitfall log calls out — a box not added here alerts never, a box removed from the wrong list (old app1-bu) alerts forever after retirement |
|
||||
| 7 | `/root/.hermes/scripts/health-master-watchdog.py` `REMOTE_DOCKER_CONTAINERS` (lines ~108-124) | Keys for `app1`, `app2` only; comment: `# app1-bu (Hetzner standby) has no Docker; app3 is CloudPanel-only.` | Add an `app4` key with the moved containers' names once DocuSeal/TimeTrex/microbin/Uptime Kuma/Ops-Portal-backend containers exist there (P2 timing for the container list itself, but the dict key should exist and be empty-safe at P1 so the code path is proven) | Container health checks are useless if the box hosting them was never added to the watch structure |
|
||||
| 8 | `/root/.hermes/scripts/vps-threshold-check.sh` lines 23-27 | `SERVERS=("core|152.53.192.33|1" "app1|152.53.36.131|0" "app2|152.53.39.202|0" "app3|152.53.241.111|0" "app1-bu|5.161.225.131|0")` | Append `"app4|159.195.205.80|0"` and `"core-bu|159.195.204.203|0"` | Disk/RAM/CPU threshold alerting has zero coverage on unlisted hosts; this is a flat array, trivial to miss |
|
||||
| 9 | `/root/.hermes/scripts/security-compliance-check.sh` line 6 | `SERVERS="152.53.192.33:core 152.53.36.131:app1 152.53.39.202:app2 152.53.241.111:app3 5.161.225.131:app1-bu"` | Append `159.195.205.80:app4 159.195.204.203:core-bu` | Nightly compliance sweep (patches, fail2ban, failed systemd units) silently skips any host not in this string |
|
||||
| 10 | `/root/.hermes/scripts/backup-health-monitor.sh` line 774 | `for server_entry in "core:152.53.192.33" "app1:152.53.36.131" "app2:152.53.39.202" "app3:152.53.241.111"; do` (Phase 3: System Crontabs SSH check) — note app1-bu is already absent from this specific loop | Add `"app4:159.195.205.80"` and `"core-bu:159.195.204.203"` once those boxes have crontabs to audit | This SSH-crontab-audit loop already under-covers (app1-bu missing); don't propagate the gap to the two new boxes |
|
||||
| 11 | `/root/.hermes/scripts/backup-failure-check.sh` line 103 | `for server in "app1:152.53.36.131" "app2:152.53.39.202" "app3:152.53.241.111"; do` | Add app4 (and core-bu if it carries backup jobs of its own beyond the standby sync) | Same class of gap as row 10 — backup failure detection has a fixed server list |
|
||||
| 12 | `/root/.hermes/scripts/api-health-check.py` lines 62-83 | Hardcoded `(name, mode, ip, port, path)` tuples for app1/app2/app3 endpoints only | Add app4 tuples for each service once it lands there (DocuSeal :3000, TimeTrex :8085, microbin :8260, Uptime Kuma :3001, Ops Portal :8090) — this is P2-timed for the actual endpoints, but the provisioning step should reserve the pattern now | API health checks won't exist for anything moved to app4 until this file is edited |
|
||||
| 13 | `/root/.hermes/references/reserved-ports.json` | `_comment`: "Reserved service ports on core (152.53.192.33)." Lists ports 8899, 8500, 8090, 8099, 8200, 8910, 8888, 3002, 3000 all scoped to Core | When Ops Portal (:8090), DocuSeal (:3000) etc. move off Core to app4, this file's port-identity-guard becomes partially stale for Core and needs an equivalent file (or scope note) for app4 | `port-identity-guard.py` (referenced by this file) exists to catch port-squatting; if the service moves and the guard still watches Core's now-empty port, a real squat on app4 goes undetected |
|
||||
| 14 | `/root/.hermes/cron/jobs.json` — `hetzner-weekly-snapshots` job (id `faa6b8760e38`, script `snapshot-hetzner.py`) | Runs weekly Mon 5:00 AM, `no_agent: true`, targets Hetzner API broadly | No P1 edit required (this job is Hetzner-API-driven, not IP-hardcoded in jobs.json itself — confirm the script's own server-ID list separately, see row 27) | Flagged for completeness; the actual IP/server-ID list lives inside `snapshot-hetzner.py`, not jobs.json |
|
||||
| 15 | `docs/infrastructure/app4-migration-plan.md` line 70 | `**Recommendation: netcup RS 4000 G12 (12 vCPU / 32 GB DDR5 ECC / 1 TB NVMe), ~$44/mo, Manassas VA.**` | Correct to Nuremberg (Manassas RS was sold out per `standby-host-replacement-2026-09-14.md` context and the server-architecture-plan skill's 2026-09-15 update): app4 is ordered at Nuremberg, IP 159.195.205.80, rDNS `v2202609377162521278.quicksrv.de` | The plan's own location assumption is now wrong; anyone following it to provision would target the wrong DC and the wrong ordering flow |
|
||||
| 16 | `docs/infrastructure/app4-migration-plan.md` (no core-bu content at all — plan only discusses app4) | Plan silent on the new core-bu box entirely; describes only Core→app4 | Add a section (or a companion doc) covering core-bu's provisioning as Core's new-generation standby, superseding the Hetzner-only DR references throughout `hermes-standby-deployment` skill and `server-dr-plans.md` | The single biggest gap in the existing plan set — core-bu's arrival redefines the entire DR architecture (a second, same-provider-family warm standby) and nothing documents it yet |
|
||||
| 17 | `README.md` lines 81-82, 91, 94 | `**Hostname:** core-bu` / `**IP:** 5.161.225.131` ... `**Hetzner Cloud (current):** As of 2026-08-28, the Hetzner Cloud API returns exactly **one** server — **app1-bu / core-bu** (5.161.225.131...)` | This is the exact naming collision flagged by the server-architecture-plan skill's 2026-09-14 correction: "core-bu is a name for the standby role, not a reference to the Hetzner host." Update README to stop calling the Hetzner box `core-bu` and use `app1-bu` only, reserving `core-bu` for the new netcup box (159.195.204.203) | This is a live, already-existing naming defect (not just a future one) — three files (`README.md`, `docs/infrastructure/key-inventory.md` §2, `network-diagram.md`, `master-apps-services.md`) currently call the Hetzner box "core-bu", which will directly collide with the new netcup box once it's named core-bu. Fix at P1, before the second "core-bu" exists, or every future reference is ambiguous |
|
||||
| 18 | `/root/.hermes/references/ip-dns-changes.md` line 17 and `server-inventory.md` line 14 and `master-apps-services.md` line 19 and `network-diagram.md` line 54 | All four call the Hetzner box (5.161.225.131) "core-bu" | Same fix as row 17 — rename all four to `app1-bu` for the Hetzner box, and add net-new rows for the netcup core-bu (159.195.204.203) once ordered | Four separate documents currently reinforce the same naming collision; all four must be corrected together or drift resumes immediately |
|
||||
| 19 | `docs/monitoring/uptime-kuma-monitoring-plan.md` §2 (Core/App1/App2/App3/wphost02 Caddy sections) and the A1-A4 Server Health monitor rows (lines 392-395) | Monitors "Core", "app1", "app2", "app3" by IP only — no app1-bu, app4, or core-bu row exists in the Uptime Kuma monitor list | Add A5 (app4, 159.195.205.80) and A6 (core-bu, 159.195.204.203) TCP/22 health monitors at provisioning; keep app1-bu's own monitor (if any exists in live Kuma — not confirmed by this audit, Kuma admin UI not inspected) until P3 | Public-facing Uptime Kuma is the customer-visible signal; a new customer-facing host (app4) with zero monitors defeats the purpose of the migration |
|
||||
| 20 | `/root/.hermes/references/server-inventory.md` (whole file, dated July 10 2026) | Table has Core/app1/app2/app3/core-bu(Hetzner)/old-ai/docker-box/wphost02/UNMS/UniFi — badly stale even before this change (old-ai, docker box, wphost02 are all long decommissioned per `decommissioned-hosts.json`) | This file needs a full refresh regardless of app4/core-bu; at minimum add app4 and core-bu rows and correct the core-bu/app1-bu naming per row 17 | Already the most stale document found in this sweep — nothing here should be trusted as current without cross-checking `server-architecture-plan` skill or live probes |
|
||||
| 21 | `/opt/ops-portal/server.py` lines 370-374 and 649-653 and 1878-1881 | Three separate hardcoded server lists: `{"name": "Core", "ip": "152.53.192.33", ...}` etc., each listing Core/app1/app2/app3/app1-bu only, no app4/core-bu | Add app4 and core-bu entries to all three lists (they are NOT DRY — each must be edited independently, a known Caddyfile-style risk) | Ops Portal is the customer/staff-facing dashboard; three independently-maintained lists mean three places to miss, and the portal itself is moving to app4 (see row 24) which makes this doubly urgent |
|
||||
| 22 | `/opt/ops-portal/server.py` line 706 and 1912 (`REMOTE_DOCKER_CONTAINERS`-equivalent dict, `app1-bu` key) | `"app1-bu": [{"name": "caddy", "type": "systemd"}, {"name": "docker", "type": "systemd"}]` (two separate near-identical dicts at lines ~706 and ~1912) | Add `"app4"` key with its own service list once services move; keep `"app1-bu"` key until P3 | Same all-lists-must-be-edited risk as row 21 |
|
||||
| 23 | `/opt/ops-portal/static/dependency-diagram.html` line 198 | `<text ...>app1-bu.itpropartner.com</text>` (SVG dependency diagram) | Add app4.itpropartner.com and core-bu nodes to the diagram at provisioning | Static SVG diagrams don't regenerate themselves — this is hand-edited HTML that silently goes stale |
|
||||
|
||||
---
|
||||
|
||||
## P2 — Must change at service cutover (app4 takes over customer-facing services, core-bu becomes Core's warm standby)
|
||||
|
||||
| # | Artifact | Current text/value (evidence) | Required new text/value | Why it matters |
|
||||
|---|---|---|---|---|
|
||||
| 24 | `/etc/caddy/Caddyfile` (live, read 2026-09-15) — every customer-facing site block | `default_bind 152.53.192.33` (global block) plus individual `reverse_proxy 127.0.0.1:PORT` blocks for `sign.itpropartner.com`, `ops.itpropartner.com`, `my.itpropartner.com`, `status.itpropartner.com`, `uptimekuma.itpropartner.com`, `voice.itpropartner.com`, `voice-open.itpropartner.com`, `auth.itpropartner.com`, and the `my.itpropartner.com` block's `reverse_proxy http://152.53.241.111:8090` calls to app3's backup-restore API | Per `app4-migration-plan.md` §2.2/§4 Phase 4: after data migration and DNS cutover, remove these site blocks from Core's Caddyfile (targeted edits + caddy-audit hook, never full rewrite) and stand up an equivalent Caddyfile on app4 with `default_bind 159.195.205.80`. Core's Caddyfile keeps only `core.itpropartner.com` and Core-local routes | This is the single largest, most consequential P2 change — every customer domain's TLS termination and backend routing moves. Getting the Caddyfile edit wrong breaks live customer traffic |
|
||||
| 25 | `dns-records.md` §PRODUCTION table (lines 21-43) | `ops.itpropartner.com → 152.53.192.33 → Core`, `my.itpropartner.com → 152.53.192.33 → Core`, `sign.itpropartner.com → 152.53.192.33 → Core`, `uptimekuma.itpropartner.com → 152.53.192.33 → Core`, `status.itpropartner.com → 152.53.192.33 → Core` | Flip A records for each moved domain to 159.195.205.80 (app4) per the phased DNS cutover in `app4-migration-plan.md` §4/§6 (lower TTL first, soak, then cut). SiteGround manual panel for itpropartner.com — no API | DNS is the actual cutover mechanism; the doc must track reality or the next person "fixing" a stale doc could break live routing |
|
||||
| 26 | `backup-plan.md` §Core backup inventory (lines 12-19) and §Schedules Summary | `core-services-backup.sh` targets Grafana/Uptime-Kuma/Docker-volumes/Prometheus all on Core (152.53.192.33); Ops Portal backend backup is implied under `root-essentials-backup.sh`/`core-services-backup.sh` on Core | Per `app4-migration-plan.md` §7: repoint backup scripts and cron for every moved app (DocuSeal, TimeTrex, microbin, Uptime Kuma, Ops Portal backend, Postgres, Redis) from Core to app4. Add an "App4" backup section to `backup-plan.md` mirroring the App1/App2/App3 sections, with its own S3 prefix (`s3://hermes-vps-backups/app4/...`) | If the app4 migration plan's own §7 instruction ("confirm app4 cron owns them before removing Core entries") isn't followed, moved services lose backup coverage silently — exactly the anita-mnz `*.db`-exclusion failure mode from 2026-09-11 |
|
||||
| 27 | `/root/.hermes/scripts/uptime-kuma` backup script (`core-services-backup.sh`, per `backup-plan.md` row) and the Uptime-Kuma container itself | Uptime Kuma runs on Core per `server-architecture-plan` skill ("Uptime-Kuma (Docker, port 3001) ... on Core") and is backed up via `core-services-backup.sh` | Per `app4-migration-plan.md` §2.2/Phase 2 ("Move microbin and Uptime Kuma first"), Uptime Kuma is the first DNS flip of the whole project — update `core-services-backup.sh` to stop covering it and add it to a new `app4-backup.sh`, and update `uptimekuma.itpropartner.com` Caddy block + DNS together | Uptime Kuma is the monitoring tool watching everything else — if its own migration breaks its backup or its DNS without coordinated verification, you lose visibility into the exact moment you need it most |
|
||||
| 28 | `hermes-standby-deployment` skill and `/root/.hermes/scripts/hermes-standby-watchdog.sh` / `hermes-standby-sync.sh` / `hermes-standby-restore.sh` | `LIVE_HOST="152.53.192.33"` hardcoded in all three scripts (watchdog line 8, sync line 11, restore line 17); watchdog alert text: `Standby: app1-bu.itpropartner.com (5.161.225.131)` | When core-bu (159.195.204.203) becomes the standby role for Core (replacing/supplementing app1-bu), the design question is: does core-bu run these SAME scripts pointed at the same `LIVE_HOST=152.53.192.33`, or does app1-bu get retired and core-bu take over the whole watchdog/sync/restore stack? Either way, `LIVE_HOST` stays 152.53.192.33 (Core's IP doesn't change) but `Standby:` alert text and the deployment target change to core-bu's IP/hostname. This decision must be made explicit in a doc before P2 lands, not inferred from script diffs | Two standbys with two independently-running watchdog/sync/fencing setups is a split-brain risk if both start Hermes on Core-down; the architecture decision (core-bu replaces app1-bu vs. runs alongside it) is currently undocumented anywhere in the swept files |
|
||||
| 29 | `hermes-standby-deployment` skill "2026-09-15 correction" section — `fence_core()` | `fence_core() SSHes to the live box (PROBE_HOST/PROBE_SSH_KEY) and runs systemctl --user stop hermes-gateway` | If core-bu becomes the active standby, `PROBE_HOST` must point at Core (unchanged, 152.53.192.33) but the fencing SSH key and the box RUNNING fence_core() moves to core-bu. Update the skill's Testing/Numbers-that-matter section to reflect which box is standby-of-record | The recently-fixed failover state machine (2026-09-15 correction) was hard-won; redeploying it on a new box without re-verifying TAKEOVER_AFTER/DRYRUN/fence behavior risks reintroducing the exact bug just fixed |
|
||||
| 30 | `docs/infrastructure/standby-host-replacement-2026-09-14.md` (whole document) | Entire document's recommendation is "move app1-bu to Hetzner fsn1/nbg1 for cost + regional diversity" — written one day before the decision to build a netcup core-bu instead | This document's core recommendation is superseded by the actual decision (order a second netcup box, not relocate the Hetzner one). Add a superseding note at the top: "Superseded 2026-09-14/15 — owner ordered netcup core-bu instead of relocating app1-bu; see app4-migration-plan.md and this matrix" | Without an explicit supersession note, a future reader (or scanner) will treat this doc's Hetzner-relocation recommendation as still-live guidance, wasting effort or causing a wrong action |
|
||||
| 31 | `docs/infrastructure/key-inventory.md` §2 (Server Root Passwords) `app1-bu` row: `"Warm standby (core-bu)"` | Parenthetical calls the Hetzner box "core-bu" — same naming collision as row 17, but specifically in the credentials table | Fix at P1 per row 17, but re-verify at P2 that the role note is accurate once core-bu (netcup) is actually live and taking over standby duties — app1-bu's row should then say "legacy warm standby, superseded by core-bu, retirement candidate" | Credential documentation errors compound: an admin reading this table during an incident needs to know which box is the ACTIVE standby right now |
|
||||
| 32 | Gitea `itpp-infrastructure` repo (git.itpropartner.com, confirmed reachable, HTTP 200, API responsive) — same file set as this local clone | Remote origin is `https://ippadmin:***@git.itpropartner.com/ippadmin/itpp-infrastructure.git`; per `git remote -v` this is the only remote, and it's the push target for every doc listed above | No new remote needed; simply confirm every P1/P2 edit above gets committed and pushed to this repo (the audit found this repo push-current as of 2026-09-12 commit `199baad`). The `disaster-recovery` Gitea repo (also confirmed reachable) should receive a parallel supersession note if it duplicates any of the standby/DR content | Local-only edits that never reach Gitea are invisible to anyone cloning fresh, and to `hermes-migration`/`server-recovery-bundle` workflows that pull from Gitea, not from this specific Core filesystem |
|
||||
| 33 | Ops Portal `TRACCAR_HOST` and similar single hardcoded backend-IP constants (`server.py` line ~1348) | `TRACCAR_HOST = "152.53.39.202"` (app2, unaffected by this migration, listed here only because it lives in the same file as the app1-bu/app4 dicts flagged in rows 21-22) | No change required — app2 is out of scope. Listed for completeness so a future editor doing a "replace all IPs" pass on `server.py` doesn't accidentally touch this unrelated constant | Prevents scope creep / accidental breakage during the app4 edit pass |
|
||||
|
||||
---
|
||||
|
||||
## P3 — Must change at app1-bu retirement (every place that would break or lie once the Hetzner box is gone)
|
||||
|
||||
| # | Artifact | Current text/value (evidence) | Required new text/value | Why it matters |
|
||||
|---|---|---|---|---|
|
||||
| 34 | `/root/.hermes/references/decommissioned-hosts.json` | No entry for app1-bu / 5.161.225.131 (it is currently live, correctly absent) | Add an entry: `{"hostname": "app1-bu", "ip": "5.161.225.131", "decommissioned": "<retirement date>", "services": ["hermes-gateway", "caddy", "docker"], "note": "Hetzner CPX21 warm standby, superseded by netcup core-bu (159.195.204.203). Retired <reason>. Provider order cancellation is a USER action per server-decommissioning skill."}` | This is THE canonical source `stale-reference-verify.py` and `doc-live-verify.py` both read — everything else in this phase depends on this entry existing first |
|
||||
| 35 | `/root/.hermes/scripts/health-master-watchdog.py` `REMOTE_SERVERS` line ~94 | `("app1-bu", "5.161.225.131")` | Remove this tuple entirely (or comment it out with a decommission note, per the anita-mnz precedent at line ~92 of the same file: `# hermes-gateway-anita.service moved to the dedicated box... It is checked remotely via REMOTE_USER_UNITS below, NOT locally here.`) | A dead host left in `REMOTE_SERVERS` alerts "unreachable" every single watchdog cycle forever — this exact failure mode is documented in the `server-decommissioning` skill as the recurring pitfall |
|
||||
| 36 | `/root/.hermes/scripts/vps-threshold-check.sh` line 27 | `"app1-bu|5.161.225.131|0"` | Remove | Threshold check will error/false-alert on an unreachable host indefinitely |
|
||||
| 37 | `/root/.hermes/scripts/security-compliance-check.sh` line 6 | `...5.161.225.131:app1-bu"` (trailing entry in the SERVERS string) | Remove `5.161.225.131:app1-bu` from the string | Nightly compliance check SSHes to every listed host; a deleted server means permanent SSH-failure noise |
|
||||
| 38 | `hermes-standby-deployment` skill (entire skill file) | Skill is written entirely around app1-bu/Hetzner: "Hetzner CPX21 recommended", `enable_rescue` API calls, Hetzner-specific rescue-mode key injection, `SERVER_ID` referencing Hetzner's API, cost tables in EUR/Hetzner pricing | This skill needs either (a) a full rewrite for netcup-based core-bu deployment (no Hetzner rescue mode, no Hetzner API, different provisioning flow per `server-provisioning-standard`), or (b) an explicit "ARCHIVED — app1-bu retired, see <new skill> for core-bu" banner if a new skill is authored separately | The entire deployment runbook (registering SSH keys via Hetzner API, rescue-mode key injection, CPX11→CPX21 upgrade math) is Hetzner-specific and becomes 100% inapplicable once core-bu (netcup, no rescue mode, standard provisioning) is the standby. Leaving this skill as the "how to deploy a standby" reference after retirement would actively mislead the next person who has to redeploy or troubleshoot |
|
||||
| 39 | `/root/.hermes/references/hermes-dr-plan-v2.md` (entire document) | Describes app1-bu exclusively: `Core (netcup RS 2000 — 152.53.192.33) / app1-bu (Hetzner CPX21 — 5.161.114.8)` (line 12, itself already stale — 5.161.114.8 was replaced by 5.161.225.131 back on 2026-07-24), SSH key `itpp-infra-v2` deployed "to app1-bu via rescue", Hetzner-specific fencing script, CPX11→CPX21 upgrade tables | Full rewrite required: replace every app1-bu/Hetzner/CPX21/rescue-mode reference with core-bu/netcup/RS-2000-twin/standard-provisioning equivalents, or mark the document ARCHIVED and write a v3 | This is the master DR plan referenced by the `hermes-standby-deployment` skill's own "Related Documents" table as "the target architecture" — if it still describes a deleted box as the target, DR execution during a real incident follows a runbook for infrastructure that no longer exists |
|
||||
| 40 | `/root/.hermes/references/3-Per-Server-Runbooks.md` lines 16, 24, 35, 50-51, 54, 74 | `Hermes orchestration hub — highest priority, protected by warm standby (app1-bu)`; `app1-bu (standby), S3 backup bucket...`; `If app1-bu is healthy, use failover...`; `Both Core and app1-bu active simultaneously — mitigated by...`; `app1-bu itself is restored like any Standby-tier host if it fails: provision replacement CPX11...` | Replace every app1-bu reference with core-bu across all 6+ occurrences; the CPX11 replacement-provisioning instruction (line 74) becomes netcup RS-line provisioning instead | Per-server runbooks are the document someone opens DURING an incident — six stale references to a deleted host in the runbook for CORE's own failover path is a P0-during-an-incident risk |
|
||||
| 41 | `/root/.hermes/references/itpp-recovery-manual.md` (8 occurrences across lines 17, 48, 52, 72, 272-326, 676-679, 1031-1037, 1185-1327) | Extensive: table of contents anchor `#4-app1-bu-standby-hetzner--51611148` (itself referencing the OLD retired IP 5.161.114.8 in the anchor text — a pre-existing stale anchor found during this audit), full §4 "app1-bu Standby (Hetzner — 5.161.225.131)" section with SSH commands hardcoded to `root@5.161.225.131`, StrongSwan/L2TP fallback note "runs on both Core and app1-bu" | Full section rewrite: new §4 "core-bu Standby (netcup — 159.195.204.203)" with corrected SSH commands, updated StrongSwan/L2TP note, and the TOC anchor fixed (it currently embeds a wrong IP even for app1-bu) | This is the master recovery manual — 39,541 bytes, 1351 lines, the document row 4 of the key-inventory's "Recovery Priority" list implicitly assumes is accurate. Every `ssh root@5.161.225.131` command in it will fail post-retirement, and following it during a real outage wastes the exact minutes DR is supposed to save |
|
||||
| 42 | `/root/.hermes/references/restore-runbooks.md` line 5 | `**Rollback plan:** Revert DNS to standby core-bu (5.161.225.131). Wipe server and restart runbook.` | This line ALREADY uses "core-bu" for the Hetzner IP (the naming collision from row 17, found here too) — at retirement, either this line's target no longer exists (if app1-bu/old-core-bu is deleted) or it needs redirecting to the NEW core-bu at 159.195.204.203 | A rollback plan pointing at a deleted IP is worse than no rollback plan — it will be trusted and fail silently mid-incident |
|
||||
| 43 | `/root/.hermes/references/restore-plan-2026-07-11.md` (8+ occurrences, e.g. lines 11, 38, 151-210, 332-333) | Entire "Recovery Path B: Failover to Warm Standby (app1-bu)" section with live `ssh -i /root/.ssh/itpp-infra root@5.161.225.131` commands | This is a dated incident-response document (July 11, 2026) describing a past incident — per the `server-decommissioning` skill's "Stale-IP cleanup in docs" guidance, dated historical incident reports are typically left as-is with an inline decommission annotation rather than rewritten, since they document what WAS done, not current procedure. Recommend annotating the top of the doc: "Historical — app1-bu (5.161.225.131) referenced below was retired <date>; current standby is core-bu (159.195.204.203)" rather than rewriting the incident narrative | Rewriting historical incident reports to reflect infrastructure that didn't exist at the time falsifies the record; annotation preserves both truth and utility |
|
||||
| 44 | `docs/infrastructure/key-inventory.md` lines 24 (SSH deployment scope), 42 (root password table), 221 (Unknown/Not Found §13 duplicate app1-bu row) | `All servers (Core, app1, app2, app3, app1-bu, home router)` / `app1-bu \| 5.161.225.131 \| Hetzner CPX21 \| itpp-infra SSH key \| Warm standby (core-bu)` / duplicate row 221: `**app1-bu** \| 5.161.225.131 \| Hetzner CPX21 — accessed via itpp-infra SSH key only.` (note: this row is misplaced inside §13 "Unknown/Not Found", itself a pre-existing doc defect) | Remove app1-bu from the SSH deployment scope line; remove or annotate-as-historical both the §2 row and the misplaced §13 duplicate row | Three separate mentions of app1-bu in one document, one of them filed under the wrong section header — all three need the retirement edit or two of three will be missed |
|
||||
| 45 | `docs/infrastructure/key-inventory.md` §11 Tailscale table, `app1-bu` row | `app1-bu \| 100.112.23.21 \| Linux \| ⚠️ Offline (7d)` — confirmed still true live via `tailscale status` on 2026-09-15 (shows `100.112.23.21 app1-bu ... offline, last seen 61d ago`, plus a SECOND stale node `100.95.212.28 app1-bu-1` currently `idle`) | At retirement: remove the `app1-bu` (100.112.23.21) row entirely; also remove/rename the `app1-bu-1` node once confirmed it's the same retired box (Tailscale auto-renamed it per the `hermes-standby-deployment` skill's Aug 7 pitfall note: "the node appears as app1-bu-1 (auto-renamed because the old app1-bu node was parked offline for 22+ days)") | Live evidence found DURING this audit: there are currently TWO Tailscale nodes for the same physical box (`app1-bu` offline 61 days, `app1-bu-1` idle/active). This is already a live inconsistency, not just a future one — flag now, clean up at retirement |
|
||||
| 46 | `README.md` (multiple: lines 81-94, 176-177, 253, 277) | `**Hostname:** core-bu` / `**IP:** 5.161.225.131` / `~~wphost02-backup~~... **REMOVED — wphost02 decommissioned**` (shows the precedent pattern to follow) / `warm-standby-sync \| Every 10 min \| core-bu ← S3 \| DR readiness` / `**Provider diversity:** core-bu stays at Hetzner specifically so a netcup outage can't kill both Core and standby simultaneously` | Follow the exact strikethrough-and-note pattern already used for wphost02-backup (line 176) for every app1-bu/old-core-bu line: strike it, note "REMOVED — app1-bu decommissioned <date>, superseded by core-bu (netcup, 159.195.204.203)". The "provider diversity" claim (line 277) becomes FALSE once core-bu is also netcup — this line must be corrected to state the diversity argument no longer applies (or reframed around Nuremberg vs Manassas regional diversity within netcup, which is weaker than true provider diversity) | Row 277 is a substantive factual claim ("provider diversity... netcup outage can't kill both") that becomes false, not just outdated, once the standby is also netcup. This is the highest-risk single line in the whole sweep — see Top 5 below |
|
||||
| 47 | `docs/infrastructure/app4-migration-plan.md` §7 line 205 | `Decide standby scope: app1-bu is a warm standby for Core, not for customer apps. app4 relies on S3 backups unless a customer-app standby is separately approved.` | Update to reference core-bu instead of app1-bu once it's the standby of record, and resolve the open question (does core-bu ALSO need to be app4's standby, or does app4 remain S3-only for DR) | This is an explicitly flagged open decision in the plan itself — retirement of app1-bu is the forcing function to finally resolve it |
|
||||
| 48 | `docs/architecture.md` line 17 | `| **app1-bu** \| 5.161.225.131 \| CPX21 (3 vCPU, 4 GB RAM, 80 GB) \| Hetzner \| Warm standby — provider diversity. Auto-restores from S3. |` | Replace the row with core-bu's specs (netcup RS 2000 G12 twin, 159.195.204.203) and drop "provider diversity" from the rationale text (see row 46) | Same false-claim risk as row 46, in the primary architecture doc |
|
||||
| 49 | `docs/infrastructure/standby-host-replacement-2026-09-14.md` §"Rebuild procedure" (final section) | `Reference the hermes-standby-deployment skill. Sequence: create CPX21 in fsn1 or nbg1, run the standby deploy script, restore from s3://hermes-vps-backups/live/, verify the failover cron...` | This entire rebuild procedure is Hetzner-specific and becomes fully inapplicable; either delete this section or replace with the netcup core-bu equivalent procedure once one exists | A "how to rebuild the standby" procedure that references a decommissioned provider's rescue-mode tooling is actively harmful if followed after retirement |
|
||||
| 50 | `/root/.hermes/references/network-diagram.md` lines 26-27, 33-39, 54, 67 | ASCII diagram box labeled `[Standby Host] (app1-bu)`; `[Legacy Net] (Hetzner)` box listing decommissioned services plus implying app1-bu lives there too; `**Standby Host (core-bu)**: 5.161.225.131 — Warm standby (CPX21), connected via Tailscale to Core.` (again the naming collision); Tailscale Mesh description says "between Primary Host (Core) and Standby Host (core-bu)" | Redraw the ASCII diagram: Standby Host box becomes core-bu at 159.195.204.203 (netcup box, not "Legacy Net"/Hetzner); remove the app1-bu box or move it to a "retired" annotation | ASCII diagrams are easy to skip during edits because they're not table rows — this file's diagram will visually lie about network topology post-retirement if not redrawn |
|
||||
| 51 | `/root/.hermes/references/master-apps-services.md` (10 occurrences: lines 19, 30, 260-271, 307, 356-367, 386, 421-423) | Extensive core-bu-as-Hetzner-name usage: `**core-bu** \| 5.161.225.131 \| CPX11 (2C/2G/40G) \| Hetzner \| Warm standby for Core`; `### core-bu (Hetzner CPX11 — 5.161.225.131) — Warm Standby`; `Tailscale \| BSD \| Core, core-bu`; `**Hermes Agent** \| Proprietary \| Core, core-bu`; `**DR Standby:** core-bu (Hetzner) boots and auto-restores from S3 if Core is down for 2+ minutes.` | Ten separate lines in one document all need the app1-bu/core-bu rename at P1 (existing collision) and then a further retirement edit at P3. This is the single most-referenced document for the naming collision found in this sweep | Ten independent edit points in one file is exactly the kind of surface a manual sweep misses one or two of — recommend a scripted find/replace pass on this file specifically, verified line-by-line afterward |
|
||||
| 52 | Gitea `disaster-recovery` repo (confirmed reachable via API, `git.itpropartner.com/ippadmin/disaster-recovery`, last commit referenced in `restore-test-log.md` line 43: `disaster-recovery.git \| 5 \| 3bc6d08 Fix: app1-bu CPX11 → CPX21 spec`) | Repo history shows this repo has previously been edited specifically for app1-bu spec corrections — implying it contains its own copy of DR content that will need the same app1-bu→core-bu treatment | Clone and sweep this repo's content directly (not done in this audit — only confirmed reachability and one commit-log reference via the local `restore-test-log.md`) before declaring the retirement documentation complete | This repo was NOT directly inspected in this pass (Core's local checkout doesn't contain it) — flagged as an open item, not a completed row |
|
||||
|
||||
---
|
||||
|
||||
## Artifacts referencing hosts that no longer exist (already stale, independent of app4/core-bu/app1-bu)
|
||||
|
||||
These were found while sweeping the same files and are stale right now, unrelated to the current provisioning work. Listed per the task's explicit request.
|
||||
|
||||
| Host / IP | Where still referenced (live, not historical-annotated) | Status per `decommissioned-hosts.json` |
|
||||
|---|---|---|
|
||||
| `wphost02` / `5.161.62.38` | `docs/monitoring/uptime-kuma-monitoring-plan.md` §2.5 header still titled "wphost02 (5.161.62.38) — RunCloud (LEGACY — being migrated)" (present tense, "being migrated", not "migrated"); `/root/.hermes/references/network-diagram.md` line 59 "wphost02: 5.161.62.38 — DECOMMISSIONING" (present-progressive, not past); `master-apps-services.md` line 25/43/70/245/275/386/421 refer to wphost02 in present tense in several spots even though the file elsewhere (line correctly) notes migration | Decommissioned 2026-08-28, confirmed in `decommissioned-hosts.json` and `doc-live-verify.py`'s `known_historical` |
|
||||
| `178.156.130.130` (old standalone Hudu) | `network-diagram.md` does not list it directly but `decommissioned-hosts.json` carries it with `"decommissioned": null` (the field is present but not dated, unlike the other entries) | Listed in graveyard but missing a decommission DATE — a data-quality gap in the source-of-truth file itself |
|
||||
| `5.161.114.8` (old app1-bu IP) | `/root/.hermes/references/hermes-dr-plan-v2.md` line 12: `app1-bu (Hetzner CPX21 — 5.161.114.8)`; `itpp-recovery-manual.md` TOC anchor `#4-app1-bu-standby-hetzner--51611148` still embeds this dead IP in the anchor slug even though the section body correctly uses 5.161.225.131 | Decommissioned 2026-07-24, correctly in `known_historical`, but doc-live-verify.py's own DR audit (dr-issue-log.md, 2026-09-15 entry) already flags this exact hit as "benign" since the file is historical — confirms the scanner is working as designed for THIS one, but the two live-doc mentions above were not caught because doc-live-verify skips files with decommission-marker words nearby, not exact-line context |
|
||||
| `178.156.167.181` (old admin-ai) | `hetzner-server-inventory.md` (already self-marked "ARCHIVED — Historical Reference Only") and `app-inventory.csv` (also self-marked ARCHIVED) — both correctly annotated | Decommissioned, correctly archived |
|
||||
| `87.99.144.163` (old app1) | `decommissioned-hosts.json` graveyard entry has `"decommissioned": null` — no date | Same data-quality gap as `178.156.130.130` |
|
||||
| `87.99.159.142` (Tony's old VPS) | `/root/.hermes/references/ip-dns-changes.md` line 95: `tony.iamgmb.com \| 87.99.159.142 \| Tony VPS \| Tony's Hermes` — this is a LIVE DNS record row in a document dated August 12, 2026, for a server the same file's own line 26 lists as "DELETED Jul 14" | Contradiction WITHIN the same document: line 26 says deleted, line 95 lists it as a current DNS target |
|
||||
|
||||
---
|
||||
|
||||
## Top 5 highest-risk stale references found
|
||||
|
||||
1. **README.md line 277 — "provider diversity" claim becomes factually false, not just outdated.** Once core-bu (netcup) replaces or joins app1-bu (Hetzner) as Core's standby, the stated rationale "core-bu stays at Hetzner specifically so a netcup outage can't kill both Core and standby simultaneously" is wrong the moment the standby is also netcup. This is a substantive risk claim in the primary README, not a cosmetic IP mismatch, and the same false claim is repeated in `docs/architecture.md` line 17 and `master-apps-services.md` line 271.
|
||||
|
||||
2. **The "core-bu" naming collision already exists across 6 files before the new box is even ordered.** `README.md`, `docs/infrastructure/key-inventory.md`, `network-diagram.md`, `master-apps-services.md`, `restore-runbooks.md`, and `ip-dns-changes.md` all currently call the Hetzner box (5.161.225.131) "core-bu" — exactly the name the owner's plan assigns to the new netcup box (per `server-architecture-plan` skill's 2026-09-14 correction note). If this isn't fixed before core-bu goes live, every future reference to "core-bu" in these six files is ambiguous between two physically different servers with different providers, specs, and failure domains.
|
||||
|
||||
3. **`itpp-recovery-manual.md` — 8 live SSH commands hardcoded to `root@5.161.225.131` inside the master recovery runbook.** This is the document opened during an actual incident. Post-retirement, every one of these commands fails, and the failure mode during a live outage (typing a command that connects to nothing, or worse, to a re-leased IP with a different owner) is worse than having no runbook at all.
|
||||
|
||||
4. **`hermes-standby-deployment` skill and `hermes-dr-plan-v2.md` are entirely Hetzner-API-specific** (rescue mode, `enable_rescue`, Hetzner SSH key registration, CPX11→CPX21 upgrade math) and become 100% inapplicable to a netcup-based core-bu. There is currently no equivalent skill or plan describing how to deploy/audit/troubleshoot a netcup-based warm standby — this is a capability gap, not just a stale reference, and it will be discovered mid-incident if not addressed before app1-bu is actually retired.
|
||||
|
||||
5. **Two live Tailscale nodes currently exist for one physical box** (`app1-bu` at 100.112.23.21, offline 61 days; `app1-bu-1` at 100.95.212.28, idle/active) — found live during this audit, not from a doc. This is a pre-existing data-quality problem in the mesh itself (documented as a known Tailscale auto-rename behavior in the `hermes-standby-deployment` skill's pitfalls, but never cleaned up) that will complicate identifying which Tailscale node to decommission when app1-bu is formally retired, and could cause the retirement checklist to miss one of the two.
|
||||
|
||||
---
|
||||
|
||||
## Counts
|
||||
|
||||
- **Total artifacts (files) touched by at least one required edit:** 33 distinct files/scripts/configs identified with concrete row-level evidence, plus 1 remote Gitea repo flagged as unswept (`disaster-recovery`), plus 1 external Uptime Kuma admin UI not inspected (config lives outside the filesystem, referenced only via `uptime-kuma-monitoring-plan.md`).
|
||||
- **P1 (provisioning-time):** 23 rows (rows 1-23)
|
||||
- **P2 (service cutover):** 10 rows (rows 24-33)
|
||||
- **P3 (app1-bu retirement):** 19 rows (rows 34-52)
|
||||
- **Already-stale references (pre-existing, independent of this migration):** 6 items (wphost02 present-tense language, 2 undated graveyard entries, 1 dead-IP anchor slug, 1 contradiction within ip-dns-changes.md re: Tony's VPS)
|
||||
|
||||
Row counts were verified against the table bodies above (23 + 10 + 19 = 52 total numbered rows, matching the highest row number used).
|
||||
@@ -0,0 +1,127 @@
|
||||
# Standby Host Replacement: app1-bu Cost and Location Review
|
||||
|
||||
**Date:** 2026-09-14
|
||||
**Status:** Recommendation, awaiting decision
|
||||
**Author:** Sho'Nuff
|
||||
**Source of truth:** Hetzner invoice 086001081978 (2026-08-27, period 07/2026) + live Hetzner Cloud API pull, 2026-09-14
|
||||
|
||||
## Verdict (BLUF)
|
||||
|
||||
The DR plan records app1-bu at roughly $14/mo. That is wrong. The real figure is **EUR 31.99/mo for the CPX21 alone, EUR 37.06/mo all in** (about USD 40.39). The fix is not a new provider. It is the **same box in a different location**: Hetzner Falkenstein or Nuremberg prices the identical CPX21 at **EUR 9.49/mo**.
|
||||
|
||||
**Recommendation: move the standby from Ashburn to fsn1 or nbg1. Saves EUR 22.50/mo (USD 294/yr) and closes a real diversity gap at the same time.**
|
||||
|
||||
## What the invoice actually shows
|
||||
|
||||
Invoice 086001081978, period 07/2026, net EUR 114.44, 0% VAT.
|
||||
|
||||
It is not one server. The July bill carried eleven servers mid-teardown plus supporting resources:
|
||||
|
||||
| Line | Item | Qty | Unit | EUR |
|
||||
|---|---|---|---|---|
|
||||
| 1 | Backup (20% of instance price) | 3 | | 7.23 |
|
||||
| 2 | CPX11 | 1 hr | 0.0280 | 0.03 |
|
||||
| 3 | CPX11 | 4 x 1,368 hr | 0.0096 | 13.13 |
|
||||
| 4 | CPX21 | 3 x 1,069 hr | 0.0192 | 20.52 |
|
||||
| 5 | CPX21 | 2 x 571 hr | 0.0513 | 29.29 |
|
||||
| 6 | CPX21 | 1 month | 11.99 | 11.99 |
|
||||
| 7 | CPX41 | 1 x 359 hr | 0.0625 | 22.44 |
|
||||
| 8 | Primary IPv4 | 11 x 3,368 hr | 0.0008 | 2.69 |
|
||||
| 9 | Primary IPv4 | 1 month | 0.5000 | 0.50 |
|
||||
| 10 | Snapshot | 462.289 GB-mo | 0.0143 | 6.61 |
|
||||
| 11-13 | Traffic | | | 0.00 |
|
||||
|
||||
The EUR 114.44 is a July artifact, not today's run rate. Most of those servers no longer exist.
|
||||
|
||||
Note line 5: two CPX21 instances at **EUR 0.0513/hr**. That is the Ashburn hourly rate, and it is the same rate the API reports today. Line 4 at EUR 0.0192/hr is the EU rate. The 3.5x US premium is visible inside a single invoice.
|
||||
|
||||
## What is actually live today
|
||||
|
||||
Live API pull, 2026-09-14. Exactly **one** server remains:
|
||||
|
||||
- **app1-bu.itpropartner.com**, id 151357515, CPX21, 3 vCPU / 4 GB / 80 GB, location **ash** (Ashburn, VA), status running, created 2026-07-15, public IPv4 5.161.225.131, backups **disabled**
|
||||
- 2 primary IPs (IPv4 + IPv6), both assigned to that server
|
||||
- 9 snapshots, 319.2 GB total
|
||||
- 0 volumes, 0 firewalls
|
||||
|
||||
### Current monthly run rate
|
||||
|
||||
| Component | EUR/mo |
|
||||
|---|---|
|
||||
| CPX21 in Ashburn | 31.99 |
|
||||
| Primary IPv4 | 0.50 |
|
||||
| Snapshot storage (319.2 GB at 0.0143) | 4.57 |
|
||||
| **Total** | **37.06** (~USD 40.39) |
|
||||
|
||||
### Same stack in an EU location
|
||||
|
||||
| Component | EUR/mo |
|
||||
|---|---|
|
||||
| CPX21 in fsn1 or nbg1 | 9.49 |
|
||||
| Primary IPv4 | 0.50 |
|
||||
| Snapshot storage | 4.57 |
|
||||
| **Total** | **14.56** (~USD 15.87) |
|
||||
|
||||
**Delta: EUR 22.50/mo, EUR 270.00/yr, about USD 294/yr.**
|
||||
|
||||
## The diversity finding
|
||||
|
||||
The rule is that a netcup outage must not take out both Core and its standby. Today both sit in the same US East corridor: Core in Manassas, VA and app1-bu in Ashburn, VA. That is roughly 150 miles and one weather system, one grid region, one set of upstream transit providers.
|
||||
|
||||
Provider diversity is satisfied. Regional diversity is not.
|
||||
|
||||
Falkenstein or Nuremberg fixes it. The standby would move from the same corridor as production to a separate continent, which is what a warm standby is supposed to be.
|
||||
|
||||
## Options
|
||||
|
||||
**Option A: Move the standby to Hetzner fsn1 or nbg1 (RECOMMENDED)**
|
||||
- Cost: EUR 14.56/mo all in, down from EUR 37.06
|
||||
- Keeps the entire runbook: Hetzner API, rescue mode with key injection, snapshot tooling, S3 restore path
|
||||
- Improves regional diversity from 150 miles to 4,000
|
||||
- Latency Core to standby rises from roughly 10ms to roughly 95ms. The 10-minute failover poll does not care, and a 1.17 GB DB sync at that latency is still tens of seconds
|
||||
- Reversibility: high. This is a rebuild from the standby deploy script plus an S3 restore. Nine snapshots and the S3 backup chain remain as the safety net throughout
|
||||
- Risk: low. Same provider, same tooling, same image lineage
|
||||
|
||||
**Option B: Move to a US provider with monthly billing and a real API (DigitalOcean 2 GiB verified at USD 12.00/mo)**
|
||||
- Keeps latency low and adds a genuinely separate provider on top of netcup
|
||||
- Breaks the API-driven recovery runbook, which is Hetzner-specific. Rescue mode, key injection, and scripted power control would all need rewriting and re-testing
|
||||
- Price per resource is worse than Hetzner EU: USD 12.00 for 2 GiB / 1 vCPU against EUR 9.49 for 4 GiB / 3 vCPU
|
||||
- Vultr and Akamai/Linode price per region, so a specific number requires a per-region pull
|
||||
|
||||
**Option C: Stay in Ashburn at EUR 31.99**
|
||||
- Rejected. Paying 3.4x for the same instance, in the same corridor as production
|
||||
|
||||
**Rejected: SSD Nodes prepaid**
|
||||
- 14-day refund window then a 1 to 3 year commitment, an overselling reputation, and a thin API. Wrong profile for the asset that has to work when everything else is down
|
||||
|
||||
## Dead storage on the account
|
||||
|
||||
Snapshot review, same API pull:
|
||||
|
||||
- 3 wphost02 snapshots, **151.6 GB, EUR 2.17/mo**, for a host decommissioned 2026-08-28
|
||||
- 1 legacy snapshot "Back-online" (2026-06-05), **112.1 GB, EUR 1.60/mo**, superseded by the current weekly chain
|
||||
- 1 legacy snapshot "migration-complete" (2025-11-23), 7.6 GB, EUR 0.11/mo
|
||||
- 4 current app1-bu weeklies, 48.0 GB, EUR 0.69/mo, KEEP
|
||||
|
||||
Removing the decommissioned and superseded snapshots saves **EUR 3.88/mo, EUR 46.55/yr**. Not proposed for silent deletion. Confirm first, per standing rule.
|
||||
|
||||
## Combined opportunity
|
||||
|
||||
| Action | EUR/mo | USD/yr |
|
||||
|---|---|---|
|
||||
| Relocate standby to EU | 22.50 | 294 |
|
||||
| Retire dead snapshots | 3.88 | 51 |
|
||||
| **Total** | **26.38** | **345** |
|
||||
|
||||
## Record correction required
|
||||
|
||||
The DR plan records the app1-bu line as a CPX21 at roughly $14/mo. Both halves are wrong:
|
||||
|
||||
1. **Price:** the Ashburn CPX21 is EUR 31.99/mo, not about $14
|
||||
2. **Spec:** CPX21 is 3 vCPU / 4 GB / 80 GB, not 4C/8G
|
||||
|
||||
Corrections should land in the DR plan v2 and any derived runbook once the location decision is made, so the standby budget is not understated again.
|
||||
|
||||
## Rebuild procedure
|
||||
|
||||
Reference the `hermes-standby-deployment` skill. Sequence: create CPX21 in fsn1 or nbg1, run the standby deploy script, restore from s3://hermes-vps-backups/live/, verify the failover cron and the 10-minute health check, then destroy the Ashburn instance. Keep the Ashburn box alive until the EU standby passes a restore test. Backed up is not finished; a confirmed restore test is.
|
||||
Reference in New Issue
Block a user