Files
itpp-infrastructure/backup-dr-audit-2026-08-10.md
T
root f5175f1ce0 Sync docs, audit artifacts, project notes, and VerdictTank proposal docs
- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
2026-08-26 02:27:28 -04:00

30 KiB
Raw Blame History

ITPP Backup & DR — Full Audit Report

Date: 2026-08-10 ~20:45 ET
Auditor: Hermes Agent (automated audit)
Scope: All ITPP infrastructure — Core, app1, app2, app3, wphost02, MikroTik CCR
Storage Verified: Wasabi S3 (hermes-vps-backups, mikrotik-ccr-backups)


Executive Summary

  • Overall grade: D — Multiple active data-loss risks, no restore testing, several services completely unprotected.
  • Critical gaps (items that risk data loss RIGHT NOW):
    1. app3 MySQL backups stopped 2026-08-08 — 9 databases (apextrackexperience, boxpilotlogistics, debtrecoveryexperts, iAmGMB, intelsight, katiewattsdesign, mainwp, vigilanttac, voipsimplicity_site) have no backup for Aug 910. Two days of production data completely unprotected.
    2. git.modelortho.com (Anita's Gitea) has a backup script and Hermes cron, but the cron has NEVER executed — the cron job exists with no Last run timestamp. Only one manual backup exists on S3.
    3. Prometheus TSDB — ZERO backups evercore/prometheus/ S3 prefix is completely empty. All monitoring history, alert rules, and dashboards are unprotected.
    4. 10+ production services have NO backup — buzz-prod (Block Buzz relay + PostgreSQL + Redis + MinIO), support-api, bookstack (×2), searxng, timetrex, microbin, camofox-browser, browserless, msp-forms, docs-auth-validator, transitpin-api, transitpin-relay.
    5. ZERO evidence of any verified restore — backups exist but nobody has ever tested if they can actually be restored.
  • High-priority improvements:
    1. Fix app3 MySQL backup immediately
    2. Get gitea-modelortho cron running (it exists but hasn't fired)
    3. Create backups for all unbacked services
    4. Stand up Prometheus backup
    5. Perform a restore test of at least one backup within 7 days

1. Backup Inventory & Verification

Legend

  • = Healthy (script exists, runs, S3 file present, size normal)
  • ⚠️ = Warning (present but suspicious — unchanging size, stale data, missing days)
  • 🔴 = Failed (missing entirely, script broken, zero files)
  • 👻 = Phantom (script in plan but doesn't exist on disk)
  • = Unknown (could not verify execution)

Core (152.53.192.33)

Target Script Script Exists? S3 Files (7d) Size Healthy? Status Notes
Hermes Full Backup hermes-backup.sh (system cron 1AM) Core Aug 410 daily, 1.21.3 GB Growing
Hermes Live Sync hermes-live-sync.sh (Hermes cron, every 15m) Core ⚠️ Only verification_evidence.db, no state.db ⚠️ Stale ⚠️ live/ has old malformed backups from Jul 9; current sync may go elsewhere
/root Essentials root-essentials-backup.sh (system cron 3AM) Core Aug 410 daily, 170236 MB Growing
Grafana core-services-backup.sh (system cron 1:30AM) Core Aug 410 daily, ~52 KB Consistent
Uptime Kuma core-services-backup.sh Core Aug 410 daily, ~100 MB Growing
Docker Volumes core-services-backup.sh Core ⚠️ vaultwarden-data only, stale since Jul 28 ⚠️ Stale ⚠️ vaultwarden-data backed up but service migrated; no other volumes visible
Prometheus core-services-backup.sh Core 🔴 EMPTY 🔴 Zero bytes 🔴 core/prometheus/ has NO files ever
Auth API auth-api-backup.sh (Hermes cron 3:15AM) Core Aug 810 daily, ~41 KB Consistent

App1 (152.53.36.131)

Target Script Script Exists? S3 Files (7d) Size Healthy? Status Notes
Open WebUI /root/backup.sh (app1 cron 2AM) app1 ⚠️ Aug 410 daily ⚠️ Identical size 5 days (1,470,272,729 bytes Aug 59) ⚠️ Size finally changed Aug 10 (1,470,769,575). Docker cp may capture stale data
LiteLLM DB litellm-backup.sh (Hermes cron 3:30AM) Core Aug 410 daily, 1521 MB, growing Growing Also has config-only files from app1 backup.sh (240 bytes)
n8n /root/backup.sh (app1 cron 2AM) app1 Aug 410 daily, ~6064 KB Consistent
MCP Configs /root/backup.sh (app1 cron 2AM) app1 Aug 410 daily, 373 bytes Consistent Very small — configs only
Vaultwarden vaultwarden-backup.sh (Hermes cron 2:30AM) Core Aug 410 daily, 520755 KB Growing
Komodo komodo-backup.sh (Hermes cron 3:45AM) Core Aug 410 daily, ~12.7 KB Consistent
DocuSeal docuseal-backup.sh (Hermes cron 4AM) Core Aug 410 daily, ~255 KB Consistent
Twenty CRM twenty-backup.sh (Hermes cron 4:15AM) Core Aug 310 daily, ~156 KB Consistent Also has twenty-files tarballs from app1 backup.sh
Kokoro TTS (stateless) N/A N/A N/A

App2 (152.53.39.202)

Target Script Script Exists? S3 Files (7d) Size Healthy? Status Notes
Traccar /root/backup.sh (app2 cron 2:30AM) app2 Aug 410 daily, ~1.8 MB Slightly growing
Gitea gitea-backup.sh (Hermes cron 8AM) Core Aug 410 daily (folders w/ db + repos) DB 2.7 MB, 30+ repos
Hudu hudu-backup.sh (Hermes cron 7AM) Core Aug 410 daily, 670714 KB Growing
UNMS unms-backup-sync.sh (Hermes cron 6AM) Core Daily since Aug 5, ~100 MB Consistent Gap Jul 31Aug 5
UniFi unifi-backup-sync.sh (Hermes cron 2AM) Core Aug 410 daily, ~884893 KB Consistent
Technitium DNS technitium-backup.sh (Hermes cron 2:45AM) Core Aug 610 daily, 369461 KB, growing Growing Also has duplicates from app2 backup.sh
Dawarich dawarich-backup.sh (Hermes cron 4AM) Core Aug 610 daily, ~7 MB Consistent Also backed up by app2 backup.sh
RAGFlow ragflow-backup.sh (Hermes cron 4:15AM) Core Aug 810 daily, ~267 KB ⚠️ Unchanging ⚠️ Identical size all 3 days — very small for MySQL

App3 (152.53.241.111)

Target Script Script Exists? S3 Files (7d) Size Healthy? Status Notes
CloudPanel DB /root/backup.sh (app3 cron 3AM) app3 Aug 410 daily, 1.41.6 MB Growing
MySQL (all DBs) /root/backup.sh (app3 cron 3AM) app3 🔴 Last: Aug 8 🔴 No Aug 910 🔴 CRITICAL: 9 DBs stopped backing up after Aug 8. No mysql lines in Aug 10 log.
WordPress Files /root/backup.sh (app3 cron 3AM) app3 Aug 9 daily, 9 sites Normal wp-www identical size for 10 days is suspicious but other sites vary
Nginx Configs /root/backup.sh (app3 cron 3AM) app3 Aug 48 daily, ~12 MB Growing ⚠️ FAILED on Aug 10 per cron log: "Nginx configs: FAILED"
Static Sites /root/backup.sh (app3 cron 3AM) app3 Aug 910 daily, 17 sites Normal Includes modelortho.com, buzz.iamgmb.com, transitpin.com, etc.
WP Snapshots (local) /opt/backup-restore/snapshot.sh (app3 cron 1AM/1PM) app3 Local only Not verified Local snapshots, not in S3
Hexclave (Stack Auth) hexclave-backup.sh (Hermes cron 3:30AM) Core Aug 810 daily, ~315 KB Consistent
modelortho.com modelortho-backup.sh (listed in plan) 👻 MISSING N/A N/A 👻 Does not exist. However static modelortho.com IS backed up via app3/static.
git.modelortho.com (Anita's Gitea) gitea-modelortho-backup.sh (Hermes cron 4:30AM) Core 🔴 Only 1 file (manual?) 🔴 Cron never ran 🔴 CRITICAL: Script exists, cron exists, but cron has NO Last run. Zero automated backups.

wphost02 (5.161.62.38)

Target Script Script Exists? S3 Files (7d) Size Healthy? Status Notes
WordPress (7 sites) /root/backup.sh (Core cron 5AM via SSH) wphost02 Aug 410 daily Normal (50900 MB per site) 7 sites + all-databases.sql + RunCloud config

MikroTik CCR

Target Script Script Exists? S3 Files (7d) Size Healthy? Status Notes
Home Gateway run-wisp-backup.sh (Hermes cron 6AM) Core Daily config .rsc Normal ⚠️ Home gateway OK, but 5 tower CCRs all timeout (T01T04, MP100)
WISP Tower CCRs Same script Core 🔴 None since Jul 7 🔴 🔴 SSH timeout to all towers. Only 1 historical config from Jul 7.

Hetzner Snapshots

Target Script Script Exists? S3 Files (7d) Status Notes
Weekly disk snapshots snapshot-hetzner.py (Hermes cron Mon 5AM) Core N/A (Hetzner API, not S3) Last run Aug 10, ok

2. Missing Backups

The backup plan's "Unbacked Services" section claims "(none)" — this is incorrect. The following services are running in production with ZERO backup coverage:

Core (152.53.192.33)

Service Type Data at Risk Priority
searxng Docker (search engine) Configuration, any cached indexes Low
timetrex Docker (time tracking) Employee time data, payroll records High
microbin Docker (pastebin) Shared text snippets Low
camofox-browser Docker (browser automation) Stateless? Low
browserless Docker (headless browser) Stateless Low
Prometheus systemd (monitoring TSDB) All metrics history, alert rules, dashboards High

App1 (152.53.36.131)

Service Type Data at Risk Priority Notes
Wazuh (SIEM/XDR) Docker (×3 containers) Security events, agent configs, alerts High Has wazuh-backup.sh on app1 and daily S3 files, but NOT in backup plan or Hermes cron. Backup works but is undocumented/unanchored.

App2 (152.53.39.202)

Service Type Data at Risk Priority
support-api Docker (support API) Support ticket data High
bookstack Docker (wiki) Documentation wiki content High
bookstack-db Docker (MariaDB for bookstack) Wiki database High
happy_rosalind Docker (bookstack duplicate?) Unknown content Medium

App3 (152.53.241.111)

Service Type Data at Risk Priority
buzz-prod-relay Docker (Block Buzz relay) Relay configuration, keys High
buzz-prod-postgres Docker (Buzz PostgreSQL) All Buzz relay state Critical
buzz-prod-redis Docker (Buzz Redis) Session/cache data Medium
buzz-prod-minio Docker (Buzz object storage) Uploaded files Medium
transitpin-api systemd Transportation API data Medium
transitpin-relay systemd Relay config Medium
msp-forms systemd Form submissions Medium
docs-auth-validator systemd Auth validation state Low

3. 3-2-1 Compliance

The 3-2-1 rule states: 3 copies of data, on 2 different media, with 1 off-site.

Service 3 Copies? 2 Media? 1 Off-Site? Verdict
Hermes Full Backup (live + daily S3 + standby) ⚠️ (all Wasabi S3) (Wasabi us-east-1) PARTIAL — single media type
App1 services (LiteLLM, n8n, Vaultwarden, etc.) (server + S3) ⚠️ (all Wasabi S3) PARTIAL
App2 services (Gitea, Hudu, Traccar, etc.) (server + S3) ⚠️ (all Wasabi S3) PARTIAL
App3 MySQL 🔴 (server only, no S3 since Aug 8) 🔴 🔴 FAIL
App3 WordPress (server + S3 + local snapshots) (S3 + local disk) PASS
App3 static sites (server + S3) ⚠️ (all Wasabi S3) PARTIAL
wphost02 WordPress (server + S3) ⚠️ (all Wasabi S3) PARTIAL
git.modelortho.com 🔴 (only server, cron never ran) 🔴 🔴 FAIL
Prometheus 🔴 (server only) 🔴 🔴 FAIL
Wazuh (server + S3) ⚠️ (all Wasabi S3) PARTIAL
Buzz / Bookstack / Support-API / etc. 🔴 (server only) 🔴 🔴 FAIL
MikroTik Home Gateway (router + S3) ⚠️ (all Wasabi S3) PARTIAL
MikroTik Tower CCRs 🔴 (router only) 🔴 🔴 FAIL

Overall 3-2-1 compliance: FAIL

The strategy relies entirely on Wasabi S3 for off-site storage — there is no second media type (no local NAS, no tape, no separate cloud provider). Every service that passes does so only by counting the production server + S3 as two copies. For true 2-media compliance, a second storage type (e.g., local NAS, separate cloud provider, or physical media) would be required.


4. RPO/RTO Assessment

Service Tier Planned RPO Current RPO Acceptable? Planned RTO Current RTO Est. Acceptable?
Hermes Agent Critical ≤1hr ~15 min (live sync) ≤4hr ~24hr
Gitea (app2) Critical ≤1hr 24hr (daily only) ⚠️ ≤4hr ~24hr
Traccar Critical ≤1hr 24hr (daily only) ⚠️ ≤4hr ~24hr
UISP (UNMS) Critical ≤1hr 24hr ⚠️ ≤4hr ~48hr ⚠️
LiteLLM High 24hr 24hr ≤8hr ~48hr
n8n High 24hr 24hr ≤8hr ~48hr
Open WebUI High 24hr 24hr ≤8hr ~48hr
Vaultwarden High 24hr 24hr ≤8hr ~24hr
Twenty CRM High 24hr 24hr ≤8hr ~48hr
Hudu Medium 24hr 24hr ≤24hr ~824hr
UniFi Medium 24hr 24hr ≤24hr ~48hr
Komodo Medium 24hr 24hr ≤24hr ~24hr
DocuSeal Medium 24hr 24hr ≤24hr ~24hr
App3 WP sites Medium 24hr 24hr ≤24hr ~824hr
Auth API Medium 24hr 24hr ≤24hr ~24hr
Hexclave (Stack Auth) Medium 24hr 24hr ≤24hr ~48hr
Grafana Low 24hr 24hr ≤48hr ~24hr
Uptime Kuma Low 24hr 24hr ≤48hr ~24hr
Prometheus Low 24hr 🔴 ∞ (no backup) 🔴 ≤48hr 🔴 Impossible 🔴
MikroTik CCR Low 24hr 24hr (home) / 🔴 ∞ (towers) 🔴 ≤48hr ~48hr (home) / 🔴 Impossible (towers) 🔴
Technitium DNS Low 24hr 24hr ≤48hr ~24hr
Dawarich Low 24hr 24hr ≤48hr ~48hr
RAGFlow Low 24hr 24hr ⚠️ ≤48hr ~48hr
Buzz / Bookstack / Support-API / etc. Unclassified N/A 🔴 ∞ (no backup) 🔴 N/A 🔴 Impossible 🔴
git.modelortho.com Unclassified N/A 🔴 ∞ (cron never ran) 🔴 N/A 🔴 Possible (script exists) 🔴
app3 MySQL Medium 24hr 🔴 ~48hr and growing 🔴 ≤24hr ⚠️ Partial (Aug 8 dump) ⚠️

RPO/RTO Assessment: FAIL for Critical tier — Gitea, Traccar, and UISP are planned for ≤1hr RPO but receive only daily backups. The plan itself overstates capabilities.


5. Restore Testing

Evidence of Verified Restores: NONE

Source What it says Evidence of execution
backup-plan.md §Restore Testing Cadence "Quarterly: Pick one random backup per tier, restore to staging, verify integrity" 🔴 No test logs anywhere
backup-plan.md "Annual: Full DR simulation" 🔴 Never executed
4-DR-Testing-Schedule.md "Every restore test gets logged with date, tester, duration, findings" 🔴 No log file exists
backup-policy.md "Monthly restore test of a randomly selected asset" 🔴 No evidence
3-Per-Server-Runbooks.md Detailed per-server restore procedures Procedures exist on paper only
dr-issue-log.md Historical DR audit issues Issues tracked; no restore tests documented
/root/.hermes/references/ Search for "restore test", "verified restore", "DR simulation" 🔴 Zero actual test results found

Conclusion: Restore procedures are well-documented on paper, but no backup has ever been restore-tested. There are no test logs, no verification records, and no evidence that any of the 30+ backup targets can actually be restored successfully.

Restore test cadence last executed: NEVER


6. Monitoring & Alerting

What exists:

Check Mechanism Status
Hermes full backup age check backup-audit-check.sh (system cron 2AM) Runs daily, checks if last backup <36hr old
Cron output logging All backup cron jobs pipe to logger Output captured in syslog
Hermes cron job status hermes cron list shows last run status Most show "ok"
Uptime Kuma monitoring External monitoring of services
Ops data collector ops-data-collector.py (every 5min) Collects metrics

What's MISSING:

Gap Impact Urgency
No alert on backup failure When backup-audit-check.sh fails, output goes to logger only. Nobody gets notified. Critical
No alert on 0-byte backup Several backups produce tiny config files (240 bytes) that could silently become 0 bytes High
No alert on backup not running Hermes cron jobs silently enter "error" state (home-router-daily-backup has been error for days) Critical
No cross-server backup health dashboard Must SSH to each server individually to check backup status High
No S3 file integrity validation Nobody checks if backup files are actually valid (corrupt tar, truncated SQL) High
Core crontab doesn't monitor remote server backup outcomes Core only checks its own full backup — app1/app2/app3 backup failures go unnoticed Critical
No automatic ticket/issue creation Backup failures should auto-create a GitHub issue or Hudu ticket Medium

The backup-audit-check.sh is the only monitoring, and it only checks Hermes full backup freshness. Everything else is silent.

Cron Job Statuses:

Job Schedule Latest Status Issues
hermes-backup (system cron) 1 AM Ran today
backup-audit-check (system cron) 2 AM OK (20h ago)
root-essentials-backup (system cron) 3 AM S3 files present
core-services-backup (system cron) 1:30 AM S3 files present Prometheus part empty
wphost02-backup (system cron via SSH) 5 AM S3 files present
vaultwarden-backup (Hermes cron) 2:30 AM ok
litellm-backup (Hermes cron) 3:30 AM ok
komodo-backup (Hermes cron) 3:45 AM ok
docuseal-backup (Hermes cron) 4:00 AM ok
twenty-backup (Hermes cron) 4:15 AM ok
auth-api-backup (Hermes cron) 3:15 AM ok
stack-auth-backup (Hermes cron) 3:15 AM ok
hexclave-backup (Hermes cron) 3:30 AM ok
technitium-backup (Hermes cron) 2:45 AM ok
dawarich-backup (Hermes cron) 4:00 AM ok
ragflow-backup (Hermes cron) 4:15 AM ok
hudu-backup (Hermes cron) 7:00 AM ok
gitea-backup (Hermes cron) 8:00 AM ok
unms-backup-sync (Hermes cron) 6:00 AM ok
unifi-backup-sync (Hermes cron) 2:00 AM ok
hetzner-weekly-snapshots (Hermes cron) Mon 5 AM ok
home-router-daily-backup (Hermes cron) 6:00 AM 🔴 error 5/6 towers timeout daily
Gitea ModelOrtho Backup (Hermes cron) 4:30 AM 🔴 NEVER RAN No Last run field
app1 /root/backup.sh (app1 crontab) 2:00 AM Not verified S3 files present but openwebui sizes suspicious
app2 /root/backup.sh (app2 crontab) 2:30 AM Not verified S3 files present
app3 /root/backup.sh (app3 crontab) 3:00 AM ⚠️ MySQL failed MySQL dumps stopped Aug 8; nginx failed Aug 10
app3 /opt/backup-restore/snapshot.sh (app3 crontab) 1AM/1PM Not verified Local-only

7. Industry Standard Gaps

What we're doing right:

  • Comprehensive backup plan — well-documented in backup-plan.md with 27 declared targets
  • 24+ Hermes cron jobs operating backup scripts — good automation
  • Daily cadence for all critical services
  • Wasabi S3 as immutable off-site storage (11-nines durability)
  • Provider diversity — Core on netcup, standby on Hetzner
  • Warm standby for Hermes (app1-bu with 7-min failover)
  • DR runbooks documented per server
  • DR issue log maintained with root cause analysis
  • Hetzner weekly snapshots as additional layer

What's missing (every competent shop has these):

Gap Severity Industry Expectation
Restore testing Critical Monthly restore tests are table stakes. "Untested backups are not backups."
Backup failure alerting Critical PagerDuty/OpsGenie/Telegram alert on ANY backup job failure
Second media type High NAS, tape, or second cloud provider for 3-2-1 compliance
Immutable backups High S3 Object Lock to prevent ransomware deletion
Backup integrity validation High Automated restore-and-verify pipeline (spin up temp container, restore DB, run queries)
Service discovery for backups High Automated scan that identifies running services and flags unbacked ones
Backup retention policy Medium Defined retention periods per tier (only gitea-modelortho and wphost02 have cleanup)
Encryption at rest documentation Medium Explicit documentation of which backups are encrypted
Database consistency Medium PostgreSQL/MongoDB backups should use pg_dump with consistent snapshots, not just file copies
Runbook testing Medium DR runbooks should be exercised, not just written
Cross-region replication Low S3 cross-region replication for geographic diversity
Automated documentation Low Backup inventory auto-generated from running config, not manually maintained

Single highest-risk gap:

Zero restore testing. Every script could produce corrupt archives and nobody would know until a real disaster. The backup plan itself states quarterly restore testing is required — it has never been done.


8. Risk Matrix

# Risk Likelihood Impact Urgency
R1 app3 MySQL has no backup since Aug 8 — 9 production databases at risk Certain (already happening) Critical (customer-facing WP sites, MainWP) IMMEDIATE
R2 git.modelortho.com Gitea cron never ran — Anita's repositories have no automated backup Certain (cron misconfigured) High (all Anita's repos and issues lost if disk fails) IMMEDIATE
R3 Prometheus has never been backed up — all monitoring history gone on disk failure Certain (never configured) High (lose all metrics, alerts, dashboards) Within 48hr
R4 Buzz/Bookstack/Support-API/etc. (10+ services) with zero backup Medium (disk failure) High (complete data loss for those services) Within 1 week
R5 No restore testing means all backups are unverified High (silent corruption) Critical (backups useless in real DR) Within 2 weeks
R6 Backup failure goes unnoticed — no alerting pipeline High (cron errors daily) High (data loss accumulates silently) Within 1 week
R7 MikroTik tower CCRs not backed up since Jul 7 Medium (routers stable) Medium (reconfig from scratch) Within 2 weeks
R8 app3 nginx configs backup failing since Aug 10 Certain (observed today) Medium (can rebuild, but slower) Within 1 week
R9 Wazuh backup undocumented and unanchored — could be lost in migration Low Medium (security event history lost) Within 1 month
R10 Open WebUI backup capturing potentially stale Docker data Medium (bug in docker cp caching) Medium (some chat history lost) Within 1 month
R11 No second media type — single S3 provider is single point of failure Very Low (Wasabi 11-nines) Critical (if Wasabi has outage/dataloss) Within 3 months

9. Recommendations

Priority 0 — FIX TODAY

# Action Effort Risk Addressed
1 Fix app3 MySQL backup — diagnose why mysqldump stopped after Aug 8 (likely MySQL auth, disk full, or script logic change). Re-run manually for Aug 910. 30 min R1
2 Fire gitea-modelortho-backup cron — the script exists and is correct. The cron has no Last run. Run it manually now, verify S3, then fix the cron schedule. 15 min R2
3 Verify app3 Nginx config backup — failed on Aug 10 per cron log. Diagnose and re-run. 15 min R8

Priority 1 — THIS WEEK

# Action Effort Risk Addressed
4 Add Prometheus backup — add a simple tar czf of TSDB to core-services-backup.sh or create standalone script. The data path exists, just needs to be included. 30 min R3
5 Build backup alerting — pipe backup-audit-check.sh output to Telegram/email. Add S3 file size check (flag 0-byte or <1KB files). 2 hr R6
6 Create backups for 10+ unbacked services — prioritize: buzz-prod-postgres, bookstack, support-api, timetrex. Create individual backup scripts. 4 hr R4
7 Restore test: 1 backup — pick any one backup (suggest: gitea/app2), restore to a temp location, verify data integrity. Document results. This proves the concept. 2 hr R5
8 Update backup plan — add Wazuh, fix modelortho script name (it's gitea-modelortho-backup.sh not modelortho-backup.sh), document that app servers run local cron not Core-orchestrated SSH. 1 hr Documentation

Priority 2 — THIS MONTH

# Action Effort Risk Addressed
9 Fix MikroTik tower backups — diagnose SSH timeout to 5 tower CCRs (VPN issue? IP change?). Get at least a weekly config export. 2 hr R7
10 Monthly restore test schedule — set up recurring calendar reminder. Test one random backup per month. 30 min setup R5
11 Add S3 backup integrity checks — for each service, after upload, re-download tarball and verify tar integrity or SQL dump validity. 3 hr R5
12 Investigate Open WebUI backup size freeze — check if docker cp is caching stale data; consider pg_dump approach instead. 2 hr R10
13 Add 14-day retention to all backup scripts — most scripts have no cleanup; S3 costs accrue indefinitely. 2 hr Cost

Priority 3 — THIS QUARTER

# Action Effort Risk Addressed
14 Add S3 Object Lock — enable compliance mode on Wasabi buckets to prevent ransomware deletion. 1 hr R11
15 Second media type — add local NAS backup for critical services, or replicate to a second cloud provider (Backblaze B2). 8 hr R11
16 Full DR simulation — schedule and execute a complete failover-to-standby exercise per the backup plan's annual requirement. 8 hr R5
17 Automated service discovery — script that scans all servers for running services and cross-references against backup coverage. 3 hr R4
18 Cross-server backup health dashboard — Uptime Kuma or Grafana dashboard showing last backup time and status for every service. 3 hr R6

Appendix A: Script Name Discrepancies

The backup plan lists script names that don't match reality:

Plan Says Actual Name Location
modelortho-backup.sh gitea-modelortho-backup.sh /root/.hermes/scripts/ on Core
app1-backup.sh (runs from Core via SSH) /root/backup.sh (runs locally on app1) app1 crontab
app2-backup.sh (runs from Core via SSH) /root/backup.sh (runs locally on app2) app2 crontab
app3-backup.sh (runs from Core via SSH) /root/backup.sh (runs locally on app3) app3 crontab
wphost02-backup.sh (runs from Core) ssh root@5.161.62.38 '/root/backup.sh' (Core crontab invokes wphost02's script) Both exist

Architecture reality: The backup plan describes a Core-orchestrated model where Core SSHs to each server and runs backup scripts. In reality, each app server has its own local cron job that runs /root/backup.sh independently. Core's system crontab only directly backs up: Hermes, Core services, root essentials, and wphost02 (via SSH). All other app backups are orchestrated by Hermes cron jobs that SSH from Core (vaultwarden, litellm, komodo, etc.) or by local crontabs on the app servers themselves.

Appendix B: S3 Bucket Health

Bucket Status Notes
hermes-vps-backups Active Primary backup bucket. 33 prefixes, daily uploads
mikrotik-ccr-backups ⚠️ Partial Home gateway configs OK; tower CCR configs missing since Jul 7

Appendix C: Hermes Cron Job Summary

Total Hermes cron jobs 28
Backup-related jobs 17
Jobs with "ok" status 23
Jobs with "error" status 2 (home-router-daily-backup, exotic-vehicle-scout)
Jobs with NO Last run (never executed) 1 (Gitea ModelOrtho Backup)
Jobs running in "no_agent" mode 25