Sync docs, audit artifacts, project notes, and VerdictTank proposal docs

- audit/phase-one + phase-two: security audit briefs, findings, credential-rotation plan, Docker-USER hardening scripts, rollback refs
- disaster-recovery/restore-test-log.md + backup-dr-audit-2026-08-10.md
- clients/ (modelortho SEO audit, ai-biz-dev competitive landscape), notes/ (tiktok strategy)
- projects/: front-desk-voice-agent, seo-visibility-checker product plan, hotnow-savannah HTML, resend-transactional-email, backup-dashboard-enhancements, code-review-graph, seo-ci-architecture
- proposals/verdicttank/: architecture v4.0, methodology, judge-pool review, consolidation reasoning, cross-check review
- docs/super-search/firecrawl-provider-strategy.md
- updates: CHANGELOG, model-chain, projects-master-readme, intelsight.io
- .gitignore: exclude nested standalone repos (seo-tool, venturebuilt)
This commit is contained in:
root
2026-08-26 02:27:28 -04:00
parent 23e9751d38
commit f5175f1ce0
55 changed files with 14669 additions and 3 deletions
+423
View File
@@ -0,0 +1,423 @@
# ITPP Backup & DR — Full Audit Report
> **Date:** 2026-08-10 ~20:45 ET
> **Auditor:** Hermes Agent (automated audit)
> **Scope:** All ITPP infrastructure — Core, app1, app2, app3, wphost02, MikroTik CCR
> **Storage Verified:** Wasabi S3 (hermes-vps-backups, mikrotik-ccr-backups)
---
## Executive Summary
- **Overall grade: D** — Multiple active data-loss risks, no restore testing, several services completely unprotected.
- **Critical gaps** (items that risk data loss RIGHT NOW):
1. **app3 MySQL backups stopped 2026-08-08** — 9 databases (apextrackexperience, boxpilotlogistics, debtrecoveryexperts, iAmGMB, intelsight, katiewattsdesign, mainwp, vigilanttac, voipsimplicity_site) have no backup for Aug 910. Two days of production data completely unprotected.
2. **git.modelortho.com (Anita's Gitea) has a backup script and Hermes cron, but the cron has NEVER executed** — the cron job exists with no `Last run` timestamp. Only one manual backup exists on S3.
3. **Prometheus TSDB — ZERO backups ever**`core/prometheus/` S3 prefix is completely empty. All monitoring history, alert rules, and dashboards are unprotected.
4. **10+ production services have NO backup** — buzz-prod (Block Buzz relay + PostgreSQL + Redis + MinIO), support-api, bookstack (×2), searxng, timetrex, microbin, camofox-browser, browserless, msp-forms, docs-auth-validator, transitpin-api, transitpin-relay.
5. **ZERO evidence of any verified restore** — backups exist but nobody has ever tested if they can actually be restored.
- **High-priority improvements:**
1. Fix app3 MySQL backup immediately
2. Get gitea-modelortho cron running (it exists but hasn't fired)
3. Create backups for all unbacked services
4. Stand up Prometheus backup
5. Perform a restore test of at least one backup within 7 days
---
## 1. Backup Inventory & Verification
### Legend
- ✅ = Healthy (script exists, runs, S3 file present, size normal)
- ⚠️ = Warning (present but suspicious — unchanging size, stale data, missing days)
- 🔴 = Failed (missing entirely, script broken, zero files)
- 👻 = Phantom (script in plan but doesn't exist on disk)
- ❓ = Unknown (could not verify execution)
### Core (152.53.192.33)
| Target | Script | Script Exists? | S3 Files (7d) | Size Healthy? | Status | Notes |
|--------|--------|---------------|---------------|---------------|--------|-------|
| Hermes Full Backup | `hermes-backup.sh` (system cron 1AM) | ✅ Core | ✅ Aug 410 daily, 1.21.3 GB | ✅ Growing | ✅ | |
| Hermes Live Sync | `hermes-live-sync.sh` (Hermes cron, every 15m) | ✅ Core | ⚠️ Only `verification_evidence.db`, no state.db | ⚠️ Stale | ⚠️ | `live/` has old malformed backups from Jul 9; current sync may go elsewhere |
| /root Essentials | `root-essentials-backup.sh` (system cron 3AM) | ✅ Core | ✅ Aug 410 daily, 170236 MB | ✅ Growing | ✅ | |
| Grafana | `core-services-backup.sh` (system cron 1:30AM) | ✅ Core | ✅ Aug 410 daily, ~52 KB | ✅ Consistent | ✅ | |
| Uptime Kuma | `core-services-backup.sh` | ✅ Core | ✅ Aug 410 daily, ~100 MB | ✅ Growing | ✅ | |
| Docker Volumes | `core-services-backup.sh` | ✅ Core | ⚠️ vaultwarden-data only, stale since Jul 28 | ⚠️ Stale | ⚠️ | vaultwarden-data backed up but service migrated; no other volumes visible |
| Prometheus | `core-services-backup.sh` | ✅ Core | 🔴 **EMPTY** | 🔴 **Zero bytes** | 🔴 | `core/prometheus/` has NO files ever |
| Auth API | `auth-api-backup.sh` (Hermes cron 3:15AM) | ✅ Core | ✅ Aug 810 daily, ~41 KB | ✅ Consistent | ✅ | |
### App1 (152.53.36.131)
| Target | Script | Script Exists? | S3 Files (7d) | Size Healthy? | Status | Notes |
|--------|--------|---------------|---------------|---------------|--------|-------|
| Open WebUI | `/root/backup.sh` (app1 cron 2AM) | ✅ app1 | ⚠️ Aug 410 daily | ⚠️ **Identical size 5 days** (1,470,272,729 bytes Aug 59) | ⚠️ | Size finally changed Aug 10 (1,470,769,575). Docker cp may capture stale data |
| LiteLLM DB | `litellm-backup.sh` (Hermes cron 3:30AM) | ✅ Core | ✅ Aug 410 daily, 1521 MB, growing | ✅ Growing | ✅ | Also has config-only files from app1 backup.sh (240 bytes) |
| n8n | `/root/backup.sh` (app1 cron 2AM) | ✅ app1 | ✅ Aug 410 daily, ~6064 KB | ✅ Consistent | ✅ | |
| MCP Configs | `/root/backup.sh` (app1 cron 2AM) | ✅ app1 | ✅ Aug 410 daily, 373 bytes | ✅ Consistent | ✅ | Very small — configs only |
| Vaultwarden | `vaultwarden-backup.sh` (Hermes cron 2:30AM) | ✅ Core | ✅ Aug 410 daily, 520755 KB | ✅ Growing | ✅ | |
| Komodo | `komodo-backup.sh` (Hermes cron 3:45AM) | ✅ Core | ✅ Aug 410 daily, ~12.7 KB | ✅ Consistent | ✅ | |
| DocuSeal | `docuseal-backup.sh` (Hermes cron 4AM) | ✅ Core | ✅ Aug 410 daily, ~255 KB | ✅ Consistent | ✅ | |
| Twenty CRM | `twenty-backup.sh` (Hermes cron 4:15AM) | ✅ Core | ✅ Aug 310 daily, ~156 KB | ✅ Consistent | ✅ | Also has `twenty-files` tarballs from app1 backup.sh |
| Kokoro TTS | *(stateless)* | N/A | N/A | N/A | ✅ | |
### App2 (152.53.39.202)
| Target | Script | Script Exists? | S3 Files (7d) | Size Healthy? | Status | Notes |
|--------|--------|---------------|---------------|---------------|--------|-------|
| Traccar | `/root/backup.sh` (app2 cron 2:30AM) | ✅ app2 | ✅ Aug 410 daily, ~1.8 MB | ✅ Slightly growing | ✅ | |
| Gitea | `gitea-backup.sh` (Hermes cron 8AM) | ✅ Core | ✅ Aug 410 daily (folders w/ db + repos) | ✅ DB 2.7 MB, 30+ repos | ✅ | |
| Hudu | `hudu-backup.sh` (Hermes cron 7AM) | ✅ Core | ✅ Aug 410 daily, 670714 KB | ✅ Growing | ✅ | |
| UNMS | `unms-backup-sync.sh` (Hermes cron 6AM) | ✅ Core | ✅ Daily since Aug 5, ~100 MB | ✅ Consistent | ✅ | Gap Jul 31Aug 5 |
| UniFi | `unifi-backup-sync.sh` (Hermes cron 2AM) | ✅ Core | ✅ Aug 410 daily, ~884893 KB | ✅ Consistent | ✅ | |
| Technitium DNS | `technitium-backup.sh` (Hermes cron 2:45AM) | ✅ Core | ✅ Aug 610 daily, 369461 KB, growing | ✅ Growing | ✅ | Also has duplicates from app2 backup.sh |
| Dawarich | `dawarich-backup.sh` (Hermes cron 4AM) | ✅ Core | ✅ Aug 610 daily, ~7 MB | ✅ Consistent | ✅ | Also backed up by app2 backup.sh |
| RAGFlow | `ragflow-backup.sh` (Hermes cron 4:15AM) | ✅ Core | ✅ Aug 810 daily, ~267 KB | ⚠️ Unchanging | ⚠️ | Identical size all 3 days — very small for MySQL |
### App3 (152.53.241.111)
| Target | Script | Script Exists? | S3 Files (7d) | Size Healthy? | Status | Notes |
|--------|--------|---------------|---------------|---------------|--------|-------|
| CloudPanel DB | `/root/backup.sh` (app3 cron 3AM) | ✅ app3 | ✅ Aug 410 daily, 1.41.6 MB | ✅ Growing | ✅ | |
| MySQL (all DBs) | `/root/backup.sh` (app3 cron 3AM) | ✅ app3 | 🔴 **Last: Aug 8** | 🔴 **No Aug 910** | 🔴 | **CRITICAL**: 9 DBs stopped backing up after Aug 8. No mysql lines in Aug 10 log. |
| WordPress Files | `/root/backup.sh` (app3 cron 3AM) | ✅ app3 | ✅ Aug 9 daily, 9 sites | ✅ Normal | ✅ | `wp-www` identical size for 10 days is suspicious but other sites vary |
| Nginx Configs | `/root/backup.sh` (app3 cron 3AM) | ✅ app3 | ✅ Aug 48 daily, ~12 MB | ✅ Growing | ⚠️ | **FAILED on Aug 10** per cron log: "Nginx configs: FAILED" |
| Static Sites | `/root/backup.sh` (app3 cron 3AM) | ✅ app3 | ✅ Aug 910 daily, 17 sites | ✅ Normal | ✅ | Includes modelortho.com, buzz.iamgmb.com, transitpin.com, etc. |
| WP Snapshots (local) | `/opt/backup-restore/snapshot.sh` (app3 cron 1AM/1PM) | ✅ app3 | ❓ Local only | ❓ Not verified | ❓ | Local snapshots, not in S3 |
| Hexclave (Stack Auth) | `hexclave-backup.sh` (Hermes cron 3:30AM) | ✅ Core | ✅ Aug 810 daily, ~315 KB | ✅ Consistent | ✅ | |
| modelortho.com | `modelortho-backup.sh` (listed in plan) | 👻 **MISSING** | N/A | N/A | 👻 | **Does not exist**. However static modelortho.com IS backed up via app3/static. |
| git.modelortho.com (Anita's Gitea) | `gitea-modelortho-backup.sh` (Hermes cron 4:30AM) | ✅ Core | 🔴 **Only 1 file** (manual?) | 🔴 **Cron never ran** | 🔴 | **CRITICAL**: Script exists, cron exists, but cron has NO `Last run`. Zero automated backups. |
### wphost02 (5.161.62.38)
| Target | Script | Script Exists? | S3 Files (7d) | Size Healthy? | Status | Notes |
|--------|--------|---------------|---------------|---------------|--------|-------|
| WordPress (7 sites) | `/root/backup.sh` (Core cron 5AM via SSH) | ✅ wphost02 | ✅ Aug 410 daily | ✅ Normal (50900 MB per site) | ✅ | 7 sites + all-databases.sql + RunCloud config |
### MikroTik CCR
| Target | Script | Script Exists? | S3 Files (7d) | Size Healthy? | Status | Notes |
|--------|--------|---------------|---------------|---------------|--------|-------|
| Home Gateway | `run-wisp-backup.sh` (Hermes cron 6AM) | ✅ Core | ✅ Daily config .rsc | ✅ Normal | ⚠️ | Home gateway OK, but **5 tower CCRs all timeout** (T01T04, MP100) |
| WISP Tower CCRs | Same script | ✅ Core | 🔴 **None since Jul 7** | 🔴 | 🔴 | SSH timeout to all towers. Only 1 historical config from Jul 7. |
### Hetzner Snapshots
| Target | Script | Script Exists? | S3 Files (7d) | Status | Notes |
|--------|--------|---------------|---------------|--------|-------|
| Weekly disk snapshots | `snapshot-hetzner.py` (Hermes cron Mon 5AM) | ✅ Core | N/A (Hetzner API, not S3) | ✅ | Last run Aug 10, ok |
---
## 2. Missing Backups
The backup plan's "Unbacked Services" section claims "(none)" — this is **incorrect**. The following services are running in production with ZERO backup coverage:
### Core (152.53.192.33)
| Service | Type | Data at Risk | Priority |
|---------|------|-------------|----------|
| searxng | Docker (search engine) | Configuration, any cached indexes | Low |
| timetrex | Docker (time tracking) | Employee time data, payroll records | **High** |
| microbin | Docker (pastebin) | Shared text snippets | Low |
| camofox-browser | Docker (browser automation) | Stateless? | Low |
| browserless | Docker (headless browser) | Stateless | Low |
| Prometheus | systemd (monitoring TSDB) | **All metrics history, alert rules, dashboards** | **High** |
### App1 (152.53.36.131)
| Service | Type | Data at Risk | Priority | Notes |
|---------|------|-------------|----------|-------|
| Wazuh (SIEM/XDR) | Docker (×3 containers) | Security events, agent configs, alerts | **High** | Has `wazuh-backup.sh` on app1 and daily S3 files, but NOT in backup plan or Hermes cron. **Backup works but is undocumented/unanchored.** |
### App2 (152.53.39.202)
| Service | Type | Data at Risk | Priority |
|---------|------|-------------|----------|
| support-api | Docker (support API) | Support ticket data | **High** |
| bookstack | Docker (wiki) | Documentation wiki content | **High** |
| bookstack-db | Docker (MariaDB for bookstack) | Wiki database | **High** |
| happy_rosalind | Docker (bookstack duplicate?) | Unknown content | Medium |
### App3 (152.53.241.111)
| Service | Type | Data at Risk | Priority |
|---------|------|-------------|----------|
| buzz-prod-relay | Docker (Block Buzz relay) | Relay configuration, keys | **High** |
| buzz-prod-postgres | Docker (Buzz PostgreSQL) | **All Buzz relay state** | **Critical** |
| buzz-prod-redis | Docker (Buzz Redis) | Session/cache data | Medium |
| buzz-prod-minio | Docker (Buzz object storage) | Uploaded files | Medium |
| transitpin-api | systemd | Transportation API data | Medium |
| transitpin-relay | systemd | Relay config | Medium |
| msp-forms | systemd | Form submissions | Medium |
| docs-auth-validator | systemd | Auth validation state | Low |
---
## 3. 3-2-1 Compliance
The 3-2-1 rule states: **3 copies** of data, on **2 different media**, with **1 off-site**.
| Service | 3 Copies? | 2 Media? | 1 Off-Site? | Verdict |
|---------|-----------|----------|-------------|---------|
| Hermes Full Backup | ✅ (live + daily S3 + standby) | ⚠️ (all Wasabi S3) | ✅ (Wasabi us-east-1) | **PARTIAL** — single media type |
| App1 services (LiteLLM, n8n, Vaultwarden, etc.) | ✅ (server + S3) | ⚠️ (all Wasabi S3) | ✅ | **PARTIAL** |
| App2 services (Gitea, Hudu, Traccar, etc.) | ✅ (server + S3) | ⚠️ (all Wasabi S3) | ✅ | **PARTIAL** |
| App3 MySQL | 🔴 (server only, no S3 since Aug 8) | 🔴 | 🔴 | **FAIL** |
| App3 WordPress | ✅ (server + S3 + local snapshots) | ✅ (S3 + local disk) | ✅ | **PASS** |
| App3 static sites | ✅ (server + S3) | ⚠️ (all Wasabi S3) | ✅ | **PARTIAL** |
| wphost02 WordPress | ✅ (server + S3) | ⚠️ (all Wasabi S3) | ✅ | **PARTIAL** |
| git.modelortho.com | 🔴 (only server, cron never ran) | 🔴 | 🔴 | **FAIL** |
| Prometheus | 🔴 (server only) | 🔴 | 🔴 | **FAIL** |
| Wazuh | ✅ (server + S3) | ⚠️ (all Wasabi S3) | ✅ | **PARTIAL** |
| Buzz / Bookstack / Support-API / etc. | 🔴 (server only) | 🔴 | 🔴 | **FAIL** |
| MikroTik Home Gateway | ✅ (router + S3) | ⚠️ (all Wasabi S3) | ✅ | **PARTIAL** |
| MikroTik Tower CCRs | 🔴 (router only) | 🔴 | 🔴 | **FAIL** |
**Overall 3-2-1 compliance: FAIL**
The strategy relies entirely on Wasabi S3 for off-site storage — there is no second media type (no local NAS, no tape, no separate cloud provider). Every service that passes does so only by counting the production server + S3 as two copies. For true 2-media compliance, a second storage type (e.g., local NAS, separate cloud provider, or physical media) would be required.
---
## 4. RPO/RTO Assessment
| Service | Tier | Planned RPO | Current RPO | Acceptable? | Planned RTO | Current RTO Est. | Acceptable? |
|---------|------|-------------|-------------|-------------|-------------|-----------------|-------------|
| Hermes Agent | Critical | ≤1hr | ~15 min (live sync) | ✅ | ≤4hr | ~24hr | ✅ |
| Gitea (app2) | Critical | ≤1hr | 24hr (daily only) | ⚠️ | ≤4hr | ~24hr | ✅ |
| Traccar | Critical | ≤1hr | 24hr (daily only) | ⚠️ | ≤4hr | ~24hr | ✅ |
| UISP (UNMS) | Critical | ≤1hr | 24hr | ⚠️ | ≤4hr | ~48hr | ⚠️ |
| LiteLLM | High | 24hr | 24hr | ✅ | ≤8hr | ~48hr | ✅ |
| n8n | High | 24hr | 24hr | ✅ | ≤8hr | ~48hr | ✅ |
| Open WebUI | High | 24hr | 24hr | ✅ | ≤8hr | ~48hr | ✅ |
| Vaultwarden | High | 24hr | 24hr | ✅ | ≤8hr | ~24hr | ✅ |
| Twenty CRM | High | 24hr | 24hr | ✅ | ≤8hr | ~48hr | ✅ |
| Hudu | Medium | 24hr | 24hr | ✅ | ≤24hr | ~824hr | ✅ |
| UniFi | Medium | 24hr | 24hr | ✅ | ≤24hr | ~48hr | ✅ |
| Komodo | Medium | 24hr | 24hr | ✅ | ≤24hr | ~24hr | ✅ |
| DocuSeal | Medium | 24hr | 24hr | ✅ | ≤24hr | ~24hr | ✅ |
| App3 WP sites | Medium | 24hr | 24hr | ✅ | ≤24hr | ~824hr | ✅ |
| Auth API | Medium | 24hr | 24hr | ✅ | ≤24hr | ~24hr | ✅ |
| Hexclave (Stack Auth) | Medium | 24hr | 24hr | ✅ | ≤24hr | ~48hr | ✅ |
| Grafana | Low | 24hr | 24hr | ✅ | ≤48hr | ~24hr | ✅ |
| Uptime Kuma | Low | 24hr | 24hr | ✅ | ≤48hr | ~24hr | ✅ |
| Prometheus | Low | 24hr | 🔴 **∞ (no backup)** | 🔴 | ≤48hr | 🔴 **Impossible** | 🔴 |
| MikroTik CCR | Low | 24hr | 24hr (home) / 🔴 **∞ (towers)** | 🔴 | ≤48hr | ~48hr (home) / 🔴 **Impossible** (towers) | 🔴 |
| Technitium DNS | Low | 24hr | 24hr | ✅ | ≤48hr | ~24hr | ✅ |
| Dawarich | Low | 24hr | 24hr | ✅ | ≤48hr | ~48hr | ✅ |
| RAGFlow | Low | 24hr | 24hr | ⚠️ | ≤48hr | ~48hr | ✅ |
| **Buzz / Bookstack / Support-API / etc.** | **Unclassified** | N/A | 🔴 **∞ (no backup)** | 🔴 | N/A | 🔴 **Impossible** | 🔴 |
| **git.modelortho.com** | **Unclassified** | N/A | 🔴 **∞ (cron never ran)** | 🔴 | N/A | 🔴 **Possible** (script exists) | 🔴 |
| **app3 MySQL** | **Medium** | 24hr | 🔴 **~48hr and growing** | 🔴 | ≤24hr | ⚠️ **Partial** (Aug 8 dump) | ⚠️ |
**RPO/RTO Assessment: FAIL for Critical tier** — Gitea, Traccar, and UISP are planned for ≤1hr RPO but receive only daily backups. The plan itself overstates capabilities.
---
## 5. Restore Testing
### Evidence of Verified Restores: **NONE**
| Source | What it says | Evidence of execution |
|--------|-------------|----------------------|
| `backup-plan.md` §Restore Testing Cadence | "Quarterly: Pick one random backup per tier, restore to staging, verify integrity" | 🔴 No test logs anywhere |
| `backup-plan.md` | "Annual: Full DR simulation" | 🔴 Never executed |
| `4-DR-Testing-Schedule.md` | "Every restore test gets logged with date, tester, duration, findings" | 🔴 No log file exists |
| `backup-policy.md` | "Monthly restore test of a randomly selected asset" | 🔴 No evidence |
| `3-Per-Server-Runbooks.md` | Detailed per-server restore procedures | ✅ Procedures exist on paper only |
| `dr-issue-log.md` | Historical DR audit issues | ✅ Issues tracked; no restore tests documented |
| `/root/.hermes/references/` | Search for "restore test", "verified restore", "DR simulation" | 🔴 Zero actual test results found |
**Conclusion:** Restore procedures are well-documented on paper, but **no backup has ever been restore-tested**. There are no test logs, no verification records, and no evidence that any of the 30+ backup targets can actually be restored successfully.
**Restore test cadence last executed:** NEVER
---
## 6. Monitoring & Alerting
### What exists:
| Check | Mechanism | Status |
|-------|-----------|--------|
| Hermes full backup age check | `backup-audit-check.sh` (system cron 2AM) | ✅ Runs daily, checks if last backup <36hr old |
| Cron output logging | All backup cron jobs pipe to `logger` | ✅ Output captured in syslog |
| Hermes cron job status | `hermes cron list` shows last run status | ✅ Most show "ok" |
| Uptime Kuma monitoring | External monitoring of services | ✅ |
| Ops data collector | `ops-data-collector.py` (every 5min) | ✅ Collects metrics |
### What's MISSING:
| Gap | Impact | Urgency |
|-----|--------|---------|
| **No alert on backup failure** | When `backup-audit-check.sh` fails, output goes to `logger` only. Nobody gets notified. | **Critical** |
| **No alert on 0-byte backup** | Several backups produce tiny config files (240 bytes) that could silently become 0 bytes | **High** |
| **No alert on backup not running** | Hermes cron jobs silently enter "error" state (home-router-daily-backup has been error for days) | **Critical** |
| **No cross-server backup health dashboard** | Must SSH to each server individually to check backup status | **High** |
| **No S3 file integrity validation** | Nobody checks if backup files are actually valid (corrupt tar, truncated SQL) | **High** |
| **Core crontab doesn't monitor remote server backup outcomes** | Core only checks its own full backup — app1/app2/app3 backup failures go unnoticed | **Critical** |
| **No automatic ticket/issue creation** | Backup failures should auto-create a GitHub issue or Hudu ticket | Medium |
**The `backup-audit-check.sh` is the only monitoring, and it only checks Hermes full backup freshness. Everything else is silent.**
### Cron Job Statuses:
| Job | Schedule | Latest Status | Issues |
|-----|----------|---------------|--------|
| hermes-backup (system cron) | 1 AM | ✅ Ran today | |
| backup-audit-check (system cron) | 2 AM | ✅ OK (20h ago) | |
| root-essentials-backup (system cron) | 3 AM | ✅ S3 files present | |
| core-services-backup (system cron) | 1:30 AM | ✅ S3 files present | Prometheus part empty |
| wphost02-backup (system cron via SSH) | 5 AM | ✅ S3 files present | |
| vaultwarden-backup (Hermes cron) | 2:30 AM | ✅ ok | |
| litellm-backup (Hermes cron) | 3:30 AM | ✅ ok | |
| komodo-backup (Hermes cron) | 3:45 AM | ✅ ok | |
| docuseal-backup (Hermes cron) | 4:00 AM | ✅ ok | |
| twenty-backup (Hermes cron) | 4:15 AM | ✅ ok | |
| auth-api-backup (Hermes cron) | 3:15 AM | ✅ ok | |
| stack-auth-backup (Hermes cron) | 3:15 AM | ✅ ok | |
| hexclave-backup (Hermes cron) | 3:30 AM | ✅ ok | |
| technitium-backup (Hermes cron) | 2:45 AM | ✅ ok | |
| dawarich-backup (Hermes cron) | 4:00 AM | ✅ ok | |
| ragflow-backup (Hermes cron) | 4:15 AM | ✅ ok | |
| hudu-backup (Hermes cron) | 7:00 AM | ✅ ok | |
| gitea-backup (Hermes cron) | 8:00 AM | ✅ ok | |
| unms-backup-sync (Hermes cron) | 6:00 AM | ✅ ok | |
| unifi-backup-sync (Hermes cron) | 2:00 AM | ✅ ok | |
| hetzner-weekly-snapshots (Hermes cron) | Mon 5 AM | ✅ ok | |
| home-router-daily-backup (Hermes cron) | 6:00 AM | 🔴 **error** | 5/6 towers timeout daily |
| **Gitea ModelOrtho Backup** (Hermes cron) | 4:30 AM | 🔴 **NEVER RAN** | No Last run field |
| app1 `/root/backup.sh` (app1 crontab) | 2:00 AM | ❓ Not verified | S3 files present but openwebui sizes suspicious |
| app2 `/root/backup.sh` (app2 crontab) | 2:30 AM | ❓ Not verified | S3 files present |
| app3 `/root/backup.sh` (app3 crontab) | 3:00 AM | ⚠️ MySQL failed | **MySQL dumps stopped Aug 8; nginx failed Aug 10** |
| app3 `/opt/backup-restore/snapshot.sh` (app3 crontab) | 1AM/1PM | ❓ Not verified | Local-only |
---
## 7. Industry Standard Gaps
### What we're doing right:
-**Comprehensive backup plan** — well-documented in backup-plan.md with 27 declared targets
-**24+ Hermes cron jobs** operating backup scripts — good automation
-**Daily cadence** for all critical services
-**Wasabi S3** as immutable off-site storage (11-nines durability)
-**Provider diversity** — Core on netcup, standby on Hetzner
-**Warm standby** for Hermes (app1-bu with 7-min failover)
-**DR runbooks** documented per server
-**DR issue log** maintained with root cause analysis
-**Hetzner weekly snapshots** as additional layer
### What's missing (every competent shop has these):
| Gap | Severity | Industry Expectation |
|-----|----------|---------------------|
| **Restore testing** | Critical | Monthly restore tests are table stakes. "Untested backups are not backups." |
| **Backup failure alerting** | Critical | PagerDuty/OpsGenie/Telegram alert on ANY backup job failure |
| **Second media type** | High | NAS, tape, or second cloud provider for 3-2-1 compliance |
| **Immutable backups** | High | S3 Object Lock to prevent ransomware deletion |
| **Backup integrity validation** | High | Automated restore-and-verify pipeline (spin up temp container, restore DB, run queries) |
| **Service discovery for backups** | High | Automated scan that identifies running services and flags unbacked ones |
| **Backup retention policy** | Medium | Defined retention periods per tier (only gitea-modelortho and wphost02 have cleanup) |
| **Encryption at rest documentation** | Medium | Explicit documentation of which backups are encrypted |
| **Database consistency** | Medium | PostgreSQL/MongoDB backups should use pg_dump with consistent snapshots, not just file copies |
| **Runbook testing** | Medium | DR runbooks should be exercised, not just written |
| **Cross-region replication** | Low | S3 cross-region replication for geographic diversity |
| **Automated documentation** | Low | Backup inventory auto-generated from running config, not manually maintained |
### Single highest-risk gap:
**Zero restore testing.** Every script could produce corrupt archives and nobody would know until a real disaster. The backup plan itself states quarterly restore testing is required — it has never been done.
---
## 8. Risk Matrix
| # | Risk | Likelihood | Impact | Urgency |
|---|------|-----------|--------|---------|
| R1 | app3 MySQL has no backup since Aug 8 — 9 production databases at risk | **Certain** (already happening) | **Critical** (customer-facing WP sites, MainWP) | **IMMEDIATE** |
| R2 | git.modelortho.com Gitea cron never ran — Anita's repositories have no automated backup | **Certain** (cron misconfigured) | **High** (all Anita's repos and issues lost if disk fails) | **IMMEDIATE** |
| R3 | Prometheus has never been backed up — all monitoring history gone on disk failure | **Certain** (never configured) | **High** (lose all metrics, alerts, dashboards) | **Within 48hr** |
| R4 | Buzz/Bookstack/Support-API/etc. (10+ services) with zero backup | **Medium** (disk failure) | **High** (complete data loss for those services) | **Within 1 week** |
| R5 | No restore testing means all backups are unverified | **High** (silent corruption) | **Critical** (backups useless in real DR) | **Within 2 weeks** |
| R6 | Backup failure goes unnoticed — no alerting pipeline | **High** (cron errors daily) | **High** (data loss accumulates silently) | **Within 1 week** |
| R7 | MikroTik tower CCRs not backed up since Jul 7 | **Medium** (routers stable) | **Medium** (reconfig from scratch) | **Within 2 weeks** |
| R8 | app3 nginx configs backup failing since Aug 10 | **Certain** (observed today) | **Medium** (can rebuild, but slower) | **Within 1 week** |
| R9 | Wazuh backup undocumented and unanchored — could be lost in migration | **Low** | **Medium** (security event history lost) | **Within 1 month** |
| R10 | Open WebUI backup capturing potentially stale Docker data | **Medium** (bug in docker cp caching) | **Medium** (some chat history lost) | **Within 1 month** |
| R11 | No second media type — single S3 provider is single point of failure | **Very Low** (Wasabi 11-nines) | **Critical** (if Wasabi has outage/dataloss) | **Within 3 months** |
---
## 9. Recommendations
### Priority 0 — FIX TODAY
| # | Action | Effort | Risk Addressed |
|---|--------|--------|---------------|
| 1 | **Fix app3 MySQL backup** — diagnose why mysqldump stopped after Aug 8 (likely MySQL auth, disk full, or script logic change). Re-run manually for Aug 910. | 30 min | R1 |
| 2 | **Fire gitea-modelortho-backup cron** — the script exists and is correct. The cron has no Last run. Run it manually now, verify S3, then fix the cron schedule. | 15 min | R2 |
| 3 | **Verify app3 Nginx config backup** — failed on Aug 10 per cron log. Diagnose and re-run. | 15 min | R8 |
### Priority 1 — THIS WEEK
| # | Action | Effort | Risk Addressed |
|---|--------|--------|---------------|
| 4 | **Add Prometheus backup** — add a simple `tar czf` of TSDB to `core-services-backup.sh` or create standalone script. The data path exists, just needs to be included. | 30 min | R3 |
| 5 | **Build backup alerting** — pipe `backup-audit-check.sh` output to Telegram/email. Add S3 file size check (flag 0-byte or <1KB files). | 2 hr | R6 |
| 6 | **Create backups for 10+ unbacked services** — prioritize: buzz-prod-postgres, bookstack, support-api, timetrex. Create individual backup scripts. | 4 hr | R4 |
| 7 | **Restore test: 1 backup** — pick any one backup (suggest: gitea/app2), restore to a temp location, verify data integrity. Document results. This proves the concept. | 2 hr | R5 |
| 8 | **Update backup plan** — add Wazuh, fix modelortho script name (it's `gitea-modelortho-backup.sh` not `modelortho-backup.sh`), document that app servers run local cron not Core-orchestrated SSH. | 1 hr | Documentation |
### Priority 2 — THIS MONTH
| # | Action | Effort | Risk Addressed |
|---|--------|--------|---------------|
| 9 | **Fix MikroTik tower backups** — diagnose SSH timeout to 5 tower CCRs (VPN issue? IP change?). Get at least a weekly config export. | 2 hr | R7 |
| 10 | **Monthly restore test schedule** — set up recurring calendar reminder. Test one random backup per month. | 30 min setup | R5 |
| 11 | **Add S3 backup integrity checks** — for each service, after upload, re-download tarball and verify tar integrity or SQL dump validity. | 3 hr | R5 |
| 12 | **Investigate Open WebUI backup size freeze** — check if docker cp is caching stale data; consider pg_dump approach instead. | 2 hr | R10 |
| 13 | **Add 14-day retention to all backup scripts** — most scripts have no cleanup; S3 costs accrue indefinitely. | 2 hr | Cost |
### Priority 3 — THIS QUARTER
| # | Action | Effort | Risk Addressed |
|---|--------|--------|---------------|
| 14 | **Add S3 Object Lock** — enable compliance mode on Wasabi buckets to prevent ransomware deletion. | 1 hr | R11 |
| 15 | **Second media type** — add local NAS backup for critical services, or replicate to a second cloud provider (Backblaze B2). | 8 hr | R11 |
| 16 | **Full DR simulation** — schedule and execute a complete failover-to-standby exercise per the backup plan's annual requirement. | 8 hr | R5 |
| 17 | **Automated service discovery** — script that scans all servers for running services and cross-references against backup coverage. | 3 hr | R4 |
| 18 | **Cross-server backup health dashboard** — Uptime Kuma or Grafana dashboard showing last backup time and status for every service. | 3 hr | R6 |
---
## Appendix A: Script Name Discrepancies
The backup plan lists script names that don't match reality:
| Plan Says | Actual Name | Location |
|-----------|-------------|----------|
| `modelortho-backup.sh` | `gitea-modelortho-backup.sh` | `/root/.hermes/scripts/` on Core |
| `app1-backup.sh` (runs from Core via SSH) | `/root/backup.sh` (runs locally on app1) | app1 crontab |
| `app2-backup.sh` (runs from Core via SSH) | `/root/backup.sh` (runs locally on app2) | app2 crontab |
| `app3-backup.sh` (runs from Core via SSH) | `/root/backup.sh` (runs locally on app3) | app3 crontab |
| `wphost02-backup.sh` (runs from Core) | `ssh root@5.161.62.38 '/root/backup.sh'` (Core crontab invokes wphost02's script) | Both exist |
**Architecture reality:** The backup plan describes a Core-orchestrated model where Core SSHs to each server and runs backup scripts. In reality, each app server has its own local cron job that runs `/root/backup.sh` independently. Core's system crontab only directly backs up: Hermes, Core services, root essentials, and wphost02 (via SSH). All other app backups are orchestrated by Hermes cron jobs that SSH from Core (vaultwarden, litellm, komodo, etc.) or by local crontabs on the app servers themselves.
## Appendix B: S3 Bucket Health
| Bucket | Status | Notes |
|--------|--------|-------|
| hermes-vps-backups | ✅ Active | Primary backup bucket. 33 prefixes, daily uploads |
| mikrotik-ccr-backups | ⚠️ Partial | Home gateway configs OK; tower CCR configs missing since Jul 7 |
## Appendix C: Hermes Cron Job Summary
| Total Hermes cron jobs | 28 |
|------------------------|-----|
| Backup-related jobs | 17 |
| Jobs with "ok" status | 23 |
| Jobs with "error" status | 2 (home-router-daily-backup, exotic-vehicle-scout) |
| Jobs with NO `Last run` (never executed) | 1 (Gitea ModelOrtho Backup) |
| Jobs running in "no_agent" mode | 25 |