Backup and recovery statement

Status: operative document. Version 1.1, 2026-09-28. Owner: Yannick M. Richard. Review cycle: annual, or on material change. This statement publishes the basis of control A.8.13 in the statement of applicability. The operational runbook it summarises stays in the private configuration repository, because it contains host-specific recovery steps where the order matters. Not a certification document.

1. What is backed up, and where it goes

The three production hosts — the public edge, the mail host and the management host — back up their own state to a dedicated Cloudflare R2 bucket over the S3 API using restic, one repository per host. The hosts that carry a database run a second, independent path:

What How often Where it lands
Host file state Every 15 minutes, with a short randomised delay The host’s restic repository in R2
PostgreSQL logical dump (pg_dumpall, gzipped, written atomically and read only once it is complete) Hourly; the runs inside the hour reuse it Inside the same restic snapshot
PostgreSQL physical backup and write-ahead log Daily, with continuous WAL archiving (bounded at 300 seconds) pgBackRest, in the same bucket under its own prefix

The configuration is public: modules/fleet-backup.nix for the file path and modules/pgbackrest.nix for the database path.

The two paths are not redundant copies of one another. A restic snapshot is a tree: it restores a host’s state as it was, including files nobody would otherwise remember to recreate. The pgBackRest repository is a point in time: it restores the database to any moment already inside the shipped WAL window — bounded by the 300-second archive interval, so the last few minutes may not be in the repository yet — which the tree cannot, because an hourly dump is only an hourly dump.

2. Encryption, and what the storage provider holds

Everything is encrypted on the host, before it is uploaded. Cloudflare holds ciphertext and object metadata, nothing else. It cannot read a snapshot, and it cannot produce one.

The consequence is worth stating plainly, because it is also why this document exists in its present shape: a backup only the archive can decrypt is not a backup. The passphrase that opens the repositories is therefore a secret in the store, deliberately shared across the fleet — any host must be able to restore any other host’s repository during recovery, which is precisely the situation where one host is gone. The backup does not contain the key that opens it: a host’s age identity is plaintext on disk, so a raw copy would sit in the archive readable by anything holding that shared passphrase. It is excluded and escrowed instead, and recovery from total host loss uses the escrow rather than the archive. That dependency chain is the subject of the key that decrypts everything.

3. What is deliberately not backed up

4. How long a restore point survives

Retention is bounded and short — the last snapshot of each period:

Period Window
Hour 24 hours
Day 14 days
Week 4 weeks
Month about a month

Nothing is kept indefinitely, and there is no configuration under which a snapshot survives on age alone.

The database path is bounded on the same principle and a shorter clock: pgBackRest keeps 7 full backups and 7 archive (WAL) periods, applied after each full, so the point-in-time window is about a week — the figure §6 quotes.

Two properties of the mechanism matter more than the numbers in that table:

5. The 2026-09-27 retention defect

The rules in §4 had been configured for months and were not doing what they said. restic’s default grouping is (host, paths), so a group is defined by the pair — and any change to a host’s backed-up path list starts a new group, in which the oldest snapshot is that group’s newest. Snapshots in a retired group are never pruned, because each of them is the most recent member of its own group. The rule was applied, ran clean, and expired nothing.

Measured on the management host’s repository on 2026-09-27, with identical keep rules and only the grouping changed:

Grouping Result
Default (host, paths) keep 90, remove 0
host keep 38, remove 52

Across the fleet, most snapshots sat in groups that could never expire, and one host’s repository had reached 14.6 GB while its newest snapshot was 0.5 GB. Three things followed, in order:

  1. Group by host — one group per repository, which is what makes the window in §4 the real window. The first run after the change removed 53, 81 and 61 snapshots on the three hosts.
  2. A one-time rewrite of the existing snapshots to drop the rebuildable caches from history (restic rewrite --exclude --forget), preserving every timestamp and every restore point: 36, 26 and 10 snapshots modified, zero errors. The bucket went from 32.39 GB to 12.63 GB.
  3. A bound on pruning (§4), after the first unbounded prune showed what an edge host does when a backup is allowed to care.

Step 2 is recorded because a retention policy that has never been tested against the data it governs is an assertion; what can be relied on is the measurement and the date. It is not part of routine operation, and nothing in §4 depends on it.

6. What recovery looks like, and what has been exercised

The runbook (private repository, docs/disaster-recovery.md) is: provision the instance, restore the age identity from escrow, apply the configuration, restore state from the repository, rejoin the tunnel, repoint DNS, verify. Its stated objectives are:

Objective Value
Recovery point — file state 15 minutes
Recovery point — database, via the logical dump 1 hour
Recovery point — database, via pgBackRest WAL about 5 minutes, inside a 7-day window
Recovery time — rebuild a host from configuration under an hour (a target, not a measurement)

Three of those four are enforced by the mechanism: the 15 minutes is the schedule, the hour is the dump’s refresh interval, the five minutes is the WAL archiving bound. The rebuild figure is an estimate from the runbook’s step count, and no full host rebuild has been performed. That is the significant gap in this document, and it is why the statement of applicability marks the continuity controls partial rather than claiming them.

File state, 2026-09-27, all three hosts

What has been performed:

Host Repository integrity Restore drill
Management restic check exit 0, no errors, 39 snapshots Restored /var/lib/wireguard-ui from the newest snapshot: 7 files; diff -r against the live tree clean
Public edge restic check exit 0, no errors, 39 snapshots Restored /var/lib/matrix-synapse from the newest snapshot: 829 files, 106 MB; diff -r clean
Mail restic check exit 0, no errors, 39 snapshots Restored fail2ban.sqlite3 (708 KiB) from the newest snapshot; differs only in ban entries written since it was taken

restic check verifies a repository’s structure, indexes and pack metadata. It does not re-read the content of every blob, and this document does not claim it does. The drill is a file-level restore: it shows that a repository can be opened with the fleet passphrase, that a snapshot reconstructs a real tree, and that the result matches the machine it came from. It does not show that a host can be rebuilt, which is the gap named above. The database path carries its own automated check — pgBackRest’s daily check runs against the repository — in addition to the daily backup.

The database path, 2026-09-28, both database hosts

Each host’s newest full backup was restored into a scratch directory beside the running cluster, WAL was replayed to a timestamp chosen before the run, and the recovered cluster was queried:

Host Backup set restored Restore Replay to target Assertion
Public edge 20260927-213048F 713.1 MB, 5,056 files in 6 min 20 s about 20 s the row written before the target is present, the change made after it is not
Mail 20260927-222614F 221.6 MB, 1,742 files in 1 min 58 s about 10 s the same

The point-in-time half is the half worth reading twice: the recovered cluster’s own log records it stopping one commit after the chosen timestamp, so this is recovery to a named moment rather than to the end of the archive. Recovery also had to fetch WAL from R2 as the database’s own recovery process — the same credential path the daily backup uses, exercised in the direction that matters, where a failure means the restore is complete and the database still will not open.

Two limits, stated so they are not assumed away. The restore lands beside a running host in a scratch path: it shows that the repository and the WAL chain reconstruct a working cluster, and it is still not a host rebuild (§8). And the drill is scripted (scripts/pg-pitr-drill.sh in the private repository) so it can be re-run; its dated output is recorded in that repository’s docs/dr-drill-evidence.md. The date is the evidence — the script only makes the next date cheap.

7. When a backup fails

A failed run leaves the unit in a failed state on the host, and a backup failure is classified S2 in the incident response plan: same-day response, published. There is no dedicated alert on backup failure — nothing pages when a run fails, and nothing pages when a timer is quietly disabled. Detection is the host’s unit state and the review cycle. That is a limitation of a one-operator fleet rather than a control, and it is recorded here rather than presented as coverage.

8. What this document does not claim

9. Review

Date Version Change
2026-09-27 1.0 Initial statement. Records the retention-grouping defect and its fix, the one-time historical rewrite, the cache exclusions, the bound on pruning, and the first file-level restore drill across the three hosts.
2026-09-28 1.1 Adds the database path’s own retention window (7 full backups and 7 WAL periods), the shipped-WAL bound on its point-in-time claim, and the physical restore and point-in-time drill performed on both database hosts with its measured durations. No claim retracted; no full host rebuild has been performed.
󰣨 ymrtech@ymrtech | 󰌠 NixOS | 󰍢 UTF-8