Backup and Recovery Statement
Backup and recovery statement
Status: operative document. Version 1.1, 2026-09-28. Owner: Yannick M. Richard. Review cycle: annual, or on material change. This statement publishes the basis of control A.8.13 in the statement of applicability. The operational runbook it summarises stays in the private configuration repository, because it contains host-specific recovery steps where the order matters. Not a certification document.
1. What is backed up, and where it goes
The three production hosts — the public edge, the mail host and the management host — back up their own state to a dedicated Cloudflare R2 bucket over the S3 API using restic, one repository per host. The hosts that carry a database run a second, independent path:
| What | How often | Where it lands |
|---|---|---|
| Host file state | Every 15 minutes, with a short randomised delay | The host’s restic repository in R2 |
PostgreSQL logical dump (pg_dumpall, gzipped, written atomically and read only once it is complete) |
Hourly; the runs inside the hour reuse it | Inside the same restic snapshot |
| PostgreSQL physical backup and write-ahead log | Daily, with continuous WAL archiving (bounded at 300 seconds) | pgBackRest, in the same bucket under its own prefix |
The configuration is public: modules/fleet-backup.nix for the
file path and modules/pgbackrest.nix for the database path.
The two paths are not redundant copies of one another. A restic snapshot is a tree: it restores a host’s state as it was, including files nobody would otherwise remember to recreate. The pgBackRest repository is a point in time: it restores the database to any moment already inside the shipped WAL window — bounded by the 300-second archive interval, so the last few minutes may not be in the repository yet — which the tree cannot, because an hourly dump is only an hourly dump.
2. Encryption, and what the storage provider holds
Everything is encrypted on the host, before it is uploaded. Cloudflare holds ciphertext and object metadata, nothing else. It cannot read a snapshot, and it cannot produce one.
The consequence is worth stating plainly, because it is also why this document exists in its present shape: a backup only the archive can decrypt is not a backup. The passphrase that opens the repositories is therefore a secret in the store, deliberately shared across the fleet — any host must be able to restore any other host’s repository during recovery, which is precisely the situation where one host is gone. The backup does not contain the key that opens it: a host’s age identity is plaintext on disk, so a raw copy would sit in the archive readable by anything holding that shared passphrase. It is excluded and escrowed instead, and recovery from total host loss uses the escrow rather than the archive. That dependency chain is the subject of the key that decrypts everything.
3. What is deliberately not backed up
- Customer traffic. The architecture does not store it, so there is no copy to back up and none to restore. See what we log.
- Raw PostgreSQL data directories. Version-fragile; the logical dump and the pgBackRest repository are the authoritative forms.
- Plaintext secret material. Escrowed separately, as above.
- Rebuildable caches. Go build,
uv, npm, Playwright browser and Nix evaluation caches are excluded, and so is the mail host’s Attic blob store. This is the largest single thing that used to be in the archive: on the management host the caches were 9.0 GB of a 13.7 GB uncompressed payload. What they cost to restore is a rebuild that a scheduled job does anyway; what they cost to keep is size, bandwidth and time on every run. The mail host’s blob store is 4 GB of nar archives that rebuild themselves on first request, and its metadata travels in the database dump, so a restored host knows what it is missing.
4. How long a restore point survives
Retention is bounded and short — the last snapshot of each period:
| Period | Window |
|---|---|
| Hour | 24 hours |
| Day | 14 days |
| Week | 4 weeks |
| Month | about a month |
Nothing is kept indefinitely, and there is no configuration under which a snapshot survives on age alone.
The database path is bounded on the same principle and a shorter clock: pgBackRest keeps 7 full backups and 7 archive (WAL) periods, applied after each full, so the point-in-time window is about a week — the figure §6 quotes.
Two properties of the mechanism matter more than the numbers in that table:
- Grouping by host. Snapshots are grouped by host alone. That reads like a detail; §5 is why it is not one.
- Bounded reclaim. A prune repacks at most 500 MiB per run. Unused pack objects are deleted regardless — the bound limits how much rewriting one run does, so a backlog cannot turn a maintenance job into an outage. The need for it was measured rather than assumed: the first post-fix prune on the mail host ran about 80 minutes, and sshd stopped answering banners for the last ten of them.
5. The 2026-09-27 retention defect
The rules in §4 had been configured for months and were not doing what they
said. restic’s default grouping is (host, paths), so a group is defined by
the pair — and any change to a host’s backed-up path list starts a new group,
in which the oldest snapshot is that group’s newest. Snapshots in a retired group
are never pruned, because each of them is the most recent member of its own
group. The rule was applied, ran clean, and expired nothing.
Measured on the management host’s repository on 2026-09-27, with identical keep rules and only the grouping changed:
| Grouping | Result |
|---|---|
Default (host, paths) |
keep 90, remove 0 |
host |
keep 38, remove 52 |
Across the fleet, most snapshots sat in groups that could never expire, and one host’s repository had reached 14.6 GB while its newest snapshot was 0.5 GB. Three things followed, in order:
- Group by host — one group per repository, which is what makes the window in §4 the real window. The first run after the change removed 53, 81 and 61 snapshots on the three hosts.
- A one-time rewrite of the existing snapshots to drop the rebuildable
caches from history (
restic rewrite --exclude --forget), preserving every timestamp and every restore point: 36, 26 and 10 snapshots modified, zero errors. The bucket went from 32.39 GB to 12.63 GB. - A bound on pruning (§4), after the first unbounded prune showed what an edge host does when a backup is allowed to care.
Step 2 is recorded because a retention policy that has never been tested against the data it governs is an assertion; what can be relied on is the measurement and the date. It is not part of routine operation, and nothing in §4 depends on it.
6. What recovery looks like, and what has been exercised
The runbook (private repository, docs/disaster-recovery.md) is: provision the
instance, restore the age identity from escrow, apply the configuration, restore
state from the repository, rejoin the tunnel, repoint DNS, verify. Its stated
objectives are:
| Objective | Value |
|---|---|
| Recovery point — file state | 15 minutes |
| Recovery point — database, via the logical dump | 1 hour |
| Recovery point — database, via pgBackRest WAL | about 5 minutes, inside a 7-day window |
| Recovery time — rebuild a host from configuration | under an hour (a target, not a measurement) |
Three of those four are enforced by the mechanism: the 15 minutes is the schedule, the hour is the dump’s refresh interval, the five minutes is the WAL archiving bound. The rebuild figure is an estimate from the runbook’s step count, and no full host rebuild has been performed. That is the significant gap in this document, and it is why the statement of applicability marks the continuity controls partial rather than claiming them.
File state, 2026-09-27, all three hosts
What has been performed:
| Host | Repository integrity | Restore drill |
|---|---|---|
| Management | restic check exit 0, no errors, 39 snapshots |
Restored /var/lib/wireguard-ui from the newest snapshot: 7 files; diff -r against the live tree clean |
| Public edge | restic check exit 0, no errors, 39 snapshots |
Restored /var/lib/matrix-synapse from the newest snapshot: 829 files, 106 MB; diff -r clean |
restic check exit 0, no errors, 39 snapshots |
Restored fail2ban.sqlite3 (708 KiB) from the newest snapshot; differs only in ban entries written since it was taken |
restic check verifies a repository’s structure, indexes and pack metadata. It
does not re-read the content of every blob, and this document does not claim it
does. The drill is a file-level restore: it shows that a repository can be
opened with the fleet passphrase, that a snapshot reconstructs a real tree, and
that the result matches the machine it came from. It does not show that a host
can be rebuilt, which is the gap named above. The database path carries its own
automated check — pgBackRest’s daily check runs against the repository — in
addition to the daily backup.
The database path, 2026-09-28, both database hosts
Each host’s newest full backup was restored into a scratch directory beside the running cluster, WAL was replayed to a timestamp chosen before the run, and the recovered cluster was queried:
| Host | Backup set restored | Restore | Replay to target | Assertion |
|---|---|---|---|---|
| Public edge | 20260927-213048F |
713.1 MB, 5,056 files in 6 min 20 s | about 20 s | the row written before the target is present, the change made after it is not |
20260927-222614F |
221.6 MB, 1,742 files in 1 min 58 s | about 10 s | the same |
The point-in-time half is the half worth reading twice: the recovered cluster’s own log records it stopping one commit after the chosen timestamp, so this is recovery to a named moment rather than to the end of the archive. Recovery also had to fetch WAL from R2 as the database’s own recovery process — the same credential path the daily backup uses, exercised in the direction that matters, where a failure means the restore is complete and the database still will not open.
Two limits, stated so they are not assumed away. The restore lands beside a running
host in a scratch path: it shows that the repository and the WAL chain reconstruct
a working cluster, and it is still not a host rebuild (§8). And the drill is
scripted (scripts/pg-pitr-drill.sh in the private repository) so it can be re-run;
its dated output is recorded in that repository’s docs/dr-drill-evidence.md. The
date is the evidence — the script only makes the next date cheap.
7. When a backup fails
A failed run leaves the unit in a failed state on the host, and a backup failure is classified S2 in the incident response plan: same-day response, published. There is no dedicated alert on backup failure — nothing pages when a run fails, and nothing pages when a timer is quietly disabled. Detection is the host’s unit state and the review cycle. That is a limitation of a one-operator fleet rather than a control, and it is recorded here rather than presented as coverage.
8. What this document does not claim
- No full disaster-recovery drill has been performed. The rebuild time is a target from a runbook.
- One storage provider. R2 is the only off-site copy. The second copy is the configuration and the escrow in git, which can rebuild a host but cannot restore its data. Recorded as residual risk R-08/R-09.
- The drills are not yet on a cycle. The standing rule is that each internal review includes one, and no review has been recorded yet; 2026-09-27 (file state, three hosts) and 2026-09-28 (the database path, both database hosts) are off-cycle exercises and the only dates this document can offer.
- Losing the passphrase loses the archive. It is in the secrets store, and the store is recoverable through the escrow, which makes the escrow the step that matters. The escrow is encrypted to the owner and the operator only.
- The operator’s workstation is out of scope. It is not part of the fleet backup path; this statement covers the three production hosts.
- The 12.63 GB figure is a dated measurement, not a standing one.
9. Review
| Date | Version | Change |
|---|---|---|
| 2026-09-27 | 1.0 | Initial statement. Records the retention-grouping defect and its fix, the one-time historical rewrite, the cache exclusions, the bound on pruning, and the first file-level restore drill across the three hosts. |
| 2026-09-28 | 1.1 | Adds the database path’s own retention window (7 full backups and 7 WAL periods), the shipped-WAL bound on its point-in-time claim, and the physical restore and point-in-time drill performed on both database hosts with its measured durations. No claim retracted; no full host rebuild has been performed. |