16 August 2026
We killed our platform on purpose: a disaster-recovery drill
Rebuilding an entire self-hosted platform from nothing but an offline secret list, two git repos and a bucket — measured end to end at 1 h 23 m 39 s.
disaster-recoverykubernetestalosgitopsbackupssre
Every backup strategy contains an untested assumption: that the pieces fit back together. Individually, each mechanism was proven — the database restores, the files restore, the secrets restore. What nobody had checked was whether a human with a stopwatch could assemble them into a working platform starting from an empty Hetzner account.
So we found out. New project, no cluster, no secrets in memory — just an offline list, two git repositories, and a bucket.
1 hour, 23 minutes, 39 seconds from "disaster declared" to apps serving traffic, against a four-hour target.
The last post established that each backup mechanism restores what it owns. This one asks the only question that actually matters.
What you'll get from this post
The reusable artifact first — the closed list of secrets that must live outside your platform, which is the thing to copy even if you never read another word. Then the measured rebuild: where the time actually went, and the four findings that would have turned a drill into a very bad day.
The list that has to live outside
Vaultwarden runs on the platform. OpenBao runs on the platform. Neither can be the anchor of its own recovery — you can't unlock a safe with the key inside it.
So there's a closed inventory of secrets kept in a personal password manager that is emphatically not hosted here. Nine items, and the goal is that this list plus the repos plus the bucket is sufficient to rebuild everything:
| # | Secret | Why it can't live on the platform |
|---|---|---|
| 1 | SOPS age private key | Decrypts the boundary secrets in git; its only in-cluster copy dies with the cluster |
| 2 | OpenBao unseal key + root token | A restored snapshot carries the production barrier — the backup is unreadable without these |
| 3 | Hosting account login + 2FA recovery codes | Recreate the project and mint an API token before any infrastructure exists |
| 4 | Object storage access keys | Reach the bucket before any infrastructure exists — and they can't be minted by automation |
| 5 | State encryption passphrase | The stored infrastructure state is client-side encrypted; without this it's ciphertext |
| 6 | Domain registrar login | DNS records are code, but the registration sits above the platform |
| 7 | Git host owner credentials + 2FA recovery | The repos are the platform |
| 8 | Vaultwarden master password(s) | Client-side encrypted — unrecoverable by design, no matter how good the backups are |
| 9 | Email/DNS provider account login | Outbound mail and the DNS zone |
Two things about this list are more important than its contents.
It records names and locations, never values. A value in this file is a review-blocker and an immediate rotation.
It stays small and closed, by rule. Anything not on it must be recoverable from the vault snapshots, the database backups, the file backups, or git — and if a new secret can't be, it joins this list in the same pull request that introduces it. Without that rule the inventory silently rots, which means you discover the missing row during the rebuild.
Item 5 is the proof the rule works. The state encryption passphrase wasn't in the original list; it got added because someone asked "could we actually decrypt the state?" The honest answer was no. Every other item was already known — that one was found by writing the procedure down.
Where the time actually went
| Phase | Time |
|---|---|
Substrate — new project to Ready cluster | 6 min |
| Bootstrap → dependent resources | 2 min |
| Vault — restore snapshot, unseal | 4 min |
| Database — recovery, including one reset | 13 min |
| File and manifest restore | 4 min |
| Verification and fixing friction live | 51 min |
Look at the last row. The restores took 29 minutes. Everything else took 51.
That ratio is the actual finding of the drill, and it's the one I'd expect to generalise to anybody's platform. The mechanisms are fast — they're automated, they've been tested individually, they do what they say. What consumes a recovery window is a human confirming that each layer really came back, and fixing the small things nobody knew were broken until the rebuild surfaced them.
Which means the useful question isn't "how fast is your restore?" It's "how long does it take you to believe your restore?"
Data loss, measured rather than promised: ≤ 5 minutes for the databases (continuous WAL archiving), 5 h 19 m for the secrets (the nightly vault snapshot, and the disaster was declared five hours after it ran), and zero for anything declarative, because git is the source of truth.
That secrets number is the honest weak spot. A nightly snapshot means up to 24 hours of exposure, and this drill happened to land at five. Whether that's acceptable depends on how often your secrets change — for this platform they change rarely, so it is. Say the number out loud rather than averaging it away.
Four findings that mattered
Patching only the root application is silently defeated. To point the
cluster at a rebuild branch, the obvious move is to change the root
app-of-apps. It reverts within one sync cycle, because every child application
pins its own revision and the root manages itself. The fix is a single
sed across the root plus every child — but the failure mode is what's
instructive: nothing errors, the change simply evaporates, and you conclude
you mistyped something.
Your server type may not exist any more. On drill day, both the production machine type and its obvious replacement were unorderable in the target region. Cloud providers retire plans, and you find out at the exact moment you need capacity. Checking availability is now step one of the runbook, before anything else happens.
The database refused to corrupt itself. When the cluster came up empty under the production name, the backup tool declined to write into the existing backup chain — a "foreign chain" refusal. That's a safety mechanism most people never see, and it's the thing standing between a fumbled recovery and silently destroying the backups you're recovering from. It fired here, correctly, and I'm glad it exists.
"Already exists" is what success looks like. Restoring into a cluster that GitOps already manages produced dozens of skip warnings and zero errors. If you don't know that in advance, a wall of warnings at minute 70 of a rebuild reads like the restore failed.
What this really buys you
Not confidence — a document.
The drill's output isn't the number in the headline; it's a runbook where every friction point that was hit has been folded back into the steps, so the next execution doesn't rediscover them. A recovery plan that has never been executed is a design document. Executing it once converts it into a procedure, and the conversion is the whole value.
The secondary output is knowing which promises are real. "Four hour RTO" was an aspiration written in a table. It's now a measured 84 minutes with a breakdown, which means it can be reasoned about, budgeted against, and told to a customer without crossing your fingers.
Gotchas
Every one of these cost real time.
- A rebuild branch must be applied to every application, not just the root. Child applications pin their own revision, and the self-managing root reverts partial changes within one sync cycle.
- Check machine-type availability before you start. Two candidate types were unorderable on the day. This belongs at step one, not at minute forty.
- A "restore immediately" flag fires once per object lifetime, not once per apply. After a delete-and-recreate cycle, delete the schedule object too or the immediate backup silently never happens.
- Apps that install plugins at runtime need a second restart. After the database was replaced, the file-sync app's configuration hook had to be re-triggered through a sync (deleting the job does nothing), and the web pod needed bouncing before its own routes stopped returning 404.
- Private CA trust may be honoured on one code path and ignored on another. Sign-in worked; session verification didn't. The fix was additive trust at the container level, not disabling verification — at hour two, the wrong fix is extremely tempting.
- Verify the list monthly, and record the date. A break-glass inventory nobody opens is a list of secrets that were correct once. A row that fails verification is a live incident, not a chore.
Where this leaves us
The platform can be destroyed and rebuilt in under ninety minutes from an offline list, two repositories and a bucket. That's the last thing standing between "a cluster I'm running" and "infrastructure I'd put someone else's business on."
Which is exactly what happens next. The remaining two posts stop building the platform and start using it: shipping a real SaaS onto it, and then charging money for a feature — with the entitlement state living in our own database rather than a payment provider's.
I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.