12 August 2026
Back up everything, then prove it: Velero + restore day
Four backup mechanisms, one bucket, and a table of what each one does not cover — plus a restore day with real timings and the storage bug that made file backups silently impossible.
velerobackupskubernetestalosdisaster-recoveryhetzner
The most dangerous backup is the one that runs successfully every night and cannot be restored.
It's dangerous specifically because it generates evidence of safety. Green ticks, growing bucket, a job that exits zero. Every signal you'd think to check says you're covered, and none of them are the thing you actually need, which is a restore.
The last post put company files and a password vault on this platform — the first data whose loss would genuinely hurt. So this post is where the project pays that debt: what backs up what, what each mechanism explicitly does not cover, and a restore day with real numbers, including the one that took six minutes instead of the fifty-four seconds I expected.
What you'll get from this post
The complete backup responsibility split for a small self-hosted platform — four mechanisms, one bucket, and the boundaries between them — plus the storage-layer bug that made file backups impossible for a month while reporting success, and what an actual restore drill measured.
Four mechanisms, and what each one refuses to do
The table matters more than any individual tool, because almost every backup disaster is a boundary problem: two mechanisms that each assumed the other had it.
| Data | Mechanism | Does not cover |
|---|---|---|
| All Postgres databases | CNPG + Barman (WAL + nightly base, PITR) | anything outside Postgres |
| PVC file data + manifests | Velero + node-agent file backup (kopia), nightly | CNPG data volumes — deliberately excluded |
| All runtime secrets | OpenBao raft snapshots, nightly | anything not in the vault |
| Everything declarative | git itself | runtime data of any kind |
The row that gets people is the second one. Velero never touches the
Postgres volumes, and that exclusion is structural rather than a namespace
somebody forgot: the platform-pg namespace isn't in the schedule, and volume
backup is opt-in by annotation, so a new database would have to be actively
opted in to break the rule.
The reason is worth stating plainly, because "restore the PVC" is the obvious instinct: a file-level copy of a running database's data directory is not a backup, it's a torn snapshot. Restoring one gives you a database that starts, passes a smoke test, and is subtly corrupt. Barman and PITR are the only correct path back for Postgres.
Everything lands in one bucket under different prefixes, which is a deliberate and slightly contrarian call. Object storage charges per bucket, and the credentials are project-wide anyway, so a second bucket buys separation that the credential model doesn't actually deliver — while definitely buying another line item. A second bucket needs a written decision, not a convenience.
And the schedules are staggered: Postgres at 01:15, files at 02:15, the vault at 03:15. One box, one uplink; three simultaneous uploads is a self-inflicted incident.
App-consistent, or it doesn't count
Copying Nextcloud's files while somebody is uploading gives you a backup of a half-written file and a database that disagrees with it. So the backup brackets itself with maintenance mode:
pre:
- exec:
command: [..., "occ maintenance:mode --on"]
onError: Fail
post:
- exec:
command: [..., "occ maintenance:mode --off"]
onError: Continue
The asymmetry is intentional. If the app can't be quiesced, fail the backup — an inconsistent backup that reports success is worse than no backup. If the post-hook fails, continue anyway, because the backup data is already good and a failed backup record would hide that; the cost is that the site can be left in maintenance mode, so the escape-hatch command lives in the runbook.
That hook also produced the best small bug of the epic. The first live run came
back PartiallyFailed with three hook failures reading no such container: nextcloud. The label selector matched app.kubernetes.io/name: nextcloud —
and so did the chart's completed cron Job pods, which carry the same label
and contain no nextcloud container. Adding app.kubernetes.io/component: app fixed it. A backup that reports partial failure because of pods that
aren't the app is exactly the kind of noise that trains you to ignore backup
alerts.
The month the file backups couldn't have worked
Here is the part I'd most want another self-hoster to read.
Velero's node-agent refuses hostPath volumes outright. The cluster's
storage class at the time was a local path provisioner, so every PVC on the
platform was ineligible for file backup. Manifests backed up fine. File data —
the actual company documents — was not being backed up at all.
The obvious workaround is to convert the volumes to the in-tree local type,
which the node-agent does accept. Tried it. It resolves empty through
Talos's containerised kubelet: the restore completes, reports success, and
produces nothing. That failure mode is worse than the refusal, because the
refusal at least tells you.
Then, while investigating, the genuinely alarming discovery. Nextcloud's
entire 949 MB dataset — including the OIDC app's code — was not on the host at
all. It lived inside the kubelet container's overlay filesystem, while the
PV's host path was an empty directory. The kubelet resolves hostPath sources
in its own mount namespace, so both views were internally consistent and one
of them was a lie.
That data was one Talos upgrade away from ceasing to exist. Not "unbacked up." Gone.
The fix was a real CSI driver, whose privileged node pod does the mounting so the kubelet only ever sees standard CSI paths. Worth noting for anyone about to reach for the obvious alternative: OpenEBS LocalPV-Hostpath is not CSI — it emits in-tree-plugin PVs and would have reproduced the failure exactly.
Restore day, with numbers
Manifests only (before the storage fix): a Nextcloud namespace restored
into a scratch namespace in ~54 seconds, with Deployments, Services and
ExternalSecrets recreated and re-synced from the vault. Nextcloud itself
stayed Pending on a missing PVC — the correct, documented outcome of
excluding volumes, not a failure.
With real file data (after the fix): backup in 40 seconds, including a
920 MB Nextcloud volume in 22 seconds. The restore took about six
minutes — kopia downloading the whole volume — finishing at 919,966,681
bytes, byte-identical to the backup. Verified by content, not by status:
the OIDC app's directory, config.php, and real user files all present.
Six minutes against 54 seconds is the number worth internalising. Manifest restores are nearly free; data restores are bounded by your downlink, and that ratio is what actually shapes a recovery plan.
One more thing that looks alarming and isn't: restoring into a cluster that Argo CD manages produces roughly 55 "already exists" warnings and zero errors. That's the skip working correctly — git already owns those objects.
What this really buys you
An untested backup is a belief. The value of restore day isn't the timings, it's that a human has walked the path, hit the sharp edges, and written them into a runbook — so the bad day is a procedure rather than an improvisation.
It also produced the more uncomfortable insight: for a month, the thing that was broken wasn't the backup tool. It was an assumption underneath it, in a layer nobody thinks of as part of the backup system at all.
Gotchas
Every one of these cost real time.
kubectl delete backupdeletes almost nothing. No finalizer intercepts it, and the sync period resurrects the object from the bucket's metadata within about a minute, complete with a fresh UID. Use aDeleteBackupRequestif you actually want the data gone.kubectl get backupis ambiguous wherever CNPG and Velero coexist — both register the same short name. Fully qualify, always. This one has now bitten across three separate epics.- Backup alerts must be tested by breaking something on purpose — but break something isolated. The test used a throwaway backup location pointed at a nonexistent bucket, never touching the real one. Confirming the email actually arrives is the entire point of having the alert.
- Beware the metric that sits at zero forever. One Postgres backup metric never populates with the plugin in use, so a comparison-based alert rule fired permanently off a long-resolved historical failure. Rewritten as a recency window on the failure timestamp. Verify metric names against real data before writing PromQL, not against documentation.
- A
PartiallyFailedbackup usually means your selector is too broad. Completed Job pods sharing an app's labels will happily fail your exec hooks. - Succeeded cron pods pin a PVC's finalizer. Leftover completed pods keep
pvc-protectionalive and the PVC hangs inTerminatingforever. Delete the pods. - Never restore a Postgres volume from a file backup. Yes, it's in the table above too. It is worth saying twice.
Where this leaves us
Every category of data on this platform now has an owner, a schedule, a retention policy, and — the part that took real work — a demonstrated restore with a timing next to it.
Which raises the only question left, and it's the biggest one in the series. All of this assumes the cluster still exists. What happens when it doesn't? Next post, we delete the entire platform on purpose and rebuild it from nothing, with a stopwatch running.
I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.