12 August 2026

Back up everything, then prove it: Velero + restore day

Four backup mechanisms, one bucket, and a table of what each one does not cover — plus a restore day with real timings and the storage bug that made file backups silently impossible.

velerobackupskubernetestalosdisaster-recoveryhetzner


The most dangerous backup is the one that runs successfully every night and cannot be restored.

It's dangerous specifically because it generates evidence of safety. Green ticks, growing bucket, a job that exits zero. Every signal you'd think to check says you're covered, and none of them are the thing you actually need, which is a restore.

The last post put company files and a password vault on this platform — the first data whose loss would genuinely hurt. So this post is where the project pays that debt: what backs up what, what each mechanism explicitly does not cover, and a restore day with real numbers, including the one that took six minutes instead of the fifty-four seconds I expected.

What you'll get from this post

The complete backup responsibility split for a small self-hosted platform — four mechanisms, one bucket, and the boundaries between them — plus the storage-layer bug that made file backups impossible for a month while reporting success, and what an actual restore drill measured.

Four mechanisms, and what each one refuses to do

The table matters more than any individual tool, because almost every backup disaster is a boundary problem: two mechanisms that each assumed the other had it.

DataMechanismDoes not cover
All Postgres databasesCNPG + Barman (WAL + nightly base, PITR)anything outside Postgres
PVC file data + manifestsVelero + node-agent file backup (kopia), nightlyCNPG data volumes — deliberately excluded
All runtime secretsOpenBao raft snapshots, nightlyanything not in the vault
Everything declarativegit itselfruntime data of any kind

The row that gets people is the second one. Velero never touches the Postgres volumes, and that exclusion is structural rather than a namespace somebody forgot: the platform-pg namespace isn't in the schedule, and volume backup is opt-in by annotation, so a new database would have to be actively opted in to break the rule.

The reason is worth stating plainly, because "restore the PVC" is the obvious instinct: a file-level copy of a running database's data directory is not a backup, it's a torn snapshot. Restoring one gives you a database that starts, passes a smoke test, and is subtly corrupt. Barman and PITR are the only correct path back for Postgres.

Everything lands in one bucket under different prefixes, which is a deliberate and slightly contrarian call. Object storage charges per bucket, and the credentials are project-wide anyway, so a second bucket buys separation that the credential model doesn't actually deliver — while definitely buying another line item. A second bucket needs a written decision, not a convenience.

And the schedules are staggered: Postgres at 01:15, files at 02:15, the vault at 03:15. One box, one uplink; three simultaneous uploads is a self-inflicted incident.

App-consistent, or it doesn't count

Copying Nextcloud's files while somebody is uploading gives you a backup of a half-written file and a database that disagrees with it. So the backup brackets itself with maintenance mode:

pre:
  - exec:
      command: [..., "occ maintenance:mode --on"]
      onError: Fail
post:
  - exec:
      command: [..., "occ maintenance:mode --off"]
      onError: Continue

The asymmetry is intentional. If the app can't be quiesced, fail the backup — an inconsistent backup that reports success is worse than no backup. If the post-hook fails, continue anyway, because the backup data is already good and a failed backup record would hide that; the cost is that the site can be left in maintenance mode, so the escape-hatch command lives in the runbook.

That hook also produced the best small bug of the epic. The first live run came back PartiallyFailed with three hook failures reading no such container: nextcloud. The label selector matched app.kubernetes.io/name: nextcloud — and so did the chart's completed cron Job pods, which carry the same label and contain no nextcloud container. Adding app.kubernetes.io/component: app fixed it. A backup that reports partial failure because of pods that aren't the app is exactly the kind of noise that trains you to ignore backup alerts.

The month the file backups couldn't have worked

Here is the part I'd most want another self-hoster to read.

Velero's node-agent refuses hostPath volumes outright. The cluster's storage class at the time was a local path provisioner, so every PVC on the platform was ineligible for file backup. Manifests backed up fine. File data — the actual company documents — was not being backed up at all.

The obvious workaround is to convert the volumes to the in-tree local type, which the node-agent does accept. Tried it. It resolves empty through Talos's containerised kubelet: the restore completes, reports success, and produces nothing. That failure mode is worse than the refusal, because the refusal at least tells you.

Then, while investigating, the genuinely alarming discovery. Nextcloud's entire 949 MB dataset — including the OIDC app's code — was not on the host at all. It lived inside the kubelet container's overlay filesystem, while the PV's host path was an empty directory. The kubelet resolves hostPath sources in its own mount namespace, so both views were internally consistent and one of them was a lie.

That data was one Talos upgrade away from ceasing to exist. Not "unbacked up." Gone.

The fix was a real CSI driver, whose privileged node pod does the mounting so the kubelet only ever sees standard CSI paths. Worth noting for anyone about to reach for the obvious alternative: OpenEBS LocalPV-Hostpath is not CSI — it emits in-tree-plugin PVs and would have reproduced the failure exactly.

Restore day, with numbers

Manifests only (before the storage fix): a Nextcloud namespace restored into a scratch namespace in ~54 seconds, with Deployments, Services and ExternalSecrets recreated and re-synced from the vault. Nextcloud itself stayed Pending on a missing PVC — the correct, documented outcome of excluding volumes, not a failure.

With real file data (after the fix): backup in 40 seconds, including a 920 MB Nextcloud volume in 22 seconds. The restore took about six minutes — kopia downloading the whole volume — finishing at 919,966,681 bytes, byte-identical to the backup. Verified by content, not by status: the OIDC app's directory, config.php, and real user files all present.

Six minutes against 54 seconds is the number worth internalising. Manifest restores are nearly free; data restores are bounded by your downlink, and that ratio is what actually shapes a recovery plan.

One more thing that looks alarming and isn't: restoring into a cluster that Argo CD manages produces roughly 55 "already exists" warnings and zero errors. That's the skip working correctly — git already owns those objects.

What this really buys you

An untested backup is a belief. The value of restore day isn't the timings, it's that a human has walked the path, hit the sharp edges, and written them into a runbook — so the bad day is a procedure rather than an improvisation.

It also produced the more uncomfortable insight: for a month, the thing that was broken wasn't the backup tool. It was an assumption underneath it, in a layer nobody thinks of as part of the backup system at all.

Gotchas

Every one of these cost real time.

  • kubectl delete backup deletes almost nothing. No finalizer intercepts it, and the sync period resurrects the object from the bucket's metadata within about a minute, complete with a fresh UID. Use a DeleteBackupRequest if you actually want the data gone.
  • kubectl get backup is ambiguous wherever CNPG and Velero coexist — both register the same short name. Fully qualify, always. This one has now bitten across three separate epics.
  • Backup alerts must be tested by breaking something on purpose — but break something isolated. The test used a throwaway backup location pointed at a nonexistent bucket, never touching the real one. Confirming the email actually arrives is the entire point of having the alert.
  • Beware the metric that sits at zero forever. One Postgres backup metric never populates with the plugin in use, so a comparison-based alert rule fired permanently off a long-resolved historical failure. Rewritten as a recency window on the failure timestamp. Verify metric names against real data before writing PromQL, not against documentation.
  • A PartiallyFailed backup usually means your selector is too broad. Completed Job pods sharing an app's labels will happily fail your exec hooks.
  • Succeeded cron pods pin a PVC's finalizer. Leftover completed pods keep pvc-protection alive and the PVC hangs in Terminating forever. Delete the pods.
  • Never restore a Postgres volume from a file backup. Yes, it's in the table above too. It is worth saying twice.

Where this leaves us

Every category of data on this platform now has an owner, a schedule, a retention policy, and — the part that took real work — a demonstrated restore with a timing next to it.

Which raises the only question left, and it's the biggest one in the series. All of this assumes the cluster still exists. What happens when it doesn't? Next post, we delete the entire platform on purpose and rebuild it from nothing, with a stopwatch running.


I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.