2 August 2026

Sovereign Postgres: CNPG with backups you can actually restore

CloudNativePG on a single cheap box, backed up to EU object storage with the Barman Cloud Plugin — and the two restore drills, with timings, that decide whether any of it counted.

postgrescloudnativepgbackupskuberneteshetznerdisaster-recovery


Everybody has backups. Far fewer people have restores.

The distinction sounds pedantic right up until the morning you need it. A backup is a file in a bucket and a green tick in a dashboard. A restore is a running database with your data in it. The gap between the two is where companies die, and the only way to know which side you're on is to actually perform the restore, on purpose, before you need it.

In the last post the platform got a vault, which means it can now be handed a credential without anyone making a judgement call about where to put it. That unblocks the thing every real workload needs next: a database somebody else operates. In this platform's case, "somebody else" is an operator running on the same box — and the contract says apps speak Postgres over the standard wire protocol and nothing more exotic, so what's underneath is free to change.

What you'll get from this post

The full shape of Postgres as a platform service: one shared CloudNativePG cluster, per-app databases and roles declared in git, continuous WAL archiving plus nightly base backups into EU object storage — and then the part that matters, two restore drills with real measured timings, including the failure mode that made the first point-in-time recovery attempt abort.

One cluster, many databases

The obvious design is a Postgres cluster per app. It's cleaner on paper: blast-radius isolation, independent upgrades, no noisy neighbours.

It loses here on arithmetic. Every CNPG cluster is at least one pod with its own shared buffers, its own WAL, its own connection overhead and its own backup schedule. Four apps means four of everything on a machine with 8 GB of RAM. At this tier that's not isolation, it's just less headroom for the actual workloads — and headroom on a single node is the resource you miss first.

So: one shared cluster, platform-pg, and isolation at the database and role level instead.

spec:
  instances: 1
  enableSuperuserAccess: false
  storage:
    storageClass: local-hostpath
    size: 10Gi

instances: 1 looks like a shortcut and isn't — high availability is a number in this file, not a rewrite. The property worth protecting is that going to three instances later changes nothing above the line: same service name, same connection string, same app code. That's the whole reason the contract is small.

enableSuperuserAccess: false is the clause I'd keep even at gunpoint. No app ever gets the superuser, and disabling it at the cluster level turns that from a convention someone can forget into a thing that isn't available to forget.

Each app then gets a role and a database, declared:

apiVersion: postgresql.cnpg.io/v1
kind: Database
metadata:
  name: platform-pg-keycloak     # cluster-prefixed: CNPG convention
  namespace: platform-pg
spec:
  name: keycloak                 # the literal PostgreSQL database name
  owner: keycloak
  cluster:
    name: platform-pg

Adding an app is one more Database, one more entry under managed.roles, and one more ExternalSecret. No CREATE DATABASE typed at a prompt, ever.

The role passwords come straight from the vault built in the last post: the password of record lives in OpenBao at platform/<app>/pg, and an ExternalSecret renders it into a kubernetes.io/basic-auth Secret labelled cnpg.io/reload: "true". That label is what makes rotation declarative — change the value in OpenBao and CNPG reconciles the role password itself.

The in-tree backup config is deprecated — use the plugin

This is the part that's likely to bite anyone following an older tutorial. CNPG's original backup support was spec.backup.barmanObjectStore, configured inline on the Cluster. It has been deprecated since CNPG 1.26, and the supported path is now the Barman Cloud Plugin — a separate CNPG-I component that ships its own ObjectStore CRD.

Practically, backup configuration moves out of the Cluster into its own object:

apiVersion: barmancloud.cnpg.io/v1
kind: ObjectStore
metadata:
  name: platform-pg
spec:
  retentionPolicy: "30d"
  configuration:
    destinationPath: s3://dreamcodefactory-platform/backups/pg
    endpointURL: https://nbg1.your-objectstorage.com
    wal:
      compression: zstd
    data:
      compression: gzip

and the Cluster just points at it:

  plugins:
    - name: barman-cloud.cloudnative-pg.io
      isWALArchiver: true
      parameters:
        barmanObjectName: platform-pg

Two details in there are not obvious. The compression algorithms deliberately differ: zstd for WAL, because it has the best ratio-to-speed tradeoff on a stream of small files, and gzip for base backups, because data compression doesn't support zstd at all and gzip has the fastest decompression of the supported options — which is to say it's chosen for RTO, not for bucket size. The thing you optimise in a backup is the restore.

And isWALArchiver: true is what marks this store as the WAL destination rather than a read-only recovery source. The restore drill's cluster deliberately omits it, which is what keeps a drill from writing into the production backup chain.

WAL archiving gives continuous point-in-time recovery; the nightly base backup just bounds how much WAL has to be replayed. With archive_timeout at five minutes, the worst-case data loss is ≤ 5 minutes.

spec:
  schedule: "0 15 1 * * *"
  immediate: true

That runs at 01:15 UTC — a deliberate slot. Postgres at 01:15, files at 02:15, the vault at 03:15: one backup job per hour, because this is one box on one uplink and three simultaneous uploads is a self-inflicted incident.

The drills

Everything above is a claim. Here's the part that turns it into a fact.

Two drills, both against a throwaway cluster in a scratch namespace, neither ever touching production. The first proves the chain is restorable at all. The second proves the point-in-time half, which is a genuinely different claim.

The full restore. Write a marker row, force a base backup so it's captured immediately, then build a brand-new Cluster from the bucket and check the row comes back:

$ kubectl -n pg-restore-drill get cluster platform-pg-restore
NAME                  INSTANCES   READY   STATUS                     PRIMARY
platform-pg-restore   1           1       Cluster in healthy state   platform-pg-restore-1

~79 seconds, manifest apply to Cluster in healthy state, with the marker row's written_at matching production byte for byte. Not "a file exists in the bucket" — an actual database, rebuilt from object storage, serving the row.

The point-in-time restore. Two markers with a recovery target set between them; only the earlier one should survive. ~51 seconds, with the target set 32 seconds before the write that correctly did not come back.

Both numbers deserve their caveat, and it's a big one: this is a single-instance launch-tier database with a small dataset on local NVMe. These are not enterprise restore times and nobody should quote them as such. What they establish isn't speed — it's that the path works end to end, and roughly how long the pieces take, which is the number you actually need when you're deciding whether to restore or to keep debugging.

The first PITR attempt failed outright, which turned out to be the most useful thing that happened all day. See the gotchas.

What this really buys you

An untested backup is not a backup. It's a belief about a bucket.

The point of the drills isn't the timings — it's that the runbook has been executed by a human who hit the sharp edges, wrote them down, and appended a row to a log. What you want on the bad day is not a backup strategy. It's a procedure somebody has already followed once, with the surprises removed.

Gotchas

Every one of these cost real time.

  • Point-in-time recovery needs a transaction after your target. The first PITR drill aborted with FATAL: recovery ended before configured recovery target was reached — even though replay had reached exactly the right point. If the last transaction Postgres can find is chronologically before the target and no WAL exists past it, Postgres won't stop cleanly, because it has no evidence time moved past the target. Write a throwaway buffer transaction well after the target, force a WAL switch, and give the archiver a few seconds before restoring.
  • kubectl get backup is ambiguous on a cluster that also runs Velero. Both backups.postgresql.cnpg.io and backups.velero.io register the same short name. You will eventually watch the wrong resource stay empty and conclude your backups are broken. Always fully qualify.
  • Multiple statements in one psql -c run as one implicit transaction. Semicolon-joining them means a later error silently rolls back the earlier ones — including the marker INSERT your entire drill depends on. Pass each statement as its own -c.
  • enableSuperuserAccess: false doesn't lock you out. pg_switch_wal() needs superuser, which no app role has by design, but an operator can still reach it over the local socket with kubectl exec into the primary. The flag disables the managed password-based superuser Secret, not local peer-trust. That distinction is the difference between a drill you can run and one you can't.
  • The schedule is a six-field cron, seconds first. CNPG uses Go's robfig/cron dialect, so "0 15 1 * * *" is 01:15:00 — not 15:01 as a standard five-field crontab line would read. An off-by-one-field schedule runs at a plausible-looking wrong time, which is the worst kind of wrong.
  • immediate: true fires once per ScheduledBackup lifetime, not once per apply. If you delete and recreate the cluster during a failure, delete the ScheduledBackup too, or the immediate base backup silently doesn't re-fire.
  • Hetzner Object Storage is Ceph RGW, not S3. Recent AWS SDK data-integrity headers trip its checksum handling, so the sidecar needs AWS_REQUEST_CHECKSUM_CALCULATION=when_required and the matching response flag. Same quirk every other consumer of the bucket already needed.
  • Generate database passwords alphanumeric-only. One containing a / parsed fine everywhere except one app's URL-based connection string, which failed in a way that pointed at everything except the password.
  • Never restore a Postgres data volume from a file-level backup. The file-backup tooling in a later post explicitly excludes CNPG volumes. Barman and PITR are the only correct path back; a filesystem copy of a running database is how you turn an outage into corruption.

Where this leaves us

The platform now has a database it operates for you: declared in git, isolated per app, never handing out the superuser, backed up continuously to EU object storage — and, more to the point, restored twice with the timings written down.

That's two of the three things every real workload needs. It has somewhere to put secrets and somewhere to put data. The third is knowing who the user is — so next we self-host identity: Keycloak, with the realm itself as code, and the admin console demoted to read-only by convention.


I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.