29 July 2026
Your platform needs a vault: OpenBao, ESO, and the SOPS you keep
A runtime secret store on your own cluster, bridged into Kubernetes by External Secrets — and the two secrets that can never live in it, written down as an exhaustive list.
openbaoexternal-secretssopskubernetessecretseu-sovereignty
In August 2023, HashiCorp moved Vault to the Business Source License. Nothing
broke that day. Every running cluster kept running, every vault binary kept
working, and for most people the change was an item in a newsletter.
That's exactly what makes it interesting. A licence isn't an outage — it's a dependency that changes shape while you're not looking. If your pitch is you can always leave, then the thing holding every credential you own is a strange place to accept a vendor's terms as a permanent given.
This is the first post in a new series. The earlier posts built a platform that runs a website from a git commit with nobody touching the cluster. Now it has to run real workloads — databases, identity, office apps, a SaaS with paying customers — and that starts with the boring question every one of those depends on: where do the secrets live?
What you'll get from this post
The full shape of a self-hosted secrets layer: OpenBao as the runtime store, External Secrets Operator bridging it into Kubernetes so app manifests contain no secret material at all, and — the part these write-ups usually skip — the handful of secrets that cannot live in the vault, why they can't, and how to keep that list from quietly growing.
A licence is a dependency
The steelman for Vault is real and I'll make it properly: it's more mature, it has a far larger ecosystem of docs and integrations, it offers HSM and cloud-KMS auto-unseal, and there's a much bigger community to debug against at 2am. On features, it wins.
It lost anyway, because the moat here is jurisdictional sovereignty, and a component whose licensing terms can shift under us is a soft spot in exactly the place that's supposed to be solid.
OpenBao is the Linux Foundation fork of the last
MPL-2.0 Vault release: OSI-licensed, community-governed, and API-compatible
in every place this platform touches it — KV v2, Kubernetes auth, integrated
raft storage, and sys/storage/raft/snapshot. At this tier, choosing it costs
nothing functional and buys back the open-source guarantee.
The more interesting rejection was Infisical, which is a genuinely nice piece of software. Its self-hosted mode needs Postgres and Redis, and that breaks a property that only becomes visible when you think about disasters: the secret store must be restorable before the database layer exists. In a bare-metal rebuild you need secrets to bring up the database — so a secret store that depends on a database is a circular dependency you'll discover on the worst possible day. OpenBao's integrated raft storage has no such dependency. (That instinct turned out to be right for reasons I'll get to in a later post, when we deliberately destroy the whole platform to find out.)
The shape: OpenBao holds it, ESO hands it over
Apps never talk to the vault. They consume ordinary Kubernetes Secrets, and
External Secrets Operator keeps those in sync from OpenBao. That indirection
is what keeps app manifests completely free of both secret material and vault
ceremony — a Deployment looks the same whether its password came from a vault,
a sealed secret, or a sticky note.
Everything hangs off one ClusterSecretStore:
apiVersion: external-secrets.io/v1
kind: ClusterSecretStore
metadata:
name: openbao
spec:
provider:
vault:
server: https://openbao.openbao.svc:8200
path: platform
version: v2
auth:
kubernetes:
mountPath: kubernetes
role: external-secrets
serviceAccountRef:
name: external-secrets
namespace: external-secrets
caProvider:
type: Secret
name: openbao-server-tls
key: ca.crt
namespace: openbao
Two details in there earn their place. ESO authenticates with its own
ServiceAccount through OpenBao's Kubernetes auth method, so there's no static
token to rotate — the thing that usually becomes the worst secret in the
system. And the CA trust comes from caProvider reading the ca.crt key out
of the same Secret cert-manager already wrote for OpenBao's own TLS listener.
No separate CA distribution, no copy of a certificate drifting out of date in
a ConfigMap somewhere.
From there, one path convention — platform/<app>/<thing> — and every app gets
the same four-line shape:
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: keycloak-pg-app
namespace: keycloak
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
name: keycloak-pg-app
template:
type: kubernetes.io/basic-auth
dataFrom:
- extract:
key: keycloak/pg
That's an entire app's database credential. The manifest is committable, reviewable, and contains nothing you'd mind a stranger reading.
The vault has no front door
OpenBao runs as a single replica with no public route at all — no
HTTPRoute, no Gateway listener, ClusterIP only. Operator access is
kubectl port-forward or nothing. It's the one component on the platform
that never touches the public edge, and keeping it that way costs nothing
because humans should be interacting with it approximately never.
Storage is raft rather than the chart's file default, specifically because
raft is what makes bao operator raft snapshot save possible — and a secret
store you can't snapshot is a secret store you're going to lose. Those
snapshots run nightly at 03:15 UTC, slotted into a staggered schedule
(Postgres 01:15, file backups 02:15, OpenBao 03:15) because this is one box
with one uplink and three backup tools that would otherwise all wake up at
once.
One line in the config looks like a smell and isn't:
disable_mlock = true
mlock stops secrets being paged to disk, so disabling it reads like sloppy
security. It's required here: the pod's securityContext drops all
capabilities, IPC_LOCK included, so mlock cannot succeed. The choice
isn't "mlock or not", it's "drop every capability, or keep one to enable
mlock" — and on a node with no swap, dropping the capabilities is the better
trade. Worth writing the reasoning next to the line, because a future reader
(me, at 2am) will otherwise assume it's a shortcut.
Pod restart means sealed
Here's the part that will annoy you, and it's deliberate.
There's no auto-unseal. Every restart — chart bump, node reschedule, crash —
leaves the vault sealed until an operator runs bao operator unseal with
a key that exists only offline. Shamir shares are set to 1-of-1, because at
one operator, splitting into 3-of-5 adds ceremony on every restart without
adding security: there's no second human to hold the second share.
Auto-unseal was rejected on sovereignty grounds. Every viable backend — AWS KMS, GCP KMS, Azure Key Vault — is a non-EU dependency sitting underneath the component that holds every credential on the platform, which rather defeats the exercise. The remaining option, transit auto-unseal, needs a second running OpenBao to unseal the first, doubling the footprint on a single-node cluster to solve a problem that one command already solves.
The honest consequence, stated precisely, because "it's fine" is not an
engineering claim: a sealed vault degrades the ClusterSecretStore, which
stalls Argo CD's root sync at wave 5. New platform/ changes can't land while
it's sealed. Running workloads are unaffected — child Applications keep
self-syncing, and pods already holding their Secrets carry on — so a sealed
vault is a deployment freeze, not an outage. That distinction is the whole
reason it's an acceptable trade at this tier.
The part nobody writes about
Every secrets tutorial ends at "and now everything is in the vault." Nothing is ever everything. There is always a set of secrets that has to exist before the vault does, and the quality of a secrets architecture is mostly determined by whether that set is written down or discovered.
Ours is written down, in an ADR, as an exhaustive table. Three structural entries:
- The age private key itself. It's what decrypts everything else. It can't be turtles all the way down — something has to be the bottom, and the bottom lives offline.
- Argo CD's git repository credential. Argo CD has to authenticate to git before it can reconcile anything, and "anything" includes OpenBao's own manifests. This is the single bootstrap secret the entire GitOps model rests on.
- OpenBao's own S3 backup credentials. OpenBao cannot be the store for the credential that restores OpenBao from a snapshot when OpenBao is the thing that's gone.
Each of those is genuinely circular. None of them is "convenient to keep in
SOPS." Those three are encrypted with SOPS + age and committed as
*.sops.yaml files, which is fine — the encrypted files belong in git; the
private key never does.
There's a fourth row in the table that's deliberately empty: a category placeholder for "anything ESO needs before OpenBao is reachable," which currently has no members because ESO's trust is a cert-manager certificate rather than a secret. An empty row looks like clutter and isn't — it's the difference between "we thought about this category and it's empty" and "we never thought about it."
The rule that keeps the list honest: extending it is a pull request to the ADR. Add a row, state why the secret can't wait for OpenBao, and give it a sunset date unless it's structural. That's not bureaucracy for its own sake; it's the specific mechanism that stops "SOPS for now, migrate later" from becoming the permanent default, which is how these lists rot.
It has already been exercised. The SMTP credential needed to exist before OpenBao did, so it went in with a time-box and a named ticket to remove it. When that ticket landed, the credential moved to OpenBao and the row was deleted from the table — not struck through, not marked "done." The point of a sunset is that the row stops existing.
And the whole thing is auditable with one command:
$ find . -name '*.sops.yaml'
./addons/argocd-config/repo-platform-gitops.sops.yaml
./addons/openbao-backup/openbao-backup-s3.sops.yaml
Two files, matching two rows. (The age key is row one and isn't a file in the repo at all — that's the point of it.) Anything else appearing in that output is a bug, not a judgement call, and that's what makes the boundary a rule rather than an aspiration.
What this really buys you
A secrets architecture isn't a choice of vault. It's knowing precisely which secrets can't be in the vault, having written that down somewhere with teeth, and being able to check it in one command.
The vault is the easy half — it's a Helm chart and an afternoon. The boundary is the half that decides whether, two years from now, you have a secret store or a secret store plus fourteen encrypted files nobody remembers the reason for.
Gotchas
Every one of these cost real time.
- There is no
opensslin the OpenBao pod. The obvious seeding command,bao kv put ... password="$(openssl rand -base64 24)", runs in busyboxashwhereopenssldoesn't exist. The substitution expands to an empty string,baostorespassword="", and it reports success. This crash-looped Keycloak twice behind the wonderfully unhelpful error "bootstrap-admin-username available only when bootstrap admin password is set" — the password was set, to nothing. Usehead -c 24 /dev/urandom | base64, and verify withbao kv get -format=jsonbefore you walk away. - ESO's
template.datareplaces the entire Secret by default. The defaultmergePolicyisReplace, meaning values pulled bydata:ordataFrom:are available as template inputs but do not survive into the output Secret. Adding one non-secret literal alongside five real credentials would have silently shipped a Secret containing only the literal.mergePolicy: Mergeis the fix; noticing you need it is the hard part, because nothing errors. - Adding an extract to a live
ExternalSecretis a production change. AlldataFromextracts share one template namespace, and a missing key fails the whole Secret — not just the new key. Merging a new extract before seeding its path would have dropped the database URL, the OIDC client secret and the session secret off a running application. Seed the path first, then merge the manifest. Never the other way round. refreshInterval: 1hmeans an hour. After seeding a new value, force the read withkubectl annotate externalsecret <name> force-sync=$(date +%s) --overwriteinstead of staring at a Secret wondering why it's empty.envFrominjects environment variables at container start. When ESO rewrites a Secret, a running pod keeps the old values indefinitely. Delete the pod — do notrollout restart, which edits the pod template and puts you in a fight with Argo CD's selfHeal that you will lose.- Split a memory budget by component, not by arithmetic. ESO ships three Deployments, and dividing one capacity row evenly between them OOM-looped the cert-controller continuously — 1,707 restarts in six days — while the secret store stayed nominally healthy, so nothing alerted. Its limit went from 48Mi to 128Mi, making it the largest of the three despite doing the least visible work. Measured reality beat the tidy split.
Where this leaves us
There's a vault on the cluster, apps read credentials from it without containing any, and the exceptions are a two-line command rather than an act of memory. The platform can now be handed a secret without anybody making a judgement call about where to put it.
Which unblocks the thing every real workload actually needs: a database that somebody else operates. Next post is Postgres — CloudNativePG, backups to object storage, and the restore timings that decide whether any of it counted.
I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.