19 August 2026
Shipping a SaaS on the platform: the contract in practice
Taking the four-clause contract from theory to a real product running in production — one pull request, and the friction list from doing it the first time.
kubernetesgitopssaasplatform-engineeringnextjspostgres
Six posts of platform-building all rest on one promise made right at the start: that an application meeting a deliberately small contract can be run, observed, backed up and moved without its author thinking about any of it.
Promises like that are cheap until something real is standing on them. So this post is the bill arriving — an actual product, deployed to production through the same pipeline as everything else, and an honest list of what got in the way.
The previous post proved the platform survives being destroyed. This one asks whether it's pleasant to use, which is a different and slightly more embarrassing question.
What you'll get from this post
The four-clause contract demonstrated end to end: what a new application actually costs to onboard, the one moment where the platform visibly earned its design, and the complete friction list from the first real run — including a genuinely nasty authentication bug that anyone building on the same stack will hit.
One pull request
The paved path is a directory to copy and a token to rename:
$ cp -r apps/_template apps/starter
$ sed -i '' 's/myapp/starter/g' apps/starter/*.yaml
$ git mv apps/starter/application.yaml platform/starter.yaml
$ grep -rn myapp apps/starter platform/starter.yaml # MUST print nothing
That third line is load-bearing. Nothing under apps/ is synced — the
application definition only becomes real once it's moved into platform/.
Leave a copy behind and you get two controllers managing the same resources,
which is a confusing afternoon. Hence the two verification commands: one that
must print nothing, and one that must fail.
Beyond the copy, onboarding is four edits to shared files — a TLS listener, an identity client, the database claim, and the backup namespace list — plus two secrets seeded into the vault and one DNS record. That's the whole cost. One pull request in the workloads repo, one in infrastructure for DNS.
The template also carries a rule that turned out to matter more than the manifests: if a step isn't covered by the template, fix the template. That got exercised almost immediately — the authentication library needed a fourth secret the template didn't wire up. The fix was a template change, so the next app inherits it. The alternative, patching one app and moving on, is how paved paths quietly become goat tracks.
The eighty seconds that justified the design
Here's the moment the contract stopped being theory.
On first deploy, the application came up before its database role and schema had finished provisioning. The pod correctly reported not ready — traffic was routed away, the ingress held off — for about eighty seconds. Then Postgres became reachable, readiness flipped, and traffic flowed.
Zero restarts.
That is the entire reason the contract insists on two distinct health checks rather than one. With a single "health" endpoint, that same eighty seconds is a crash loop: the container gets killed for being unhealthy, restarts into the same unready state, gets killed again, and backs off — turning a brief dependency delay into an outage plus a red dashboard.
Two checks, two meanings. Am I alive is answered by the process existing. Am I ready is answered by the database being reachable. A dependency being slow is a routing decision, never a restart decision.
Then the boring confirmation that matters: write a row, restart the pod deliberately, read the row back on the fresh pod. Real persistence in the shared cluster, not state hiding in memory.
The bug worth the whole post
The application signs users in through the platform's identity provider and keeps a small amount of per-user state, keyed on the user's identity.
Signing in and out repeatedly with the same account produced three separate rows for one human.
The cause is a genuinely surprising piece of upstream behaviour. The auth library, when configured with no database adapter — which is the database-agnostic default, and deliberate here — assigns each sign-in a fresh random UUID as the user id. It's intentional: the comment upstream explains that an adapter is expected to override it by looking up the stored identity. With no adapter, that randomness silently propagates into the session token's subject claim and stays there forever.
So the identity looked stable. It rendered correctly. It just wasn't the same value twice.
The fix is one line — explicitly set the token's subject from the identity provider's own subject claim, which is mandatory and stable by specification. The reason it's worth an entire section is the failure shape: nothing errors, and everything looks right until you check whether two sign-ins produced one row or two. Any application built on this pattern that persists state keyed on user identity has this bug right now and doesn't know.
What this really buys you
The measure of a platform isn't whether an app can be deployed to it. It's what the second app costs.
Here the answer is a directory copy, a rename, four small edits and two secrets — because every hard decision was already made and encoded: where secrets live, which database, how identity works, what gets backed up, which probes mean what. The application author inherits all of it by meeting four clauses and never thinking about any of them again.
And the friction that remains is now written down, which is what turns a one-off success into a repeatable one.
Gotchas
Every one of these cost real time.
- Container registry access has two separate controls. A package being
linked to its repository and that repository's automation being allowed to
write to the package are different settings, and the second has no API
or CLI path — a human clicks it in a web UI. The first push fails with a
bare
403that suggests a token problem it isn't. Any package seeded by a manual push before its first automated build hits this. - The GitOps controller's manifest cache goes stale, per application. It compared an old revision while child apps synced the new one — four times in one epic. A hard-refresh annotation fixes it in seconds, but it has to be applied to each affected app, not just the root, which is the part that wastes the time.
- Stacked pull requests plus squash merges conflict. Once the base squashes,
the child's history no longer matches.
git rebase --ontois the recovery; not stacking in the first place is the fix. - Never let a debug image mount real secrets. A generic echo container will happily dump every environment variable it was given to anyone who asks. It ran with the secret unmounted, and that was luck plus a checklist rather than a guarantee.
- Check live capacity, not the projected table. A live check found memory requests already at 87.8% — past the documented 75% resize trigger, and entirely independent of the small app being added that day. Projections drift from reality in the direction that hurts.
- Verify multi-arch before pinning an image, and note that inspecting a private package needs an authenticated client — an anonymous check just returns an authorization error, which is easy to misread as "not multi-arch."
- Tearing an app back down is part of the paved path. Namespace, database, role, identity client, orphaned TLS secret, vault entries, DNS. If removal isn't documented, every experiment leaves sediment.
Where this leaves us
There's a real product on the platform now, deployed by merging a pull request, authenticating against self-hosted identity, persisting to a database that's backed up and restore-tested, and behaving correctly when its dependencies are slow.
Which leaves exactly one thing between "an app that runs" and "a business": getting paid. The final post takes money for a feature — and keeps the entitlement state in our own database rather than a payment provider's, for reasons that a bug found in production makes very concrete.
I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.