22 July 2026

Talos on Hetzner with OpenTofu, from zero

A reproducible single-node Talos Linux cluster on a cheap Hetzner box in the EU — no SSH, no cloud-init, nothing clicked in a console.

taloshetzneropentofukubernetesiac


Most "spin up a cluster" guides end with you SSHed into a box, having typed a dozen commands you'll never remember, with no way to reproduce what you just did. That isn't infrastructure — it's a sandcastle.

This is the first hands-on post in the series, and the rule from here on is absolute: if it isn't in code, it doesn't exist. By the end you'll have a working Kubernetes cluster on a Hetzner box, created entirely by OpenTofu, that you can destroy and rebuild identically on demand — and the node won't even have SSH for you to drift on.

In the last post I described the contract that keeps apps portable. The platform that honours that contract has to come from somewhere equally reproducible — so we start at the very bottom of the stack: the substrate it all runs on.

What you'll get from this post

A single-node Kubernetes cluster on Talos Linux, running on a roughly €8.49/month Hetzner cx33 in Helsinki, Finland, provisioned entirely by OpenTofu. The operating system is immutable and has no shell; everything is configured over an API. Nothing clicked in a console, reproducible from zero, disposable on demand.

Why the operating system has no shell

The obvious way to build this is Ubuntu plus a cloud-init script that installs k3s — it works, and it's what most guides teach. But every general purpose OS under a cluster carries the same liability: a shell. The day something is weird, you SSH in "just to look," and from that moment the box and the code describing it start to diverge. "Nothing clicked" has to include "nothing SSHed."

Talos Linux closes that door completely. It's a minimal, immutable OS built to do exactly one thing — run Kubernetes. There is no shell, no SSH daemon, no package manager. The entire machine is described by one declarative config document, applied over an API (TCP 50000). To change the machine, you change the document and re-apply; to upgrade, you boot a new signed image. And the versions are pinned like any other dependency: Talos v1.13.5, which ships Kubernetes v1.36.2.

The cluster stops being a pet you tend over SSH and becomes what it should have been all along: an artifact of the code that declares it.

The wrinkle: Hetzner has no Talos image

Hetzner Cloud can't boot custom images you upload — and Talos isn't in its image catalogue in a shape you can pin. The established workaround sounds crude and works beautifully: boot a throwaway server into Hetzner's rescue system, stream the official Talos disk image straight onto its disk, snapshot that disk, delete the server. Every cluster node then boots from the snapshot.

The image itself comes from Talos's Image Factory, pinned to an exact version and schematic, and the whole build is one idempotent script driven by OpenTofu — its heart is a single pipe:

# on the rescue system of the throwaway builder (condensed)
curl -fsSL "https://factory.talos.dev/image/$SCHEMATIC_ID/v1.13.5/hcloud-amd64.raw.xz" \
  | xz -d | dd of=/dev/sda bs=4M conv=fsync status=progress

The script checks first whether a snapshot with the expected labels already exists and no-ops if so. That makes the snapshot a durable, versioned artifact: tofu destroy doesn't delete it, rebuilds reuse it, and a completely fresh account recreates it on the first apply.

The layout: a seam you'll thank yourself for

platform-infra/
  modules/
    talos-image/   # builds the bootable Talos snapshot (rescue + dd, idempotent)
    cluster/       # network, firewall, node, machine config — all over the API
  launch/          # thin caller: the launch tier is just values

The split matters more than it looks. modules/cluster is the platform layer; launch/ calls it with the launch tier's values (one cheap node). When this platform later grows a high-availability production tier, that tier calls the same module with different values — the tiers differ only in values, never in code. Workloads don't live here at all; they arrive via GitOps in the next post, on the other side of the seam.

Step 1 — a firewall with no port 22

The network setup is ordinary: a private network for future node-to-node traffic, and a cloud firewall on the public interface. What's unusual is what isn't there — port 22 doesn't exist, because there's nothing listening on it.

resource "hcloud_firewall" "cluster" {
  name = var.cluster_name

  rule {                       # Talos machine API — operator only
    direction  = "in"
    protocol   = "tcp"
    port       = "50000"
    source_ips = var.talos_allowed_cidrs
  }

  rule {                       # HTTP — public web
    direction  = "in"
    protocol   = "tcp"
    port       = "80"
    source_ips = ["0.0.0.0/0", "::/0"]
  }

  rule {                       # HTTPS — public web
    direction  = "in"
    protocol   = "tcp"
    port       = "443"
    source_ips = ["0.0.0.0/0", "::/0"]
  }
}

Port 50000 is the Talos machine API — the only management surface the node has, and it's allow-listed to the operator's CIDR because OpenTofu itself configures the node over it. The Kubernetes API (6443) stays closed by default; you can open it to your own CIDR later for direct kubectl, but as you'll see, you don't need it to get the kubeconfig.

Step 2 — the node boots from the snapshot

resource "hcloud_server" "node" {
  name         = "${var.cluster_name}-cp-1"
  server_type  = "cx33"                          # 4 vCPU, 8 GB — ~€8.49/mo
  image        = module.talos_image.snapshot_id  # the snapshot built above
  location     = "hel1"                          # Helsinki, Finland
  firewall_ids = [hcloud_firewall.cluster.id]
}

No user_data, no cloud-init, no SSH key. The firewall is bound at creation so the box is never briefly open. Talos boots from the snapshot into maintenance mode and waits to be told who it is.

Step 3 — configure over the API, not over SSH

This is where cloud-init would have been — replaced by the siderolabs/talos provider doing everything over TCP 50000:

resource "talos_machine_secrets" "this" {          # cluster PKI + tokens
  talos_version = var.talos_version
}

data "talos_machine_configuration" "controlplane" {
  cluster_name       = var.cluster_name
  cluster_endpoint   = "https://${local.node_ip}:6443"
  machine_type       = "controlplane"
  machine_secrets    = talos_machine_secrets.this.machine_secrets
  talos_version      = var.talos_version
  kubernetes_version = var.kubernetes_version
  config_patches     = local.config_patches        # install disk, certSANs,
}                                                   # allow workloads on the CP

resource "talos_machine_configuration_apply" "this" {   # push the config
  client_configuration        = talos_machine_secrets.this.client_configuration
  machine_configuration_input = data.talos_machine_configuration.controlplane.machine_configuration
  endpoint                    = local.node_ip
  node                        = local.node_ip
}

resource "talos_machine_bootstrap" "this" {         # bootstrap etcd, once
  depends_on           = [talos_machine_configuration_apply.this]
  client_configuration = talos_machine_secrets.this.client_configuration
  endpoint             = local.node_ip
  node                 = local.node_ip
}

resource "talos_cluster_kubeconfig" "this" {        # kubeconfig — over 50000,
  depends_on           = [talos_machine_bootstrap.this]   # not 6443
  client_configuration = talos_machine_secrets.this.client_configuration
  endpoint             = local.node_ip
  node                 = local.node_ip
}

The config patches do three small things: install to the Hetzner root disk from the Image Factory installer, add the node's public IP to the certificate SANs (so a remote kubectl and talosctl actually verify TLS), and allow workloads on the control plane — it's a single-node cluster, there are no workers to schedule onto.

The machine secrets, kubeconfig, and talosconfig come back as sensitive outputs. They live in encrypted state, never on the node's user_data, never in a file someone scp'd once and forgot.

Step 4 — apply and connect

export HCLOUD_TOKEN="..."     # Hetzner Cloud API token (read & write)

tofu init
tofu plan  -var-file=launch.tfvars   # review it — always
tofu apply -var-file=launch.tfvars   # 1st: builds the Talos snapshot
tofu apply -var-file=launch.tfvars   # 2nd: boots and bootstraps the node

Yes, two applies from a cold start — and I'd rather explain it than hide it. The node's image comes from a data-source lookup that can only resolve after the snapshot resource exists in state, so the first apply materialises the snapshot and stops at the node with "image not found"; the second converges. Steady-state applies are single-shot, and the snapshot build is skipped whenever the artifact already exists. It's the honest price of building a custom image during apply on a provider that can't import one.

Then verify the node — no SSH, remember:

tofu output -raw talosconfig > talosconfig && chmod 600 talosconfig
IP=$(tofu output -raw server_ipv4)
talosctl --talosconfig talosconfig -e "$IP" -n "$IP" health   # -> healthy

tofu output -raw kubeconfig > kubeconfig && chmod 600 kubeconfig
KUBECONFIG=$PWD/kubeconfig kubectl get nodes
# NAME            STATUS   ROLES           AGE   VERSION
# platform-cp-1   Ready    control-plane   2m    v1.36.2

That kubeconfig was retrieved over the Talos API — the Kubernetes port never had to be open to the internet for you to get it. (Open 6443 to your own CIDR if you want direct kubectl access; that's a one-variable change.)

Done when: kubectl get nodes shows a single Ready node — and, the real test, tofu destroy followed by re-apply rebuilds the whole thing from scratch with the same result. I've run that loop: eleven resources destroyed, two applies later the node is back, Ready, identical — same snapshot (reused, not rebuilt), fresh machine PKI, exactly as it should be.

The takeaway

The real win isn't Talos, and it isn't OpenTofu — it's the destroy-and-reapply loop. The cluster is cattle, not a pet. Nothing about it lives in your memory or your shell history, because there is no shell. And the state describing it doesn't sit on a laptop either: it lives in a remote backend on EU object storage, encrypted client-side — that setup is worth a post of its own.

A general-purpose OS gives you a thousand knobs and a shell to turn them with; every one is a place where reality can drift from the code. An immutable, API-driven OS deletes the whole category. That property — the box is the code — is what every later post quietly depends on.

What tripped me up

  • The launch box isn't the one I planned — twice over. This platform was meant to run on Hetzner's ARM instances (Ampere — cheaper, more power-efficient). When it was time to apply, Hetzner couldn't place any ARM instance in any of its three EU locations, and there's no stock API — placement is trial-only. The fix: an x86 box, container images built multi-arch (amd64 + arm64) since day one, so moving back to ARM when stock returns is a one-variable change, not a rebuild. Later, the same story played out on a different axis: growing into cx33 for observability headroom, Hetzner only had capacity in Helsinki, not Falkenstein — so the node moved region too. Because location forces a full node replacement on Hetzner (no live migration between datacentres), that one was a genuine rebuild, not just a tfvar flip.
  • The two-apply cold start described above. Inherent to building the image during apply; document it rather than fight it.
  • Snapshot labels have a 63-character limit — and the Image Factory schematic ID is 64 hex characters. The snapshot is keyed on a 16-character prefix instead. The kind of detail you only learn by hitting it.
  • Rescue mode is eventually-consistent. The rescue system's SSH takes 30–60 seconds to come up, so the builder polls with a deadline instead of connecting once and failing. And dd gets conv=fsync, because a standalone sync is unreliable there.
  • Certificate SANs. Bake the node's public IP into both the machine and API-server certs (the config patch above) or every remote talosctl and kubectl call fails with a confusing TLS error.

Where this leaves us

You now have a reproducible EU cluster that came from code and can be rebuilt at will — with no SSH surface at all. But a cluster with nothing on it is just a very quiet box. Next, it starts driving itself: GitOps with Argo CD, where the cluster continuously pulls its desired state from git and every deploy becomes a pull request. After that, a real domain, real certificates, and the first workload — the blog you're reading.


I'm building this in the open, and I'll run it for EU teams who'd rather ship product than hire a platform engineer. If that's you, follow along via RSS — every build post lands there first — or start with what this project is about.