Skip to content

Why isolated clusters for agents

PaaSbox is built on one idea about agents and DevOps: a change is tested in a cluster that exists only for the test, and every stage of every app runs on a small cluster of its own. This page explains why, what it costs, and what it does not protect against.

One large cluster, sharedSmall clusters, one per stage, app and test
An upgrade that failsreaches every stage and app on itreaches one stage of one app
A test that runs awaycompetes with production for the same nodeshas a server to itself, and is deleted
A change to something cluster-wide (a CRD, an operator, a default)applies to everyone at onceapplies to one cluster
Kubeconfigs and API rangesone cluster’s, divided by namespaces and RBACeach cluster’s own
Upgrade timingone window for alla channel and a window per cluster
A server that failsthe other servers carry onthat cluster is down until it is restored or rebuilt
Upgrades to scheduleone clusterone per cluster

An agent’s change can be wrong in ways that a test inside a shared cluster cannot contain: a CRD replaced with an older version, a webhook that rejects every Pod, a job that fills the disk. In a cluster built for the test, such a mistake breaks only that cluster, and the cluster is deleted afterwards.

A new cluster also starts from nothing but your files. If the app comes up there, the files are complete; if it needs something that was set up by hand somewhere else, the test shows it.

This only works if a cluster is cheap in time. In the lab of 2026-10-10 a cluster was ready 2 minutes 46 seconds after the request, 69 seconds of it the copy of the node image into a project that had none, and a delete took 14 seconds, its volumes included, with no labelled resource left in the project. In progress The REST API and the MCP tools that let an agent run this loop itself have not run against a real cluster yet. Learn step 3 walks through it.

A stage of an app gets a single-node cluster of its own: its own server, private network, firewall, Kubernetes API name, kubeconfigs and snapshots. Three things follow.

  • A mistake stays small. An upgrade, a restore or a runaway test reaches one stage of one app. A restore in place takes one cluster back to its snapshot, 160 seconds in the lab, and no other cluster is part of the operation.
  • Staging goes first. Each cluster has its own channel and maintenance window. Staging can take each release on early and be checked, while production waits on stable for its own window.
  • Access follows the stage. A kubeconfig for staging opens staging. The API of production can be open to a smaller set of addresses than that of a test cluster.

Kubernetes is what makes these environments repeatable: an app with its database, certificates and configuration, created from the same files every time. Planned ownpaas is to bring the same environments to a local lab on your own machine, before a Hetzner server is involved.

  • Money. Each cluster costs €29 incl. VAT a month, plus its server at Hetzner. A cluster counts from the moment it is first ready until you delete it. In progress The billing through Paddle is being built; as written, a cluster added next to others is charged prorated at once, and one removed is credited for the rest of the period, so a throwaway cluster next to a cluster you keep costs its share of the month. The team’s last cluster is not credited: the subscription runs to the end of the period already paid, and a cluster created before then uses it without a new charge. After that end, the team’s next cluster starts a new subscription in the portal’s checkout, which an agent cannot do through the API. Hetzner bills the server by the hour, at its own price. Costs and billing has the details.
  • A limit. A team may have 10 clusters at once, and staging, production and every running test count against them: two apps with staging and production leave six for tests. Ask for a higher limit when your tests need it.
  • Upgrades to watch. Every cluster takes every release, and every upgrade reboots its node: in the lab the Kubernetes API was down for 34 seconds. With a cluster per stage and app there are more upgrades to schedule and check. The recipe An agent upgrades staging first hands that to an agent.
  • A failing server. A cluster is one server. If it fails, the stage on it is down until the server is back, a snapshot is restored, or a new cluster is built. There is no SLA, and clusters with three servers are Planned.
  • A key that reaches too far. An API key’s scopes hold for every cluster of the team: a key that may delete a test cluster may also delete production. The typed cluster name stops a mix-up, not a decision. Keep production in a team the key does not belong to. Guardrails for agents lists what each limit stops.
  • A shared Hetzner project. The cloud controller inside a cluster holds a Hetzner token, and a Hetzner token opens its whole project. Clusters in one project are separate in Kubernetes, not in Hetzner: give production a project of its own, and test clusters another.
  • Missing data. A throwaway cluster has the data your test puts there, not production’s. A snapshot holds the cluster’s state, not the contents of your volumes; a database needs backups of its own.