Upgrade a cluster
An upgrade moves a cluster to a newer release: a tested, signed set of the node image with its k3s, pbx-agent and the add-ons. This page starts one, now or in the cluster’s maintenance window, sets the window and the channel, and says what happens when an upgrade fails.
On a single node every upgrade reboots the server. The Kubernetes API and your apps are down while it restarts, until their pods have started again. Choose the moment, or let the window choose it.
What you need
Section titled “What you need”- To be a team admin. Members see the page but cannot start an upgrade.
- A cluster in the state ready, with no other upgrade queued or running.
- A release offered for it. The Upgrades page says “No newer release is offered for this cluster on its channel” when there is none.
Start an upgrade
Section titled “Start an upgrade”-
On the cluster’s page, open Upgrades. It shows the release the cluster runs and the next release offered on its channel, with its notes and a badge: patch: same k3s minor, or new k3s minor.
-
Choose when:
- In the next window, with the window’s start next to it, or
- Now:
pbx-agenttakes it at its next sync, about 30 seconds later.
-
Watch it. The page shows the upgrade as queued, then each step as
pbx-agentreports it; Operations has the full log. When it is done, the page shows the new release, and Upgrade history keeps the run. -
Check your apps.
A cluster is offered only a release that names the cluster’s current one among the releases it was tested from. An upgrade queued for the window that misses it, because pbx-agent was not reachable in time, moves to the next window; after 14 days it expires.
What happens during an upgrade
Section titled “What happens during an upgrade”k3s is part of the node image, and the image’s root file system is read-only. So an upgrade of k3s, of the operating system or of both is the same thing: the whole image is replaced, in place, on the same server. The server is not rebuilt, and the data on its disk stays.
- Preflight: the node is healthy, has room on its disk, and the release’s signature verifies.
- Snapshot:
pbx-agenttakes the snapshotpbx-pre-upgrade-<release>and confirms it reached your bucket. - Fetch: it downloads the new image and checks it against the signed release.
- Stage: it writes the new image next to the running one, as a second boot entry.
- Reboot: it boots the new image once.
- Health gate: it waits for the Kubernetes API, for the node to be Ready on the new k3s and running the new entry, and for every Deployment in
kube-systemto be available. - Commit: the new image becomes the default. The node keeps two entries: the new one and the one it came from.
In the lab run of 2026-10-10 these steps took 83 seconds from scheduling to success, Garden Linux 2150.6.0 to 2150.11.0, and the API was down for 34 seconds while the node rebooted. Before that, on 2026-09-24, the image’s own test passed 24 of 24 checks for an update in place and a return to the previous image, in a virtual machine: the volumes on the node’s disk kept their identity, and a Postgres table and a file read back unchanged after both.
When an upgrade fails
Section titled “When an upgrade fails”- If the node is not healthy within 20 minutes of the reboot,
pbx-agentreboots it into the previous image, and the upgrade is reported as failed. - If the node comes up on the new image with the wrong k3s version,
pbx-agentdoes not wait: it reboots into the previous image at once, and the upgrade fails. - If the node comes up on the previous image, the new one did not boot; the upgrade fails at once, with nothing to reboot.
- In progress These returns to the previous image have not been triggered on a real node yet.
- If only Commit fails, the node runs the new image, but its next reboot would start the old one. The upgrade’s message says so.
In progress The team’s admins get a mail about a failed upgrade. The snapshot from step 2 stays in your bucket; a failed upgrade alone does not need a restore. Troubleshooting has the details.
The maintenance window
Section titled “The maintenance window”On the cluster’s Settings, under Maintenance window, choose the Days, From, To and the Time zone, then Save. The default is Saturday and Sunday, 02:00 to 05:00, Europe/Berlin. The days are the days a window starts on; an end at or before the start runs past midnight. The window keeps its wall-clock times across daylight-saving changes, so on those nights it can be an hour shorter or longer.
In progress Three things wait for the window:
- an upgrade you scheduled In the next window,
- a patch release that installs without a click (below),
- a restart of k3s that a settings change needs, such as a new snapshot schedule.
pbx-agentwrites the change at once and restarts k3s at the next moment inside the window, never while an operation runs.
Channel and patch releases
Section titled “Channel and patch releases”Under Upgrades → Channel, choose the Release channel and whether to Install patch releases in the window without a click, then Save. The portal checks the channel’s release every hour.
- early gets a new release first; stable gets it after it ran on early. Before a release reaches early, it runs in a lab on Hetzner and then on my own clusters for three days; it stays on early for seven days before it reaches stable (Releases and channels).
- A patch release keeps the cluster’s k3s minor version. With the box ticked (the default), it is queued for the next window on its own; untick it, and every release waits for your click.
- A new k3s minor version always waits for your click.