Troubleshooting
This page goes from a symptom to its usual causes and what to do: a cluster that stops on its way to ready, pbx-agent that does not enroll or stops reporting, an operation that fails, a credential that is refused. If none of it helps, write to me with the cluster’s name and the failed step or operation.
The cluster stops at a step and shows Failed
Section titled “The cluster stops at a step and shows Failed”The cluster’s page says Stopped at step …, with the step’s name and whether it was creating (provisioning) or deleting the cluster, and Hetzner’s or the portal’s message below. The portal deletes nothing on its own, so you can look at the resources in your project. A team admin’s Retry from … resumes at that step once the cause is fixed. Each step has a time limit; a step that runs over it fails with “step … timed out after …”.
| Step | Usual cause | What to do |
|---|---|---|
validate | The project’s Hetzner token no longer works, the project’s connection is suspended, the location is unknown, the channel offers no release, or the release has no image for the server’s architecture | Replace the token (Rotate credentials), then retry; for the architecture or the location, delete the cluster and create it with another server type or location |
image | Copying the node image into the project failed or took longer than an hour | Retry. Only the first cluster in a project waits for this step |
network, firewall, primary_ip | Hetzner refused to create the resource, for example because the project reached one of its limits | Ask Hetzner for a higher limit, or free what the project holds, then retry |
dns | The DNS service refused the API’s name, or the cluster has no public IPv4 | Retry; if it stops again, write to me |
server | Hetzner refused to create the server, for example because the type is not available in the location at the moment, or the project reached its server limit | Retry later, ask Hetzner for a higher limit, or delete the cluster and create it in another location |
enrollment | pbx-agent did not enroll within 10 minutes; see below | Retry |
api_healthy | k3s did not report a healthy API within 10 minutes | Retry; if it stops again, write to me |
A server whose enrollment token expired unused is replaced on the retry by a new one with fresh user data. An adopted server is rebuilt again instead.
A delete that stops shows the same message with deleting, and Retry from … resumes it; each pass deletes what it still finds by the cluster’s label (Detach or delete a cluster).
pbx-agent does not enroll
Section titled “pbx-agent does not enroll”pbx-agent enrolls with a single-use token that expires 30 minutes after the server was created.
- The server cannot reach the portal.
pbx-agentneeds outgoing HTTPS. The cluster’s firewall has rules for incoming traffic only; an outgoing rule you added in the Hetzner console may block it. - The token was refused (
invalid_token,token_used,token_expired,server_mismatch).pbx-agenttries three times and then idles. A server that never enrolled is replaced on the retry. - The node booted another image than its release’s (
image_mismatch). The portal compares the hash of the image the node booted with the release, and refuses.pbx-agentidles. This should not happen with the image the portal copied into your project; write to me with the cluster’s name.
Before the enrollment there is no Kubernetes API to start a pod with, so Hetzner’s rescue system is the way onto the server (Get root on the server). On a running node, journalctl -u pbx-agent shows what pbx-agent tried, and pbx-agent status its state.
pbx-agent is not reporting
Section titled “pbx-agent is not reporting”The cluster list shows agent not reporting when pbx-agent has not called the portal for five minutes. After 24 hours the team’s admins get a mail. The cluster itself is not affected: k3s keeps running and keeps taking the scheduled snapshots.
- The server is off or unreachable. Check it in the Hetzner console.
pbx-agentwas removed withpbx-agent uninstall, or its service stopped:systemctl status pbx-agent.pbx-agentwas revoked by a detach: it stops calling and leaves the cluster as it is. That is expected.- The clock is off. The portal refuses requests whose timestamp is more than 60 seconds away from its own time.
pbx-agentthen takes the portal’s time from the answer and tries once more. If it keeps happening, check the server’s clock withtimedatectl.
The Kubernetes API does not answer
Section titled “The Kubernetes API does not answer”- From your computer: your address may not be in the ranges allowed to reach port 6443. Settings shows them, and a team admin can add yours there from anywhere; the portal is not behind the cluster’s firewall (Change who can reach the API).
- On the cluster: while an upgrade reboots the node or a restore runs, the API is down by design. Otherwise the cluster’s page shows whether
pbx-agentreports a healthy API. - While the API is down,
pbx-agentaccepts only a restore and diagnostics. Other operations wait at the portal until the API is back.
A kubeconfig does not arrive
Section titled “A kubeconfig does not arrive”- The cluster is not running. The Access page says that kubeconfigs need a running cluster.
- The API is down. The request waits at the portal; after 5 minutes it expires. Ask again once the API is back.
- “The credential was already handed out or has expired.” The portal hands the encrypted kubeconfig over once and keeps it for 5 minutes at most. Ask for a new one.
Get a kubeconfig has the whole flow.
A new token or new keys are refused
Section titled “A new token or new keys are refused”- A new project token: “This token does not see the servers of …”. The token is from another Hetzner project. Create it in the project the cluster lives in.
- A new token inside the cluster: “This token does not see the cluster’s servers; is it for the same project?” The same cause.
- New bucket keys stay “Waiting for the agent to list the bucket”.
pbx-agenttries them on the bucket at its next sync. If the check failed, the list under Bucket keys showss3.verifyas failed with the S3 provider’s message, such asInvalidAccessKeyId, and the keys in force stay. Check the pair and save it again.
Rotate credentials has the steps.
An upgrade failed
Section titled “An upgrade failed”The Upgrades page shows the failed step, and the team’s admins get a mail.
- If the node did not come back healthy within 20 minutes of the reboot, or came up on the new image with the wrong k3s version,
pbx-agentrebooted it into the previous image. The cluster runs the release it ran before. In progress This return has not been triggered on a real node yet. - If the node came up on the previous image, the new one did not boot, and the upgrade failed at once.
- If only the last step (
commit) failed, the node runs the new image, but its next reboot would start the old one. The upgrade’s message says so. - The snapshot taken before the upgrade, named
pbx-pre-upgrade-<release>, is in your bucket. Restore it only if the cluster’s state is damaged; a failed upgrade alone does not need it. - An upgrade that misses its window, because
pbx-agentwas not reachable, moves to the next window and expires after 14 days.
A snapshot is missing
Section titled “A snapshot is missing”- If a scheduled snapshot is missed twice in a row, the team’s admins get a mail.
- No bucket configured: the snapshots stay on the server’s disk only. Backups shows the target.
- The keys stopped working: with keys the portal holds, save new ones under Backups → Bucket keys (above). With keys you hold, check your Secret
kube-system/pbx-etcd-s3. - The bucket’s lifecycle rules may have deleted snapshots.
A restore did not finish
Section titled “A restore did not finish”The restore’s failed step and its log are on the cluster’s Operations page. A restore is never repeated on its own after a failure in its reset step: pbx-agent reads the node’s state and reports it, and you decide. With keys you hold, pbx-agent must read them from the cluster before it stops k3s, so such a restore needs a running API. Take and restore snapshots.
An add-on shows failed
Section titled “An add-on shows failed”The Add-ons page shows the state the cluster reports for each add-on, with the reason: a missing required option, a token Hetzner refuses, an object Kubernetes refuses, an add-on it requires that is off. Fix the option and save. A failed add-on keeps the objects it already has. Choose add-ons.
The PaaSbox Platform says failed: not switched off: … when you switch it off while a SaaSApplication or a database still exists: delete them first (Choose add-ons). If it stays pending instead, a part of it may have failed once: Flux tries such a part again only after an hour. Learn step 2 shows how to ask Flux to try it now.
A change made by hand disappeared
Section titled “A change made by hand disappeared”Objects an add-on manages carry the label pbx.io/managed=true, and pbx-agent applies them again every ten minutes. To control one yourself, switch the add-on off and install your own (Choose add-ons). A rule you added to the cluster’s firewall in the Hetzner console is gone after the API ranges are saved on Settings (Change who can reach the API).
A server cannot be adopted
Section titled “A server cannot be adopted”The create form shows the reason next to each server that cannot be adopted. Use a server you already have explains each requirement. Among them: a CX server (it boots with BIOS only), rebuild protection on, or a firewall that applies to the server through a label selector.
A cluster cannot be created
Section titled “A cluster cannot be created”The create page names the first thing in the way: no Hetzner project connected, no release published yet, the team’s cluster limit reached (“Delete or detach one, or ask for a higher limit”), an unpaid invoice or an ended subscription, the terms not yet accepted, or the subscription not yet started In progress. Only team admins create clusters. Create a cluster.
Diagnostics
Section titled “Diagnostics”pbx-agent can collect diagnostics, versions, the state of its units and the last lines of the k3s and pbx-agent logs, but only after you allowed it for that one request. They never include Secrets, tokens or environment files. Planned The portal page where you give that consent is not built yet; until then, journalctl -u k3s and journalctl -u pbx-agent on the server show the same logs.