Skip to content

Run without PaaSbox

A cluster of PaaSbox Clusters runs entirely in your Hetzner project. The portal only ever answered calls from pbx-agent on your server; it never ran your cluster. When you detach, or if PaaSbox disappears, the cluster keeps running, and this page is what you need to run it on your own.

Keeps running:

  • k3s and your workloads, the server, its private network and its firewall.
  • The scheduled snapshots. k3s takes them itself and uploads them to your bucket, as long as the keys in the Secret kube-system/pbx-etcd-s3 stay valid.

Stops:

  • Upgrades of the node image and of k3s.
  • Kubeconfigs from the portal, restores from the portal, and the mails about missed snapshots, failed upgrades and expiring certificates.
  • pbx-agent’s work. Once revoked, it changes nothing more. While it merely cannot reach the portal, it keeps applying the last settings it received, so remove it (below) before you change the add-ons’ objects by hand. After that they are yours to change.
WhatWhere
Root on the serverSSH with a key, a pod on the node, or Hetzner’s rescue system: Get root on the server
A permanent admin kubeconfig/etc/rancher/k3s/k3s.yaml on the server
The k3s server token/etc/rancher/k3s/config.yaml on the server, key token; k3s also keeps it in /var/lib/rancher/k3s/server/token
The snapshotsyour bucket, and k3s’ local copies in /var/lib/rancher/k3s/server/db/snapshots/ on the server
The k3s configuration/etc/rancher/k3s/config.yaml and the drop-ins in /etc/rancher/k3s/config.yaml.d/
The node imagea snapshot in your Hetzner project, labelled pbx-image

/etc on the node is an overlay kept on the persistent /var, so changes you make there survive reboots and image updates.

If you choose the moment, do these first, while the portal still works:

  1. Get root access of your own. Add your SSH key and open port 22 for your address with a firewall of your own, as Get root on the server shows. Log in once to check it.

  2. Copy the admin kubeconfig to your computer:

    Terminal window
    scp root@<server IPv4>:/etc/rancher/k3s/k3s.yaml ./admin.kubeconfig

    Change its server: line from https://127.0.0.1:6443 to the cluster’s API name, and keep the file safe: it does not expire.

  3. Keep a copy of the server token, from the token line of /etc/rancher/k3s/config.yaml, in your password manager. A restore onto another server needs it.

  4. Take a snapshot now in the portal and check that it is in your bucket.

  5. Set up a DNS name of your own (below) while the PaaSbox name still answers.

Detach in the portal under Settings → Detach. Then, on the server:

Terminal window
pbx-agent uninstall --yes

It removes pbx-agent’s unit, its rollback timer, /var/lib/pbx-agent and /etc/pbx-agent, and leaves k3s, its configuration and every Kubernetes object alone. The namespaces pbx-access (the ServiceAccounts of the portal’s kubeconfigs) and pbx-system (pbx-agent’s lease) stay; delete them with kubectl delete namespace pbx-access pbx-system when you no longer need them.

The portal held tokens and keys for your project and your cluster. Replace them, so that its copies open nothing.

  • The Hetzner token you gave the portal. Revoke it in the Hetzner console under Security → API tokens. The cluster does not use it, with one exception: if you chose A copy of the project’s token for inside the cluster, the cluster uses that same token. Then replace the cluster’s token first (next point).

  • The Hetzner token inside the cluster. Create a new token of the same project, write it into the Secret kube-system/hcloud (key token; keep the key network), and restart the workloads that read it:

    Terminal window
    kubectl -n kube-system create secret generic hcloud \
    --from-literal=token=<new token> \
    --from-literal=network=$(kubectl -n kube-system get secret hcloud -o jsonpath='{.data.network}' | base64 -d) \
    --dry-run=client -o yaml | kubectl apply -f -
    kubectl -n kube-system get deploy,ds | grep hcloud
    kubectl -n kube-system rollout restart deployment/<each one listed> # and daemonset/<…>

    Then revoke the old token in the Hetzner console.

  • The S3 keys, if the portal held them. Create new keys at your S3 provider, put them into the Secret kube-system/pbx-etcd-s3 (keys etcd-s3-access-key and etcd-s3-secret-key), and revoke the old ones. k3s reads the Secret at every snapshot; no restart is needed.

  • The DNS tokens of the add-ons that use one: cert-manager, external-dns, and the PaaSbox Platform with wildcard certificates. Each is in the values file inside the add-on’s Secret, kube-system/pbx-addon-<name>. Replace the token there, and revoke the old one where you created it.

  • The k3s server token was held by the portal as well. k3s can replace it with k3s token rotate; snapshots taken before that need the old token to restore, so keep it.

The API’s name, <label>.<team>.k3s. in PaaSbox’s zone, stays for 30 days after a detach. If PaaSbox disappears, it goes at once. To move to a name of your own:

  1. In your own DNS zone, create an A record for a name such as k8s.example.com, pointing at the server’s public IPv4 (the Hetzner console shows it).

  2. On the server, add the name to the API’s certificate with a drop-in, /etc/rancher/k3s/config.yaml.d/60-own-name.yaml:

    tls-san+:
    - k8s.example.com

    The + appends to the names already configured instead of replacing them.

  3. Restart k3s: systemctl restart k3s. The Kubernetes API is down for the restart; your workloads keep running, because a restart of k3s leaves the containers alone. k3s adds the new name to its certificate.

  4. In your kubeconfigs, change the server: line to https://k8s.example.com:6443.

The firewall still opens port 6443 only to the ranges you chose. The portal no longer changes it; you edit its rules in the Hetzner console.

  • The schedule and the number kept are in /etc/rancher/k3s/config.yaml.d/50-pbx-backup.yaml. Edit it and restart k3s to change them.
  • The bucket and its keys are in the Secret kube-system/pbx-etcd-s3.
  • To take a snapshot by hand: k3s etcd-snapshot save --name manual. With the bucket configured, k3s uploads it there too. k3s etcd-snapshot list lists the snapshots k3s knows.
  • The data in your volumes was never part of a snapshot; your own backup of it continues as before.

This resets the cluster’s state to a snapshot on the same server, as pbx-agent’s restore does. Everything stored in the cluster after the snapshot is gone; the data in your volumes stays. As root on the server:

  1. Choose the snapshot. k3s etcd-snapshot list, or list your bucket. Names look like <name>-<node>-<unix time>.zip.

  2. Put it on the node, unpacked. k3s cannot restore a compressed snapshot from a path; I measured that on k3s v1.36.4. Unpack it on your computer: take k3s’ local copy from /var/lib/rancher/k3s/server/db/snapshots/ on the server, or download it from your bucket with any S3 client. Then copy the plain file back into that directory:

    Terminal window
    scp root@<server IPv4>:/var/lib/rancher/k3s/server/db/snapshots/<name>.zip .
    unzip <name>.zip
    scp ./<plain file> root@<server IPv4>:/var/lib/rancher/k3s/server/db/snapshots/
  3. Stop k3s.

    Terminal window
    systemctl stop k3s
  4. Reset the cluster’s state to the snapshot.

    Terminal window
    k3s server --cluster-reset \
    --cluster-reset-restore-path=/var/lib/rancher/k3s/server/db/snapshots/<plain file> \
    --etcd-s3=false

    --etcd-s3=false is needed because the configuration reads the bucket from a Kubernetes Secret, which k3s refuses to use during a restore. k3s exits when the reset is done, with a message to restart without --cluster-reset. Its exit status is not a reliable verdict: check instead that /var/lib/rancher/k3s/server/db/reset-flag was written just now; stat /var/lib/rancher/k3s/server/db/reset-flag shows the time.

  5. Start k3s and wait for the API:

    Terminal window
    systemctl start k3s
    kubectl --kubeconfig /etc/rancher/k3s/k3s.yaml get nodes

The server token in config.yaml must be the one that was in force when the snapshot was taken. On the same server, without a token rotation in between, it is.

Planned A step-by-step guide is planned. In outline: a new k3s server of the same or a newer k3s version runs the same reset with the old server token (--token=<server token>), before it first starts as a server. On the node image, the server also needs the k3s configuration pbx-agent would have written. The data on the old server’s disk does not come back.

The node image’s root file system is read-only and k3s is part of it, so k3s cannot be upgraded on its own. What exists today:

  • Stay on your release. It keeps running. Restart k3s at least once a year: k3s renews its certificates when they are within 120 days of expiry at start. k3s certificate check shows the dates.
  • Update the image the way pbx-agent does. The image boots through systemd-boot from the server’s EFI partition and keeps two entries. An update writes the new image file next to the running one, boots it once (bootctl set-oneshot), and makes it the default (bootctl set-default) when the node is healthy. That update and the way back passed 24 of 24 checks in a lab, in a virtual machine (2026-09-24). Planned A guide for doing it by hand, with where the image files are, is planned.
  • Move. Build a cluster elsewhere and move your workloads, with the snapshot or with your own deployment files.

Everything above applies, with three differences:

  • The API’s DNS name goes at once. Until your own name works, reach the API by the server’s address and tell kubectl which name the certificate carries: kubectl --server https://<server IPv4>:6443 --tls-server-name <the old API name>.
  • pbx-agent keeps trying to reach the portal, waiting up to five minutes between attempts, and keeps applying the last settings it received until you remove it.
  • The portal’s copies of your tokens and keys are out of your control, so replacing them is the first thing to do.