pbx-agent
pbx-agent is the program on every PaaSbox Clusters server that keeps the cluster in the state the portal holds for it. This page lists its commands, files, protocol, operations and failure behaviour, for when you look at a node yourself or want to know exactly what the portal can ask of it.
pbx-agent is one static Go program, run by systemd on every k3s server of a PaaSbox Clusters cluster. It is not an AI agent. It runs as root, because a restore has to stop k3s and start it again while the Kubernetes API is down. It listens on no port: it calls the portal, brings the node to the desired state it receives, and runs one operation at a time from a fixed list.
Commands
Section titled “Commands”| Command | What it does |
|---|---|
pbx-agent run | The service: enroll once, start k3s, sync with the portal, apply the desired state, run operations. This is what the systemd unit runs. |
pbx-agent status | Prints pbx-agent’s state as JSON, without any secret. |
pbx-agent uninstall | Prints what it would remove and changes nothing. With --yes it removes pbx-agent (below). |
pbx-agent version | Prints the version and the IDs of the release keys built into it. |
pbx-agent simulate | Fake agents for testing the portal, up to a load test of 1,000; touches nothing on the machine it runs on. Not used on a node. |
Files on the node
Section titled “Files on the node”| Path | Mode | What it is |
|---|---|---|
/etc/pbx-agent/config.yaml | 0600 | The portal’s address, the cluster’s ID, the node’s name, and the enrollment token until it is used. Written by cloud-init. |
/etc/systemd/system/pbx-agent.service | The unit. Written by cloud-init. | |
/var/lib/pbx-agent/ | 0700 | pbx-agent’s state directory. |
/var/lib/pbx-agent/bin/pbx-agent | The program, with pbx-agent.prev (the previous version) and pbx-agent.new (during an update) beside it. | |
/var/lib/pbx-agent/identity.json | 0600 | pbx-agent’s ID and its private keys: an Ed25519 key to sign, an X25519 key to receive sealed secrets; during a rotation also the previous keys. |
/var/lib/pbx-agent/state.json | 0600 | The last applied desired state, the running operation and its step, and the IDs of finished operations (kept 30 days). |
/var/lib/pbx-agent/bootstrap.json | 0600 | The answer to the enrollment. Secrets in it stay sealed on disk. |
/var/lib/pbx-agent/ops/<id>.log | One log per operation. The last 200 lines go into a failure report. | |
/etc/rancher/k3s/config.yaml | 0600 | k3s’s configuration, written once at enrollment. |
/etc/rancher/k3s/config.yaml.d/50-pbx-*.yaml | pbx-agent’s drop-ins: the snapshot schedule and the ingress mode. | |
/var/lib/rancher/k3s/server/manifests/pbx-*.yaml | The manifests k3s applies at its first start, such as the cloud controller’s. Once pbx-agent manages that add-on through the API, it writes pbx-<name>.yaml.skip beside the file, so k3s stops applying the first version. |
The program lives under /var because /usr is read-only on the node image; /var and the /etc overlay survive image updates.
The unit:
[Unit]Description=pbx-agentWants=network-online.targetAfter=network-online.target[Service]Type=execExecStart=/var/lib/pbx-agent/bin/pbx-agent runRestart=alwaysRestartSec=5ProtectHome=yesPrivateTmp=yes[Install]WantedBy=multi-user.targetThe logs: journalctl -u pbx-agent.
The protocol
Section titled “The protocol”pbx-agent opens every connection, to https://agents.<portal domain>/agent/v1. JSON over HTTPS, bodies of at most 256 KiB.
| Endpoint | Purpose | Authenticated by |
|---|---|---|
POST /enroll | Exchange the enrollment token for an identity and the node’s configuration | The enrollment token |
POST /sync | Report the node’s state; receive the desired state and at most one operation | pbx-agent’s signature |
POST /operations/{id}/events | Report an operation’s progress and result | pbx-agent’s signature |
GET /releases/{version}/manifest | Fetch a signed release manifest | pbx-agent’s signature |
POST /rotate-key | Register new public keys, signed with the old key | pbx-agent’s signature |
Signing. Every request after the enrollment carries pbx-agent’s ID, a timestamp and an Ed25519 signature over the method, the path, the timestamp and a SHA-256 of the body. The portal refuses a clock skew of more than 60 seconds and a signature it has seen in the last 120 seconds.
Sealing. Every secret the portal sends (k3s tokens, S3 keys, Hetzner tokens, add-on secrets) is encrypted to pbx-agent’s X25519 key with a libsodium sealed box. A secret field that arrives in clear makes pbx-agent refuse the whole document (unsealed_secret).
How pbx-agent answers the portal’s status codes
Section titled “How pbx-agent answers the portal’s status codes”| Answer | pbx-agent |
|---|---|
| 200 | applies it |
| 400 | logs it and backs off; never sends the same request more than three times. A replayed answer (a restarted pbx-agent signed the same request in the same second) is signed once more with the next second. |
| 401 | treats its identity as revoked: stops syncing, leaves the cluster alone, keeps running idle |
409 clock_skew | corrects its clock offset from the portal’s time and retries |
426 protocol_too_old | updates itself to the pbx-agent of the release the portal names, whatever the desired state says |
| 429 | waits as long as the portal says |
| 5xx, no network | backs off exponentially from 5 seconds up to 5 minutes, with jitter |
Enrollment
Section titled “Enrollment”- On first start
pbx-agentcreates its keys and stores them before it uses them. - It reads the server’s ID from Hetzner’s metadata service.
- It sends the enrollment token, its public keys, and the hash of the boot image the node actually booted. The portal checks the token (unused, unexpired, bound to this cluster and node name), checks through the Hetzner API that the server carries the cluster’s label and the node’s name, and compares the boot image’s hash with the release.
- It stores the answer and deletes the token from its configuration.
- It writes k3s’s configuration, starts k3s as a server, waits for the API, and writes the Secrets
kube-system/hcloudandkube-system/pbx-etcd-s3.
| The portal answers | Meaning | pbx-agent |
|---|---|---|
| 200 | identity and configuration | goes on |
401 invalid_token, token_used, token_expired, server_mismatch | the token is refused | tries at most three times, then idles; a new token means a new server |
409 image_mismatch | the node booted another image than its release’s | idles; the node needs a new server from the right image |
409 conflict | the configuration cannot be built yet | retries with backoff |
503 internal | the portal cannot reach the Hetzner API | retries with backoff |
| 429 | too many requests | waits as long as the portal says |
The same token with the same keys from the same server gets the same answer again, so an answer lost on the way is asked for again safely.
The sync
Section titled “The sync”About every 30 seconds (± 5) pbx-agent sends its report and receives the desired state. Each cycle it converges in this order: its own update, the host’s configuration (k3s drop-ins), the cluster’s configuration (Secrets, add-ons), then the operation. The cluster’s configuration is applied again every 10 minutes even without a change, so a change by hand to a managed object is overwritten.
A desired state is applied as a whole. Its release manifest is fetched and verified first; a manifest that no built-in key verifies changes nothing (manifest_rejected).
A drop-in that needs a restart of k3s is written at once; the restart happens at the next moment inside the maintenance window, never while an operation runs. In progress This has not run on a real node yet.
On a cluster with more than one server, one pbx-agent leads: the holder of the Lease pbx-system/pbx-agent-leader. On a single node its pbx-agent leads while the Kubernetes API is up.
Add-ons
Section titled “Add-ons”pbx-agent applies each add-on from the template in the signed release, with the options from the desired state.
| What | How |
|---|---|
| Objects it may create | HelmChart and HelmChartConfig objects, Secrets named pbx-addon-* in kube-system, and for the PaaSbox Platform one cluster-wide Platform named after the cluster |
| Labels | pbx.io/managed=true on every managed object, pbx.io/addon=<name> on an add-on’s objects |
| Secret options | Only in a pbx-addon-<name> Secret that the chart reads, never in a HelmChart: k3s encrypts Secrets at rest, not custom resources |
| Switched off | Its objects are deleted; data, volumes and CRDs stay. The Secrets hcloud and pbx-etcd-s3 are never deleted. |
| A template or an apply that fails | The add-on reports failed and keeps the objects it has: a failed render never uninstalls a running chart |
Each add-on reports one state: applied, pending, failed or removed.
| Case | State |
|---|---|
| An add-on it requires is off | failed: requires add-on <r>, which is off |
| A conflicting add-on is on and installed | failed: conflicts with add-on <c>; switch that off first |
| A conflicting add-on is off, its objects still exist | pending until they are gone |
An add-on it requires is not applied yet | pending |
| Applied, its health check not passing yet | pending; pbx-agent applies the cluster’s configuration at every cycle until it passes |
| Applied, its health check passes, or it has none | applied |
| Switched off while an add-on that requires it is on | failed: still required |
Built in a lab: the PaaSbox Platform add-on is switched off in two steps. While a SaaSApplication exists, switching it off changes nothing: the add-on reports failed with the applications’ names. Delete them first. Then pbx-agent applies a Platform that installs nothing and waits until Flux has removed what the platform installed. A Platform that refuses the removal (a database, or an application that still needs a part of it) keeps everything and reports the add-on failed with its message.
Operations
Section titled “Operations”One operation runs per cluster at a time. Each is a list of named steps that can be repeated safely; pbx-agent records the step before it starts it and resumes there after a crash or a reboot. A failed step reports its name and the last 200 lines of its log. Destructive steps are never retried by pbx-agent on its own.
| Operation | Steps | What it does |
|---|---|---|
snapshot.save | save, confirm_s3, report | Takes an etcd snapshot and confirms it reached the bucket |
cluster.restore | stop_k3s, reset_restore, start_k3s, wait_api, forget_stale_nodes | Resets the cluster’s state to a snapshot; see Take and restore snapshots |
cluster.upgrade | preflight, pre_snapshot, fetch_image, stage_slot, reboot, health_gate, commit | Updates the node image in place; see Upgrade a cluster |
access.issue | ensure_sa, token, seal | Issues a kubeconfig sealed to the key of whoever asked: your browser, or an agent’s own key |
access.revoke | delete_sa | Deletes a person’s ServiceAccounts, or everyone’s |
agent.rotate_key | generate, register, confirm | Replaces pbx-agent’s own keys; the portal asks for it when they are 90 days old |
s3.verify | list_bucket | Lists the bucket with new keys before they are used |
diag.collect | collect | Versions, unit states and the last lines of the logs of k3s and pbx-agent, only with your consent for that request; never Secrets, tokens or environment files |
server.rejoin, node.drain, node.forget | For clusters with more than one server; not used on a single node |
How the steps decide, where it matters:
reset_restoreputs the snapshot on the node itself, from k3s’s local copy or fetched from the bucket withpbx-agent’s own S3 client, unpacks it and runsk3s server --cluster-resetwith--etcd-s3=false. The S3 keys never reach a command line or k3s’s environment.forget_stale_nodesdeletes a Node other thanpbx-agent’s own whose kubelet has not reported since the restored API came up, after a grace of 2 minutes.cluster.upgradetakes the snapshotpbx-pre-upgrade-<release>, stages the new image as the boot entrypbx-<release>-<arch>.efiand keeps two entries, the new one and the one it came from. The health gate waits for the API, the NodeReadyon the new k3s, the new boot entry, and every Deployment inkube-systemavailable. A node that is not healthy within 20 minutes of the reboot is rebooted into the old entry. In progress That return has not fired on a real node yet.reset_restoreandrebootare the destructive steps. After a crash inside one,pbx-agentchecks what the node shows: not started, it runs; completed, it goes on; anything else isfailedand never run again on its own.
An unknown operation is answered as unsupported. There is no operation that runs a command you or anyone else types.
With the Kubernetes API down, pbx-agent accepts only cluster.restore and diag.collect.
When something fails
Section titled “When something fails”| Situation | pbx-agent |
|---|---|
| The portal cannot be reached | keeps the last desired state and retries with backoff; k3s keeps taking the scheduled snapshots |
| The Kubernetes API is down | reports it and accepts only host-level operations: restore and diagnostics |
| An operation step fails | stops, reports the step and the last 200 log lines; never retries a destructive step on its own |
| A release manifest’s signature is invalid | refuses it, reports manifest_rejected, changes nothing |
| A secret arrives in clear | refuses the whole document, reports unsealed_secret, changes nothing |
pbx-agent is revoked | stops syncing and leaves the cluster as it is |
Its own updates
Section titled “Its own updates”When the desired state names another version of pbx-agent, it fetches the release manifest, takes the program for its architecture, checks its SHA-256, keeps the running program as pbx-agent.prev, arms a rollback timer of 10 minutes, swaps the programs and restarts. The first successful sync with the portal disarms the timer. If it fires, the previous program is put back. A pbx-agent that finds itself rolled back reports it and does not try that version again for six hours.
pbx-agent trusts only release manifests signed with one of the two release keys built into it. Nothing at run time can change those keys. A pbx-agent built without them refuses every manifest: it never updates itself and never upgrades the node. Releases and channels has the signing.
Remove it
Section titled “Remove it”pbx-agent uninstall --yesDisables and removes the unit and the rollback timer, and removes /var/lib/pbx-agent and /etc/pbx-agent. k3s, its configuration and every Kubernetes object stay. Without --yes it prints what it would do and changes nothing. In progress Not run on a real node yet. Detach or delete a cluster says when you use it.