Skip to content

pbx-agent

pbx-agent is the program on every PaaSbox Clusters server that keeps the cluster in the state the portal holds for it. This page lists its commands, files, protocol, operations and failure behaviour, for when you look at a node yourself or want to know exactly what the portal can ask of it.

pbx-agent is one static Go program, run by systemd on every k3s server of a PaaSbox Clusters cluster. It is not an AI agent. It runs as root, because a restore has to stop k3s and start it again while the Kubernetes API is down. It listens on no port: it calls the portal, brings the node to the desired state it receives, and runs one operation at a time from a fixed list.

CommandWhat it does
pbx-agent runThe service: enroll once, start k3s, sync with the portal, apply the desired state, run operations. This is what the systemd unit runs.
pbx-agent statusPrints pbx-agent’s state as JSON, without any secret.
pbx-agent uninstallPrints what it would remove and changes nothing. With --yes it removes pbx-agent (below).
pbx-agent versionPrints the version and the IDs of the release keys built into it.
pbx-agent simulateFake agents for testing the portal, up to a load test of 1,000; touches nothing on the machine it runs on. Not used on a node.
PathModeWhat it is
/etc/pbx-agent/config.yaml0600The portal’s address, the cluster’s ID, the node’s name, and the enrollment token until it is used. Written by cloud-init.
/etc/systemd/system/pbx-agent.serviceThe unit. Written by cloud-init.
/var/lib/pbx-agent/0700pbx-agent’s state directory.
/var/lib/pbx-agent/bin/pbx-agentThe program, with pbx-agent.prev (the previous version) and pbx-agent.new (during an update) beside it.
/var/lib/pbx-agent/identity.json0600pbx-agent’s ID and its private keys: an Ed25519 key to sign, an X25519 key to receive sealed secrets; during a rotation also the previous keys.
/var/lib/pbx-agent/state.json0600The last applied desired state, the running operation and its step, and the IDs of finished operations (kept 30 days).
/var/lib/pbx-agent/bootstrap.json0600The answer to the enrollment. Secrets in it stay sealed on disk.
/var/lib/pbx-agent/ops/<id>.logOne log per operation. The last 200 lines go into a failure report.
/etc/rancher/k3s/config.yaml0600k3s’s configuration, written once at enrollment.
/etc/rancher/k3s/config.yaml.d/50-pbx-*.yamlpbx-agent’s drop-ins: the snapshot schedule and the ingress mode.
/var/lib/rancher/k3s/server/manifests/pbx-*.yamlThe manifests k3s applies at its first start, such as the cloud controller’s. Once pbx-agent manages that add-on through the API, it writes pbx-<name>.yaml.skip beside the file, so k3s stops applying the first version.

The program lives under /var because /usr is read-only on the node image; /var and the /etc overlay survive image updates.

The unit:

[Unit]
Description=pbx-agent
Wants=network-online.target
After=network-online.target
[Service]
Type=exec
ExecStart=/var/lib/pbx-agent/bin/pbx-agent run
Restart=always
RestartSec=5
ProtectHome=yes
PrivateTmp=yes
[Install]
WantedBy=multi-user.target

The logs: journalctl -u pbx-agent.

pbx-agent opens every connection, to https://agents.<portal domain>/agent/v1. JSON over HTTPS, bodies of at most 256 KiB.

EndpointPurposeAuthenticated by
POST /enrollExchange the enrollment token for an identity and the node’s configurationThe enrollment token
POST /syncReport the node’s state; receive the desired state and at most one operationpbx-agent’s signature
POST /operations/{id}/eventsReport an operation’s progress and resultpbx-agent’s signature
GET /releases/{version}/manifestFetch a signed release manifestpbx-agent’s signature
POST /rotate-keyRegister new public keys, signed with the old keypbx-agent’s signature

Signing. Every request after the enrollment carries pbx-agent’s ID, a timestamp and an Ed25519 signature over the method, the path, the timestamp and a SHA-256 of the body. The portal refuses a clock skew of more than 60 seconds and a signature it has seen in the last 120 seconds.

Sealing. Every secret the portal sends (k3s tokens, S3 keys, Hetzner tokens, add-on secrets) is encrypted to pbx-agent’s X25519 key with a libsodium sealed box. A secret field that arrives in clear makes pbx-agent refuse the whole document (unsealed_secret).

How pbx-agent answers the portal’s status codes

Section titled “How pbx-agent answers the portal’s status codes”
Answerpbx-agent
200applies it
400logs it and backs off; never sends the same request more than three times. A replayed answer (a restarted pbx-agent signed the same request in the same second) is signed once more with the next second.
401treats its identity as revoked: stops syncing, leaves the cluster alone, keeps running idle
409 clock_skewcorrects its clock offset from the portal’s time and retries
426 protocol_too_oldupdates itself to the pbx-agent of the release the portal names, whatever the desired state says
429waits as long as the portal says
5xx, no networkbacks off exponentially from 5 seconds up to 5 minutes, with jitter
  1. On first start pbx-agent creates its keys and stores them before it uses them.
  2. It reads the server’s ID from Hetzner’s metadata service.
  3. It sends the enrollment token, its public keys, and the hash of the boot image the node actually booted. The portal checks the token (unused, unexpired, bound to this cluster and node name), checks through the Hetzner API that the server carries the cluster’s label and the node’s name, and compares the boot image’s hash with the release.
  4. It stores the answer and deletes the token from its configuration.
  5. It writes k3s’s configuration, starts k3s as a server, waits for the API, and writes the Secrets kube-system/hcloud and kube-system/pbx-etcd-s3.
The portal answersMeaningpbx-agent
200identity and configurationgoes on
401 invalid_token, token_used, token_expired, server_mismatchthe token is refusedtries at most three times, then idles; a new token means a new server
409 image_mismatchthe node booted another image than its release’sidles; the node needs a new server from the right image
409 conflictthe configuration cannot be built yetretries with backoff
503 internalthe portal cannot reach the Hetzner APIretries with backoff
429too many requestswaits as long as the portal says

The same token with the same keys from the same server gets the same answer again, so an answer lost on the way is asked for again safely.

About every 30 seconds (± 5) pbx-agent sends its report and receives the desired state. Each cycle it converges in this order: its own update, the host’s configuration (k3s drop-ins), the cluster’s configuration (Secrets, add-ons), then the operation. The cluster’s configuration is applied again every 10 minutes even without a change, so a change by hand to a managed object is overwritten.

A desired state is applied as a whole. Its release manifest is fetched and verified first; a manifest that no built-in key verifies changes nothing (manifest_rejected).

A drop-in that needs a restart of k3s is written at once; the restart happens at the next moment inside the maintenance window, never while an operation runs. In progress This has not run on a real node yet.

On a cluster with more than one server, one pbx-agent leads: the holder of the Lease pbx-system/pbx-agent-leader. On a single node its pbx-agent leads while the Kubernetes API is up.

pbx-agent applies each add-on from the template in the signed release, with the options from the desired state.

WhatHow
Objects it may createHelmChart and HelmChartConfig objects, Secrets named pbx-addon-* in kube-system, and for the PaaSbox Platform one cluster-wide Platform named after the cluster
Labelspbx.io/managed=true on every managed object, pbx.io/addon=<name> on an add-on’s objects
Secret optionsOnly in a pbx-addon-<name> Secret that the chart reads, never in a HelmChart: k3s encrypts Secrets at rest, not custom resources
Switched offIts objects are deleted; data, volumes and CRDs stay. The Secrets hcloud and pbx-etcd-s3 are never deleted.
A template or an apply that failsThe add-on reports failed and keeps the objects it has: a failed render never uninstalls a running chart

Each add-on reports one state: applied, pending, failed or removed.

CaseState
An add-on it requires is offfailed: requires add-on <r>, which is off
A conflicting add-on is on and installedfailed: conflicts with add-on <c>; switch that off first
A conflicting add-on is off, its objects still existpending until they are gone
An add-on it requires is not applied yetpending
Applied, its health check not passing yetpending; pbx-agent applies the cluster’s configuration at every cycle until it passes
Applied, its health check passes, or it has noneapplied
Switched off while an add-on that requires it is onfailed: still required

Built in a lab: the PaaSbox Platform add-on is switched off in two steps. While a SaaSApplication exists, switching it off changes nothing: the add-on reports failed with the applications’ names. Delete them first. Then pbx-agent applies a Platform that installs nothing and waits until Flux has removed what the platform installed. A Platform that refuses the removal (a database, or an application that still needs a part of it) keeps everything and reports the add-on failed with its message.

One operation runs per cluster at a time. Each is a list of named steps that can be repeated safely; pbx-agent records the step before it starts it and resumes there after a crash or a reboot. A failed step reports its name and the last 200 lines of its log. Destructive steps are never retried by pbx-agent on its own.

OperationStepsWhat it does
snapshot.savesave, confirm_s3, reportTakes an etcd snapshot and confirms it reached the bucket
cluster.restorestop_k3s, reset_restore, start_k3s, wait_api, forget_stale_nodesResets the cluster’s state to a snapshot; see Take and restore snapshots
cluster.upgradepreflight, pre_snapshot, fetch_image, stage_slot, reboot, health_gate, commitUpdates the node image in place; see Upgrade a cluster
access.issueensure_sa, token, sealIssues a kubeconfig sealed to the key of whoever asked: your browser, or an agent’s own key
access.revokedelete_saDeletes a person’s ServiceAccounts, or everyone’s
agent.rotate_keygenerate, register, confirmReplaces pbx-agent’s own keys; the portal asks for it when they are 90 days old
s3.verifylist_bucketLists the bucket with new keys before they are used
diag.collectcollectVersions, unit states and the last lines of the logs of k3s and pbx-agent, only with your consent for that request; never Secrets, tokens or environment files
server.rejoin, node.drain, node.forgetFor clusters with more than one server; not used on a single node

How the steps decide, where it matters:

  • reset_restore puts the snapshot on the node itself, from k3s’s local copy or fetched from the bucket with pbx-agent’s own S3 client, unpacks it and runs k3s server --cluster-reset with --etcd-s3=false. The S3 keys never reach a command line or k3s’s environment.
  • forget_stale_nodes deletes a Node other than pbx-agent’s own whose kubelet has not reported since the restored API came up, after a grace of 2 minutes.
  • cluster.upgrade takes the snapshot pbx-pre-upgrade-<release>, stages the new image as the boot entry pbx-<release>-<arch>.efi and keeps two entries, the new one and the one it came from. The health gate waits for the API, the Node Ready on the new k3s, the new boot entry, and every Deployment in kube-system available. A node that is not healthy within 20 minutes of the reboot is rebooted into the old entry. In progress That return has not fired on a real node yet.
  • reset_restore and reboot are the destructive steps. After a crash inside one, pbx-agent checks what the node shows: not started, it runs; completed, it goes on; anything else is failed and never run again on its own.

An unknown operation is answered as unsupported. There is no operation that runs a command you or anyone else types.

With the Kubernetes API down, pbx-agent accepts only cluster.restore and diag.collect.

Situationpbx-agent
The portal cannot be reachedkeeps the last desired state and retries with backoff; k3s keeps taking the scheduled snapshots
The Kubernetes API is downreports it and accepts only host-level operations: restore and diagnostics
An operation step failsstops, reports the step and the last 200 log lines; never retries a destructive step on its own
A release manifest’s signature is invalidrefuses it, reports manifest_rejected, changes nothing
A secret arrives in clearrefuses the whole document, reports unsealed_secret, changes nothing
pbx-agent is revokedstops syncing and leaves the cluster as it is

When the desired state names another version of pbx-agent, it fetches the release manifest, takes the program for its architecture, checks its SHA-256, keeps the running program as pbx-agent.prev, arms a rollback timer of 10 minutes, swaps the programs and restarts. The first successful sync with the portal disarms the timer. If it fires, the previous program is put back. A pbx-agent that finds itself rolled back reports it and does not try that version again for six hours.

pbx-agent trusts only release manifests signed with one of the two release keys built into it. Nothing at run time can change those keys. A pbx-agent built without them refuses every manifest: it never updates itself and never upgrades the node. Releases and channels has the signing.

Terminal window
pbx-agent uninstall --yes

Disables and removes the unit and the rollback timer, and removes /var/lib/pbx-agent and /etc/pbx-agent. k3s, its configuration and every Kubernetes object stay. Without --yes it prints what it would do and changes nothing. In progress Not run on a real node yet. Detach or delete a cluster says when you use it.