Skip to content

An agent upgrades staging first

With a cluster per stage, every release has to be installed more than once. This recipe hands that to an agent, in a fixed order: when staging is offered a new release, the agent upgrades staging at once, checks the app, and only if everything passed schedules production for its next maintenance window. If staging fails, production is left alone.

  • Staging and production as separate clusters, upcheck-staging and upcheck-prod, as in Learn step 4. On each cluster’s Upgrades tab, set the Release channel: early for staging, stable for production. A release reaches early first and stable after its time there, so staging sees it before production is offered it.
  • On production, untick Install patch releases in the window without a click. Otherwise patch releases install in production’s window whether staging passed or not.
  • An API key from API keys & agents, made by a team admin, with k3s:read, k3s:write and k3s:access. It does not need k3s:destructive: an upgrade is not one of the destructive operations.
  • Your app’s checks: what has to be true after an upgrade, such as the health page answering and a login working.

Give your agent these steps and run it once a day, or when a release is announced.

  1. Look for a release. get_cluster for upcheck-staging. Under upgrade, available is the release on offer with its version, its k3s version and its notes, and patch says whether it keeps the k3s minor version. canSchedule is true when the cluster is ready and no upgrade is pending. Nothing available: stop.

  2. Upgrade staging now. schedule_upgrade with when now and release set to that version: if another version is offered by then, the portal refuses the call with 409 instead of installing it. The upgrade takes a snapshot first, writes the new image next to the running one, reboots the node into it and checks its health. In the lab it took 83 seconds, and the Kubernetes API was down for 34 seconds.

  3. Wait for the outcome. Poll get_operation until state is succeeded or failed.

  4. Check staging. It passed when all of these hold:

    • the operation succeeded;
    • get_cluster shows state ready, release the new version, and no health statement that was not there before;
    • with an admin kubeconfig from request_kubeconfig (a view one cannot list nodes), kubectl get nodes shows the node Ready and kubectl -n upcheck get saasapp upcheck shows Ready;
    • your app’s own checks pass.
  5. Schedule production, or do not. If staging passed, wait until get_cluster for upcheck-prod offers the same version, then call schedule_upgrade with when window and release set to it. The upgrade waits for production’s next maintenance window, which upgrade.nextWindow shows. If production is offered a version that staging never ran, schedule nothing and report it.

  6. Report. Which version, staging’s result with the operation’s ID and duration, and whether production is scheduled and for when.

The agent stops, schedules nothing on production, and reports the operation’s failed step and its message from get_operation. It does not try again on its own.

What staging looks like after a failure depends on the step. Before the reboot, the node still runs the release it had. After the reboot, a node that is not healthy on the new release within 20 minutes is rebooted into the previous one In progress; that fallback has not run on a real node. Either way the snapshot the upgrade took first, whose name starts with pbx-pre-upgrade-, is in your bucket, and Take and restore snapshots has the restore.

With this key the agent cannot change a cluster’s channel, its maintenance window or the patch-release setting: those are on the Upgrades and Settings tabs, and there is no tool for them. It cannot restore or delete anything either, because the key has no k3s:destructive.