An agent upgrades staging first
With a cluster per stage, every release has to be installed more than once. This recipe hands that to an agent, in a fixed order: when staging is offered a new release, the agent upgrades staging at once, checks the app, and only if everything passed schedules production for its next maintenance window. If staging fails, production is left alone.
What you need
Section titled “What you need”- Staging and production as separate clusters,
upcheck-stagingandupcheck-prod, as in Learn step 4. On each cluster’s Upgrades tab, set the Release channel: early for staging, stable for production. A release reaches early first and stable after its time there, so staging sees it before production is offered it. - On production, untick Install patch releases in the window without a click. Otherwise patch releases install in production’s window whether staging passed or not.
- An API key from API keys & agents, made by a team admin, with
k3s:read,k3s:writeandk3s:access. It does not needk3s:destructive: an upgrade is not one of the destructive operations. - Your app’s checks: what has to be true after an upgrade, such as the health page answering and a login working.
The routine
Section titled “The routine”Give your agent these steps and run it once a day, or when a release is announced.
-
Look for a release.
get_clusterforupcheck-staging. Underupgrade,availableis the release on offer with itsversion, its k3s version and its notes, andpatchsays whether it keeps the k3s minor version.canScheduleis true when the cluster is ready and no upgrade is pending. Nothing available: stop. -
Upgrade staging now.
schedule_upgradewithwhennowandreleaseset to that version: if another version is offered by then, the portal refuses the call with409instead of installing it. The upgrade takes a snapshot first, writes the new image next to the running one, reboots the node into it and checks its health. In the lab it took 83 seconds, and the Kubernetes API was down for 34 seconds. -
Wait for the outcome. Poll
get_operationuntilstateissucceededorfailed. -
Check staging. It passed when all of these hold:
- the operation
succeeded; get_clustershowsstateready,releasethe new version, and no health statement that was not there before;- with an
adminkubeconfig fromrequest_kubeconfig(aviewone cannot list nodes),kubectl get nodesshows the nodeReadyandkubectl -n upcheck get saasapp upcheckshowsReady; - your app’s own checks pass.
- the operation
-
Schedule production, or do not. If staging passed, wait until
get_clusterforupcheck-prodoffers the same version, then callschedule_upgradewithwhenwindowandreleaseset to it. The upgrade waits for production’s next maintenance window, whichupgrade.nextWindowshows. If production is offered a version that staging never ran, schedule nothing and report it. -
Report. Which version, staging’s result with the operation’s ID and duration, and whether production is scheduled and for when.
When staging fails
Section titled “When staging fails”The agent stops, schedules nothing on production, and reports the operation’s failed step and its message from get_operation. It does not try again on its own.
What staging looks like after a failure depends on the step. Before the reboot, the node still runs the release it had. After the reboot, a node that is not healthy on the new release within 20 minutes is rebooted into the previous one In progress; that fallback has not run on a real node. Either way the snapshot the upgrade took first, whose name starts with pbx-pre-upgrade-, is in your bucket, and Take and restore snapshots has the restore.
What the agent cannot do
Section titled “What the agent cannot do”With this key the agent cannot change a cluster’s channel, its maintenance window or the patch-release setting: those are on the Upgrades and Settings tabs, and there is no tool for them. It cannot restore or delete anything either, because the key has no k3s:destructive.