Skip to content

Upgrade in the maintenance window

This is the last step of the Learn path. You upgrade staging to a newer release while you watch, measure how long the Kubernetes API and your app are down, check the app, and then let production take the same release in its maintenance window.

Every upgrade reboots the server. On a single node the Kubernetes API and your apps are down while it restarts; in the lab the API was down for 34 seconds.

Staging on the newer release, with the downtime measured by you, and production queued for its window. That is the order PaaSbox is built for: a release reaches staging first, and production takes it only after staging ran it.

  • upcheck-staging and upcheck-prod from step 4. Staging is on the early channel, and its patch releases wait for your click.
  • A newer release on early. Staging’s Upgrades tab shows it as a card, Release and its version, with its notes. If it says that no newer release is offered on the channel, there is none yet; Releases and channels lists the releases.
  • An admin kubeconfig of staging, valid for an hour or more.

In a second terminal, start a loop that asks the API and the app once a second:

Terminal window
export KUBECONFIG=~/Downloads/upcheck-staging-admin.kubeconfig
while true; do
api=$(kubectl get --raw /readyz --request-timeout=2s 2>/dev/null || echo down)
app=$(curl -s -o /dev/null -m 2 -w '%{http_code}' https://upcheck-staging.example.com/)
echo "$(date +%T) api=$api app=$app"
sleep 1
done

api=ok app=200 is a healthy cluster. Leave the loop running.

  1. On staging’s Upgrades tab, read the release’s notes. The card says whether the release is a patch (the same k3s minor version) or a new k3s minor version. Choose Now. The portal queues the upgrade, and pbx-agent picks it up at its next sync; it syncs every 30 seconds.

  2. Follow the steps in the banner above the card:

    • preflight: the release’s signature verifies, the cluster may move to it from the release it runs, and the boot partition has room for the new image.
    • pre_snapshot: a snapshot of the cluster, to your bucket.
    • fetch_image and stage_slot: the new node image is downloaded, checked against the signed release, and written next to the running one as a second boot entry.
    • reboot: the server boots the new image. Your loop shows api=down and no 200 now.
    • health_gate: pbx-agent waits for the API, for the node to be Ready on the new k3s, for the new boot entry, and for every Deployment in kube-system.
    • commit: the new image becomes the default. The previous one stays as the second entry.
  3. Read the loop’s times. In the lab the API was down for 34 seconds, and the whole upgrade took 83 seconds from the click to success. Your app answers again once its pods have started; the health gate does not wait for them.

  4. The cluster page’s header now names the new release, and Upgrade history lists the upgrade as succeeded. The snapshot from pre_snapshot is in your bucket, named pbx-pre-upgrade- and the release.

If the node does not pass the health gate within 20 minutes of the reboot, pbx-agent reboots it into the previous image and reports the upgrade as failed. In progress That return has not been triggered on a real node yet. Once the node has passed the gate, nothing rolls back on its own: the snapshot from pre_snapshot is the way back for the cluster’s state, as in step 5.

Terminal window
kubectl -n upcheck get saasapp upcheck

The phase is Ready, and the loop shows app=200 again. Open the app in your browser and use it the way your users do. This check is yours: no part of PaaSbox knows what your app should answer. Stop the loop with Ctrl-C.

  1. Wait until the release reaches stable: after its seven days on early. Production’s Upgrades tab then shows the same card.

  2. With Install patch releases in the window without a click ticked, as on production, the portal queues a patch release on its own within the hour: the banner says the upgrade is queued for the window, with its start and end. A new k3s minor version waits for your click: choose In the next window.

  3. The Overview shows the Next window under Facts. The upgrade runs there, with the same steps and the same reboot. A window it misses, because pbx-agent was not reachable in time, moves it to the next one, for up to 14 days.

In progress Upgrades queued for the window have not run on a real server yet.

The whole path: a cluster in your own project, an app on it, an agent testing in a cluster of its own, staging and production apart, a restore you practised and an upgrade you watched before production took it. The next pages hand this work to an agent.