Skip to content

A nightly restore drill by an agent

A snapshot helps only if it restores, and the way to know is to restore one. This recipe has an agent prove the restore every night on staging: it writes a marker, takes a snapshot, writes a second marker, restores the snapshot in place, and checks that the first marker came back and the second did not. Then it reports.

  • Staging as a cluster of its own, upcheck-staging from Learn step 4, with its snapshots in your S3 bucket.
  • A time when staging may be down. A restore stops the cluster’s Kubernetes API while it runs. In the lab the restore took 160 seconds, with the API back about 35 seconds after the start. A restore is queued at once; it does not wait for the maintenance window, so schedule the drill inside staging’s window yourself.
  • An API key from API keys & agents, made by a team admin, with all four scopes: k3s:read to follow the operations, k3s:write for the snapshot, k3s:access for a kubeconfig and k3s:destructive for the restore.
  • A place for the report: an issue, a chat message, a mail; whatever your agent can write to.

Give your agent these steps, or write them as a script that calls the REST API. Each step names the MCP tool and the endpoint below …/k3s/clusters/upcheck-staging/.

  1. Get a kubeconfig. request_kubeconfig with role admin, then get_kubeconfig_result (POST access/, GET access/<id>/result/), decrypted locally with paasbox_kubeconfig.py. Ask for it before the snapshot: the restore takes the cluster’s state back to the snapshot, and with it the ServiceAccount behind any kubeconfig issued later.

  2. Write the first marker.

    Terminal window
    kubectl -n default create configmap drill-before --from-literal=at="$(date -u +%FT%TZ)"
  3. Take the snapshot. create_snapshot (POST snapshots/) queues it. Poll get_operation (GET operations/<id>/) until state is succeeded; in the lab a snapshot reached the bucket in 25 seconds. The newest name in list_snapshots (GET snapshots/) is the one to restore.

  4. Write the second marker, which the restore must remove:

    Terminal window
    kubectl -n default create configmap drill-after --from-literal=at="$(date -u +%FT%TZ)"
  5. Restore. restore_snapshot (POST restore/) with the snapshot’s name and confirm set to upcheck-staging, the cluster’s name. Poll get_operation until it has succeeded or failed. The steps are stop_k3s, reset_restore, start_k3s, wait_api and forget_stale_nodes; the last one waits two minutes before it removes nodes that did not come back.

  6. Check.

    Terminal window
    kubectl -n default get configmap drill-before drill-after
    kubectl -n upcheck get saasapp upcheck

    drill-before must be there with its time, drill-after must be gone, and the app must reach Ready again. get_cluster must show state ready and no health statement worse than before. Add the checks your app needs, such as a request to its health page.

  7. Report and clean up. Send what passed and failed, the snapshot’s name, the two operation IDs and how long each took. Then delete drill-before.

If a step fails, the agent stops there and reports it with the operation’s failed step and message from get_operation. It does not try again: a second restore on a cluster that failed one needs your eyes first.

At night nobody is there to say yes. The portal checks that confirm is the cluster’s name, which stops the agent from restoring a cluster it did not name; it does not check that a person agreed. A key’s scopes hold for every cluster of your team, so the key that may restore staging may also restore production. Three ways to keep that in hand:

  • Run the drill as a script you have read, with upcheck-staging written into it, and keep the key only where the script runs.
  • Or start the agent yourself each morning and confirm the restore at its prompt. Then it is a morning drill, not a nightly one.
  • And keep production in a team this key does not belong to, with staging in the key’s team: then the key cannot reach production at all.

Guardrails for agents says what each limit stops.

  • Your volumes. A snapshot holds the cluster’s state, not the data in your volumes. A restore in place leaves the server’s disk as it is: in the lab the data in a volume was intact afterwards. It was never in the snapshot.
  • Your database. Postgres on the PaaSbox Platform has backups of its own, per application. Restoring the cluster’s state does not restore a database; test that with the database’s own restore.
  • A new server. If the server is lost, a restore in place has nothing to run on. A restore onto a new server is Planned.
  • A bucket whose keys you hold yourself. The lab used a bucket whose keys the portal held; with your own keys the restore has not run yet In progress.