A nightly restore drill by an agent
A snapshot helps only if it restores, and the way to know is to restore one. This recipe has an agent prove the restore every night on staging: it writes a marker, takes a snapshot, writes a second marker, restores the snapshot in place, and checks that the first marker came back and the second did not. Then it reports.
What you need
Section titled “What you need”- Staging as a cluster of its own,
upcheck-stagingfrom Learn step 4, with its snapshots in your S3 bucket. - A time when staging may be down. A restore stops the cluster’s Kubernetes API while it runs. In the lab the restore took 160 seconds, with the API back about 35 seconds after the start. A restore is queued at once; it does not wait for the maintenance window, so schedule the drill inside staging’s window yourself.
- An API key from API keys & agents, made by a team admin, with all four scopes:
k3s:readto follow the operations,k3s:writefor the snapshot,k3s:accessfor a kubeconfig andk3s:destructivefor the restore. - A place for the report: an issue, a chat message, a mail; whatever your agent can write to.
The drill
Section titled “The drill”Give your agent these steps, or write them as a script that calls the REST API. Each step names the MCP tool and the endpoint below …/k3s/clusters/upcheck-staging/.
-
Get a kubeconfig.
request_kubeconfigwithroleadmin, thenget_kubeconfig_result(POST access/,GET access/<id>/result/), decrypted locally withpaasbox_kubeconfig.py. Ask for it before the snapshot: the restore takes the cluster’s state back to the snapshot, and with it the ServiceAccount behind any kubeconfig issued later. -
Write the first marker.
Terminal window kubectl -n default create configmap drill-before --from-literal=at="$(date -u +%FT%TZ)" -
Take the snapshot.
create_snapshot(POST snapshots/) queues it. Pollget_operation(GET operations/<id>/) untilstateissucceeded; in the lab a snapshot reached the bucket in 25 seconds. The newest name inlist_snapshots(GET snapshots/) is the one to restore. -
Write the second marker, which the restore must remove:
Terminal window kubectl -n default create configmap drill-after --from-literal=at="$(date -u +%FT%TZ)" -
Restore.
restore_snapshot(POST restore/) with the snapshot’s name andconfirmset toupcheck-staging, the cluster’s name. Pollget_operationuntil it hassucceededorfailed. The steps arestop_k3s,reset_restore,start_k3s,wait_apiandforget_stale_nodes; the last one waits two minutes before it removes nodes that did not come back. -
Check.
Terminal window kubectl -n default get configmap drill-before drill-afterkubectl -n upcheck get saasapp upcheckdrill-beforemust be there with its time,drill-aftermust be gone, and the app must reachReadyagain.get_clustermust showstatereadyand no health statement worse than before. Add the checks your app needs, such as a request to its health page. -
Report and clean up. Send what passed and failed, the snapshot’s name, the two operation IDs and how long each took. Then delete
drill-before.
If a step fails, the agent stops there and reports it with the operation’s failed step and message from get_operation. It does not try again: a second restore on a cluster that failed one needs your eyes first.
Who confirms the restore
Section titled “Who confirms the restore”At night nobody is there to say yes. The portal checks that confirm is the cluster’s name, which stops the agent from restoring a cluster it did not name; it does not check that a person agreed. A key’s scopes hold for every cluster of your team, so the key that may restore staging may also restore production. Three ways to keep that in hand:
- Run the drill as a script you have read, with
upcheck-stagingwritten into it, and keep the key only where the script runs. - Or start the agent yourself each morning and confirm the restore at its prompt. Then it is a morning drill, not a nightly one.
- And keep production in a team this key does not belong to, with staging in the key’s team: then the key cannot reach production at all.
Guardrails for agents says what each limit stops.
What the drill does not prove
Section titled “What the drill does not prove”- Your volumes. A snapshot holds the cluster’s state, not the data in your volumes. A restore in place leaves the server’s disk as it is: in the lab the data in a volume was intact afterwards. It was never in the snapshot.
- Your database. Postgres on the PaaSbox Platform has backups of its own, per application. Restoring the cluster’s state does not restore a database; test that with the database’s own restore.
- A new server. If the server is lost, a restore in place has nothing to run on. A restore onto a new server is Planned.
- A bucket whose keys you hold yourself. The lab used a bucket whose keys the portal held; with your own keys the restore has not run yet In progress.