Skip to content

Practise a restore

This is the fifth step of the Learn path. You restore staging from a snapshot once, on purpose, and see for yourself what a snapshot brings back and what it leaves as it is. Do it on staging: a restore stops the Kubernetes API, and on production your users would notice.

A restore you have done once, with its timing, and two markers that show the line a snapshot draws: the objects in the cluster come back as they were, the data in volumes does not.

  • upcheck-staging from step 4, with its snapshots going to its bucket. The portal restores only snapshots that are in S3.
  • An admin kubeconfig of staging, issued before you take the snapshot. The restore removes everything the cluster stored after the snapshot, and the ServiceAccount of a kubeconfig issued later is part of that. If yours stops working after the restore, get a new one.
  1. Save this as drill.yaml. It holds a ConfigMap, which lives in the cluster’s state, and a file in a volume on the server’s disk, written by a small pod:

    apiVersion: v1
    kind: Namespace
    metadata: {name: drill}
    ---
    apiVersion: v1
    kind: ConfigMap
    metadata: {name: marker, namespace: drill}
    data: {state: before}
    ---
    apiVersion: v1
    kind: PersistentVolumeClaim
    metadata: {name: data, namespace: drill}
    spec:
    accessModes: [ReadWriteOnce]
    storageClassName: local-lvm-thin
    resources: {requests: {storage: 1Gi}}
    ---
    apiVersion: apps/v1
    kind: Deployment
    metadata: {name: marker, namespace: drill}
    spec:
    replicas: 1
    strategy: {type: Recreate}
    selector: {matchLabels: {app: marker}}
    template:
    metadata: {labels: {app: marker}}
    spec:
    containers:
    - name: c
    image: busybox:1.37
    command: [sh, -c, "sleep 1000000"]
    volumeMounts: [{name: data, mountPath: /data}]
    volumes: [{name: data, persistentVolumeClaim: {claimName: data}}]
  2. Apply it and write the file:

    Terminal window
    export KUBECONFIG=~/Downloads/upcheck-staging-admin.kubeconfig
    kubectl apply -f drill.yaml
    kubectl -n drill rollout status deployment/marker
    kubectl -n drill exec deploy/marker -- sh -c 'echo before > /data/marker'
  1. On staging’s Overview, choose Snapshot now. Open Backups: the snapshot is listed under Snapshots, and Where says S3. In the lab a snapshot reached the bucket in 25 seconds.

  2. Change both markers, and add an object the snapshot has never seen:

    Terminal window
    kubectl -n drill patch configmap marker -p '{"data":{"state":"after"}}'
    kubectl -n drill create configmap made-after --from-literal=state=after
    kubectl -n drill exec deploy/marker -- sh -c 'echo after > /data/marker'
  1. On Backups, choose Restore next to the snapshot. The page says what follows: pbx-agent stops k3s, resets the cluster’s state to the snapshot and starts k3s again; everything stored after the snapshot is gone; data in volumes stays as it is.

  2. Type upcheck-staging under Type the cluster’s name and choose Restore.

  3. Follow the operation on Operations. Its steps are stop_k3s, reset_restore (the snapshot is staged on the node and k3s resets its database to it), start_k3s, wait_api and forget_stale_nodes, which waits two minutes for nodes that do not come back; on a single node there are none. In the lab the API answered again about 35 seconds after the start, and the restore was done after 160 seconds.

If pbx-agent stops in the middle, it resumes at the step it was on: in the lab it was killed during the last step, restarted after 5 seconds and finished the restore.

Terminal window
kubectl -n drill get configmap marker -o jsonpath='{.data.state}' # before
kubectl -n drill get configmap made-after # NotFound
kubectl -n drill exec deploy/marker -- cat /data/marker # after

The ConfigMap is back as it was, and the one made after the snapshot is gone: a snapshot holds the cluster’s state, every object in the Kubernetes API. The file still says after: the volume is not part of the snapshot, and the restore leaves it alone.

That line runs through your app too. upcheck’s Postgres keeps its data in a volume on the server’s disk, so a restore in place does not take the database back to the time of the snapshot. Going back in the database is the job of its own backups, data.postgres.backup, which step 2 recommends: the platform restores them into a new app with data.postgres.restoreFrom, optionally to a point in time. In progress

A restore onto a new server, for a server that is gone, is Planned.

When you are done, remove the drill: kubectl delete namespace drill.

A restore you have watched from the click to the check, with its figures, and the line between what a snapshot holds and what your volumes hold. Next, the other operation that stops the API: an upgrade.