A test cluster per pull request
This recipe gives every pull request a cluster of its own: CI creates a throwaway cluster, deploys the branch on it, runs the tests and deletes it again, whether the tests passed or not. It is the loop of Learn step 3, run by GitHub Actions through the REST API.
What you need
Section titled “What you need”- An API key from API keys & agents, made by a team admin, with the scopes
k3s:read,k3s:write,k3s:accessandk3s:destructive. Store it as the repository secretPAASBOX_API_KEY. It can delete any cluster of its team: keep it in CI only, and keep production in a team the key does not belong to. - A Hetzner project for test clusters, connected to the portal. Its ID is in
GET …/k3s/projects/. - Four repository variables:
PAASBOX_PORTAL(the portal’s address),PAASBOX_TEAM(your team’s slug),PAASBOX_PROJECT(the project’s ID) andACME_EMAIL(the address Let’s Encrypt registers the platform’s certificates under). API keys & agents shows the REST address the first two make up. - A team that already has its subscription. The team’s first cluster starts the subscription in the portal’s checkout and is created in the portal, not through the API (
402 payment_method_required). The same holds again once the team’s last cluster is gone and the month already paid has ended. - A release with the PaaSbox Platform add-on on the stable channel, release 2026.10.2 or later. A new cluster gets the release its channel offers, stable unless you ask for early.
- Your
upcheck.yamlfrom step 2 in the repository, a build that pushes the branch’s image, and a test command, here./smoke-test.sh.
The workflow
Section titled “The workflow”-
Add
.github/workflows/test-cluster.yml:name: test-clusteron: pull_requestjobs:test:runs-on: ubuntu-latesttimeout-minutes: 45env:PAASBOX_PORTAL: ${{ vars.PAASBOX_PORTAL }}PAASBOX_API_KEY: ${{ secrets.PAASBOX_API_KEY }}API: ${{ vars.PAASBOX_PORTAL }}/api/v1/teams/${{ vars.PAASBOX_TEAM }}/k3sAUTH: "Authorization: Bearer ${{ secrets.PAASBOX_API_KEY }}"NAME: upcheck-pr-${{ github.event.pull_request.number }}-${{ github.run_number }}IMAGE: registry.example.com/upcheck:${{ github.sha }}steps:- uses: actions/checkout@v4# Build and push $IMAGE here.- name: Create the clusterrun: |body=$(jq -n --arg n "$NAME" --argjson p "${{ vars.PAASBOX_PROJECT }}" \'{name: $n, confirm: $n, project: $p, location: "fsn1", serverType: "cpx22",hcloudToken: {mode: "copy"}}')slug=$(curl -fsS -X POST "$API/clusters/" -H "$AUTH" \-H "Content-Type: application/json" -d "$body" | jq -r .cluster.slug)echo "SLUG=$slug" >> "$GITHUB_ENV"- name: Wait for the cluster, then switch the platform onrun: |while :; dostate=$(curl -fsS -H "$AUTH" "$API/clusters/$SLUG/" | jq -r .state)[ "$state" = ready ] && break[ "$state" = failed ] && exit 1sleep 10donebody=$(jq -n --arg e "${{ vars.ACME_EMAIL }}" \'{enabled: true, options: {profile: "saas-http01", acmeEmail: $e}, also: ["flux"]}')curl -fsS -X PATCH "$API/clusters/$SLUG/addons/paasbox-platform/" -H "$AUTH" \-H "Content-Type: application/json" -d "$body"while :; dostate=$(curl -fsS -H "$AUTH" "$API/clusters/$SLUG/addons/" | jq -r \'.addons[] | select(.name == "paasbox-platform") | .reported.state')[ "$state" = applied ] && break[ "$state" = failed ] && exit 1sleep 10done- name: Get a kubeconfigrun: |curl -fsSO "$PAASBOX_PORTAL/mcp/paasbox_kubeconfig.py" && pip install cryptographypython3 paasbox_kubeconfig.py fetch --team ${{ vars.PAASBOX_TEAM }} \--cluster "$SLUG" --role admin > "$RUNNER_TEMP/kubeconfig"chmod 600 "$RUNNER_TEMP/kubeconfig"echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV"- name: Deploy the branchrun: |yq '(select(.kind == "SaaSApplication") | .spec.image) = strenv(IMAGE) | del(.spec.exposure)' \upcheck.yaml | kubectl apply -f -kubectl -n upcheck wait saasapp/upcheck --for=jsonpath='{.status.phase}'=Ready --timeout=15m- name: Testrun: |kubectl -n upcheck port-forward svc/upcheck-web 8000:80 &sleep 3./smoke-test.sh http://localhost:8000- name: Delete the clusterif: always() && env.SLUG != ''run: curl -fsS -X DELETE -H "$AUTH" "$API/clusters/$SLUG/?confirm=$NAME&finalSnapshot=false" -
Open a pull request. The job’s log shows the cluster’s state while it waits; in the lab a cluster was ready 2 minutes 46 seconds after the request. The first cluster in the test project also waits for the node image to be copied into it, 69 seconds in the lab; later ones use the copy.
What each part relies on:
confirmmust be the cluster’s name, at create and at delete; any other value is refused with422and the codeconfirmation_required. The create also runs the create page’s checks: a release on the stable channel, your team’s limit of 10 clusters (quota_exceeded), the accepted terms, billing (payment_method_requiredwhile the team has no subscription).hcloudTokencopyputs a copy of the project’s token into the cluster, so CI never handles a Hetzner token. In a project that holds only test clusters, the copy reaches nothing else.- The PATCH of
paasbox-platformswitches the PaaSbox Platform on, and withalsothe Flux add-on it requires. The platform needsacmeEmailfor its Let’s Encrypt account, and a create takes only the platform’s profile, not that option, so the platform goes on once the cluster is ready. It is ready when the add-on reportsapplied;failedends the job. In the lab, Flux and the platform switched on in one save wereappliedafter 2 minutes 53 seconds, and upcheck wasReady82 seconds afterkubectl apply. A part of the platform that fails once is retried by Flux only after an hour, a defect still open, which the job’s 45 minutes do not cover. - The
yqline sets the branch’s image on theSaaSApplicationonly, not on the Namespace the file also holds, and removesspec.exposure. Without it the app gets no Ingress, no certificate and no DNS name, and the tests reach it through a port-forward to the Serviceupcheck-web, port 80. - The API ranges stay open to everyone, because a GitHub-hosted runner has no fixed address; the API still needs the kubeconfig’s token. With a self-hosted runner, set
allowedApiRangesto its address. if: always()deletes the cluster when a step before it failed or the run was cancelled. The cluster stops counting for billing at the delete; in the lab a delete took 14 seconds.
Sweep what was left
Section titled “Sweep what was left”A run that dies before its last step, a runner that disappears for example, leaves its cluster behind. A second workflow, on a schedule, deletes every test cluster older than two hours:
cut=$(date -u -d '-2 hours' +%FT%TZ)curl -fsS -H "$AUTH" "$API/clusters/?page_size=100" \ | jq -r --arg cut "$cut" '.items[] | select((.name | startswith("upcheck-pr-")) and .createdAt < $cut and .state != "deleting") | "\(.slug) \(.name)"' \ | while read -r slug name; do curl -fsS -X DELETE -H "$AUTH" "$API/clusters/$slug/?confirm=$name&finalSnapshot=false" doneWhat it costs
Section titled “What it costs”Each run is one cluster, counted from the moment it is first ready until the delete. In progress As the Paddle billing is written, a test cluster next to a cluster the team keeps is charged prorated when it becomes ready and credited for the rest of the month when it is deleted. In a team whose only clusters are test clusters, the month is paid when the first one is created, and test clusters within that month use it; after a month with no cluster at its end, the next one needs the portal’s checkout again (Costs and billing). Hetzner bills the server by the hour, at its own price. Every running test cluster counts against your team’s limit of 10 clusters, next to staging and production: with staging and production of one app, eight pull requests can have a cluster at the same time. Ask for a higher limit when you need more. A key is also limited in how many calls it makes per minute: poll every ten seconds, not faster (Limits and quotas).
With an agent instead of CI
Section titled “With an agent instead of CI”The same steps work through the MCP tools, as in Learn step 3: give your agent the pull request’s image and the task, and keep create_cluster and delete_cluster on asking you. Then you confirm each cluster by hand, which CI cannot do.