Skip to content

A test cluster per pull request

This recipe gives every pull request a cluster of its own: CI creates a throwaway cluster, deploys the branch on it, runs the tests and deletes it again, whether the tests passed or not. It is the loop of Learn step 3, run by GitHub Actions through the REST API.

  • An API key from API keys & agents, made by a team admin, with the scopes k3s:read, k3s:write, k3s:access and k3s:destructive. Store it as the repository secret PAASBOX_API_KEY. It can delete any cluster of its team: keep it in CI only, and keep production in a team the key does not belong to.
  • A Hetzner project for test clusters, connected to the portal. Its ID is in GET …/k3s/projects/.
  • Four repository variables: PAASBOX_PORTAL (the portal’s address), PAASBOX_TEAM (your team’s slug), PAASBOX_PROJECT (the project’s ID) and ACME_EMAIL (the address Let’s Encrypt registers the platform’s certificates under). API keys & agents shows the REST address the first two make up.
  • A team that already has its subscription. The team’s first cluster starts the subscription in the portal’s checkout and is created in the portal, not through the API (402 payment_method_required). The same holds again once the team’s last cluster is gone and the month already paid has ended.
  • A release with the PaaSbox Platform add-on on the stable channel, release 2026.10.2 or later. A new cluster gets the release its channel offers, stable unless you ask for early.
  • Your upcheck.yaml from step 2 in the repository, a build that pushes the branch’s image, and a test command, here ./smoke-test.sh.
  1. Add .github/workflows/test-cluster.yml:

    name: test-cluster
    on: pull_request
    jobs:
    test:
    runs-on: ubuntu-latest
    timeout-minutes: 45
    env:
    PAASBOX_PORTAL: ${{ vars.PAASBOX_PORTAL }}
    PAASBOX_API_KEY: ${{ secrets.PAASBOX_API_KEY }}
    API: ${{ vars.PAASBOX_PORTAL }}/api/v1/teams/${{ vars.PAASBOX_TEAM }}/k3s
    AUTH: "Authorization: Bearer ${{ secrets.PAASBOX_API_KEY }}"
    NAME: upcheck-pr-${{ github.event.pull_request.number }}-${{ github.run_number }}
    IMAGE: registry.example.com/upcheck:${{ github.sha }}
    steps:
    - uses: actions/checkout@v4
    # Build and push $IMAGE here.
    - name: Create the cluster
    run: |
    body=$(jq -n --arg n "$NAME" --argjson p "${{ vars.PAASBOX_PROJECT }}" \
    '{name: $n, confirm: $n, project: $p, location: "fsn1", serverType: "cpx22",
    hcloudToken: {mode: "copy"}}')
    slug=$(curl -fsS -X POST "$API/clusters/" -H "$AUTH" \
    -H "Content-Type: application/json" -d "$body" | jq -r .cluster.slug)
    echo "SLUG=$slug" >> "$GITHUB_ENV"
    - name: Wait for the cluster, then switch the platform on
    run: |
    while :; do
    state=$(curl -fsS -H "$AUTH" "$API/clusters/$SLUG/" | jq -r .state)
    [ "$state" = ready ] && break
    [ "$state" = failed ] && exit 1
    sleep 10
    done
    body=$(jq -n --arg e "${{ vars.ACME_EMAIL }}" \
    '{enabled: true, options: {profile: "saas-http01", acmeEmail: $e}, also: ["flux"]}')
    curl -fsS -X PATCH "$API/clusters/$SLUG/addons/paasbox-platform/" -H "$AUTH" \
    -H "Content-Type: application/json" -d "$body"
    while :; do
    state=$(curl -fsS -H "$AUTH" "$API/clusters/$SLUG/addons/" | jq -r \
    '.addons[] | select(.name == "paasbox-platform") | .reported.state')
    [ "$state" = applied ] && break
    [ "$state" = failed ] && exit 1
    sleep 10
    done
    - name: Get a kubeconfig
    run: |
    curl -fsSO "$PAASBOX_PORTAL/mcp/paasbox_kubeconfig.py" && pip install cryptography
    python3 paasbox_kubeconfig.py fetch --team ${{ vars.PAASBOX_TEAM }} \
    --cluster "$SLUG" --role admin > "$RUNNER_TEMP/kubeconfig"
    chmod 600 "$RUNNER_TEMP/kubeconfig"
    echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV"
    - name: Deploy the branch
    run: |
    yq '(select(.kind == "SaaSApplication") | .spec.image) = strenv(IMAGE) | del(.spec.exposure)' \
    upcheck.yaml | kubectl apply -f -
    kubectl -n upcheck wait saasapp/upcheck --for=jsonpath='{.status.phase}'=Ready --timeout=15m
    - name: Test
    run: |
    kubectl -n upcheck port-forward svc/upcheck-web 8000:80 &
    sleep 3
    ./smoke-test.sh http://localhost:8000
    - name: Delete the cluster
    if: always() && env.SLUG != ''
    run: curl -fsS -X DELETE -H "$AUTH" "$API/clusters/$SLUG/?confirm=$NAME&finalSnapshot=false"
  2. Open a pull request. The job’s log shows the cluster’s state while it waits; in the lab a cluster was ready 2 minutes 46 seconds after the request. The first cluster in the test project also waits for the node image to be copied into it, 69 seconds in the lab; later ones use the copy.

What each part relies on:

  • confirm must be the cluster’s name, at create and at delete; any other value is refused with 422 and the code confirmation_required. The create also runs the create page’s checks: a release on the stable channel, your team’s limit of 10 clusters (quota_exceeded), the accepted terms, billing (payment_method_required while the team has no subscription).
  • hcloudToken copy puts a copy of the project’s token into the cluster, so CI never handles a Hetzner token. In a project that holds only test clusters, the copy reaches nothing else.
  • The PATCH of paasbox-platform switches the PaaSbox Platform on, and with also the Flux add-on it requires. The platform needs acmeEmail for its Let’s Encrypt account, and a create takes only the platform’s profile, not that option, so the platform goes on once the cluster is ready. It is ready when the add-on reports applied; failed ends the job. In the lab, Flux and the platform switched on in one save were applied after 2 minutes 53 seconds, and upcheck was Ready 82 seconds after kubectl apply. A part of the platform that fails once is retried by Flux only after an hour, a defect still open, which the job’s 45 minutes do not cover.
  • The yq line sets the branch’s image on the SaaSApplication only, not on the Namespace the file also holds, and removes spec.exposure. Without it the app gets no Ingress, no certificate and no DNS name, and the tests reach it through a port-forward to the Service upcheck-web, port 80.
  • The API ranges stay open to everyone, because a GitHub-hosted runner has no fixed address; the API still needs the kubeconfig’s token. With a self-hosted runner, set allowedApiRanges to its address.
  • if: always() deletes the cluster when a step before it failed or the run was cancelled. The cluster stops counting for billing at the delete; in the lab a delete took 14 seconds.

A run that dies before its last step, a runner that disappears for example, leaves its cluster behind. A second workflow, on a schedule, deletes every test cluster older than two hours:

Terminal window
cut=$(date -u -d '-2 hours' +%FT%TZ)
curl -fsS -H "$AUTH" "$API/clusters/?page_size=100" \
| jq -r --arg cut "$cut" '.items[]
| select((.name | startswith("upcheck-pr-")) and .createdAt < $cut and .state != "deleting")
| "\(.slug) \(.name)"' \
| while read -r slug name; do
curl -fsS -X DELETE -H "$AUTH" "$API/clusters/$slug/?confirm=$name&finalSnapshot=false"
done

Each run is one cluster, counted from the moment it is first ready until the delete. In progress As the Paddle billing is written, a test cluster next to a cluster the team keeps is charged prorated when it becomes ready and credited for the rest of the month when it is deleted. In a team whose only clusters are test clusters, the month is paid when the first one is created, and test clusters within that month use it; after a month with no cluster at its end, the next one needs the portal’s checkout again (Costs and billing). Hetzner bills the server by the hour, at its own price. Every running test cluster counts against your team’s limit of 10 clusters, next to staging and production: with staging and production of one app, eight pull requests can have a cluster at the same time. Ask for a higher limit when you need more. A key is also limited in how many calls it makes per minute: poll every ten seconds, not faster (Limits and quotas).

The same steps work through the MCP tools, as in Learn step 3: give your agent the pull request’s image and the task, and keep create_cluster and delete_cluster on asking you. Then you confirm each cluster by hand, which CI cannot do.