Skip to content

Agents operating your apps

“Where AI agents operate your apps” is the direction of PaaSbox. Made literal: an agent, a language model with tools, does the operations work on a running app, such as checking its health, reading its logs, taking a backup, proving that the backup restores and running an upgrade. It does that only through the actions the app’s catalog entry declares, a person approves the actions the entry marks for approval, and every call is recorded.

Where that stands, in short:

  • Built In a lab: typed actions per app served as MCP tools, refusals and approvals, an audit log, and an agent on a small model operating four apps through their operate skills.
  • In progress A REST API and MCP tools for PaaSbox Clusters, with which an agent reads your clusters, takes snapshots and schedules upgrades.
  • In progress A read proxy that gives agents kubectl for reading, with secret values redacted.
  • Planned An operations agent that runs on its own schedule, gates before every update, live figures from real operations, and app operations by agents offered as a service of its own.

“In a lab” means virtual machines on a Mac running a real Kubernetes cluster. Nothing on this page is a production result, and no AI-automated operations are sold today.

pbx-agent, the agent on a PaaSbox Clusters node, is a program, not an AI model. It runs a fixed list of operations that the portal hands it, such as a snapshot, a restore or an upgrade, and nothing else. The word “agent” on this page means an AI agent.

In progress An agent can create, check and delete PaaSbox Clusters through a REST API and MCP tools, with a key that carries only the scopes you give it: Give an agent access has the steps, and Guardrails for agents what limits it.

Why an agent’s change is tested in a cluster built for the test, and why each stage and app gets a small cluster of its own, is on Why isolated clusters for agents.

Every app in the catalog has an entry, and the entry declares what may be done to the app as typed actions: a name, a title, typed parameters, and whether the action is destructive. Shortened, from the entry of Forgejo:

operations:
actions:
- name: reset-mfa
title: Remove a user's two-factor authentication
kind: job
danger: true
params: {username: {type: string}}
command: [forgejo, admin, user, reset-mfa, --username, "{username}"]
agent:
mcp: {server: actions}
skills: [skills/operate-forgejo]
routines:
- {name: update-check, schedule: "0 6 * * 1", skill: operate-forgejo, tools: write, approval: true}

Built In the lab:

  • The actions server serves an app’s actions to an agent as MCP tools, and nothing else: no shell, no kubectl writes. Parameters are checked against their types before anything runs.
  • Destructive actions wait for a person. The server refuses them unless it was started by a client that asks a person first. In the lab, a simulated approver stands in for that person, and its answers are recorded.
  • Every call lands in an audit log, the refused ones included: the action, its parameters, the reason the agent must give for a write or a destructive action, and the result.
  • The operate skill tells the agent how to operate this app: which action answers which question, and what to check before and after.
  • The harness has the agent operate the live app through its skill and the skill’s evals, records the approvals it asked for, and grades the run.

The routines an entry declares, such as a weekly restore drill, a weekly update check and a daily health digest, are Planned to run on a schedule; today an agent runs when a person starts it.

Letting an agent touch production is a question of trust. Four things earn it, and each part of PaaSbox belongs to one of them.

1. Guardrails: an agent can only do what it may, and you see what it does

Section titled “1. Guardrails: an agent can only do what it may, and you see what it does”
  • Built Every command prints a JSON plan before anything changes, and reports typed JSON progress.
  • Built Typed actions per app, served as MCP tools: destructive ones refused unless allowed, approvals where the entry asks for one, every call in an audit log. In a lab.
  • Built Network rules generated from each app’s connections, in a lab.
  • In progress A read proxy: agents keep kubectl for reading; Secrets are refused, secret values are redacted in every answer and log line, and every read is recorded. Built on a branch and tested on a throwaway cluster.
  • Planned An operations agent that may only do what a person on call may do.

2. Tested changes: nothing reaches production untested

Section titled “2. Tested changes: nothing reaches production untested”
  • Built GitOps: every change is a commit, rendered from one file per suite of apps.
  • Built An upgrade with a rollback for each of the four golden entries, with the data checked, in a lab.
  • Planned Staging, a promotion lane, gates before every deployment and update, a gradual rollout. An agent’s change passes the same gates as yours.

3. Operations built in: every app arrives ready to run

Section titled “3. Operations built in: every app arrives ready to run”
  • Built Apps connected to their database and cache, with credentials no other app can read, in a lab.
  • Built Backups with a tested restore: restored into a scratch copy, and a data marker read back, for each golden entry, in a lab.
  • Built An operate skill per golden entry, with which an agent operates the live app.
  • Planned Monitoring and alerts per app, fed by the platform’s metrics.

4. Measured in the open: the claim becomes a number

Section titled “4. Measured in the open: the claim becomes a number”
  • Built The lab round of 2026-10-09, below.
  • Planned Live figures from PaaSbox’s own operations: the share of tasks agents handled, recovery time, change failure rate, restore drills, measured availability.
  • Planned Evidence for CIS and BSI IT-Grundschutz controls; your auditor decides what it proves.

On 2026-10-09 the four golden entries, Forgejo, Vaultwarden, Listmonk and Umami, ran every lab row on the Mac lab: install, credentials, network rules, configuration, actions, alert, restore drill, upgrade, the agent’s skill run, and a second install.

  • Vaultwarden, Listmonk and Umami passed every row they have; the alert row was skipped where the catalog allows it.
  • Forgejo passed 12 of 13 rows. Its alert row cannot pass until the platform collects the app’s metrics. No alert has fired in the lab yet.
  • Entries broken on purpose fail the right row: credentials not passed to the app fail the binding row, a backup that leaves out the app’s volume fails the restore drill, a skill that names a missing action fails the skill row.
  • The agent is cheap to run: Claude Haiku operated each live app through its operate skill; the skill rows cost $0.058 in model spend over 12 recorded runs.
  • Still open: a final full run of every entry on a fresh node, the same round on a Hetzner lab project, and an alert that actually fires.

Operated in the open has the table, row by row.

  • In progress The read proxy gets a lab row of its own, whose negative control plants a known secret in a log line, an environment value and a ConfigMap and checks that it never reaches the agent. In the lab, an agent can then be given direct access or access through the proxy, and the two diagnoses compared.
  • Planned An operations agent. It reads through the read proxy and the platform’s metrics and logs, acts only through the entries’ typed actions, and waits for a person where an action is destructive or marked for approval. It runs PaaSbox’s own clusters first.
  • Planned Update gates: checks an update must pass before it reaches a running app, the same for an agent’s change as for yours.
  • Planned App operations by agents as a service of its own: agents operate your apps under a person’s approval, priced per cluster plus per operated app. It would be a separate offer: PaaSbox Clusters is software you run your clusters with, not a managed service. No prices and no date yet.
  • Planned A demo in your own Hetzner test project: a small cluster, two apps and your own agent fixing what a scenario breaks, through typed actions and with your approval, and a course built on it. Neither can be booked.

The pieces that make the claim checkable are to be open: the action runner, the actions server, the read proxy, the harness and the four golden entries. They are planned for the third wave of publication. The open components lists them with their licences.

I build PaaSbox with agents too. They write code against a written spec, run the lab rows and drills, check dependencies and images, and draft the analysis when something breaks. Their output is a pull request, a lab record or a ticket. I decide, I merge, and I am the only one who touches a live system. Operated in the open has the charter.