Tracing for AI applications
A ManagedTracing is an OTLP endpoint for the calls your app makes to an LLM: point your model client’s OpenTelemetry exporter at it and every prompt, completion, token count and latency lands somewhere you can see it — OpenLIT (Apache-2.0) over ClickHouse, rendered directly with no separate operator to install.
Turning it on
Section titled “Turning it on”Tracing is a capability, off by default, enabled per shoot in the paasbox-paas extension’s config — the same block paasbox garden shoot create writes into your shoot manifest:
apiVersion: core.gardener.cloud/v1beta1kind: Shootspec: extensions: - type: paasbox-paas providerConfig: apiVersion: paas.paasbox.com/v1alpha1 kind: PaasConfig capabilities: [apps, postgres, valkey, tracing]Add tracing to the list and re-apply the Shoot. Enabling it costs nothing by itself — no operator is installed, only the object and a reconciler that is already running. The cost starts with your first ManagedTracing.
Creating one
Section titled “Creating one”apiVersion: paas.paasbox.com/v1alpha1kind: ManagedTracingmetadata: { name: traces, namespace: shop }spec: plan: s store: { mode: managed }status: conditions: [Reconciled, Ready, Degraded, StoreHealthy, SpansDropping] phase: Ready endpoint: otlp: http://traces.shop.svc:4318 ui: http://traces.shop.svc:3000 secretName: traces-otlpkubectl apply -f traces.yaml gets you a Ready instance in about a minute — measured 52 seconds on a paid Hetzner shoot at plan s (2026-09-16). plan sizes the OpenLIT pod; store.mode: managed renders a single ClickHouse instance alongside it, in the same namespace, that nothing else reads. retention (default 30d) is the only bound on how much a ManagedTracing keeps: ClickHouse has no per-instance quota on stored bytes.
Pointing your app at it
Section titled “Pointing your app at it”# on the Appuses: [{ name: traces, kind: ManagedTracing }]injects five environment variables: OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, OTEL_EXPORTER_OTLP_TRACES_PROTOCOL (http/protobuf) and OTEL_EXPORTER_OTLP_HEADERS (an x-api-key) as references into traces-otlp, plus two literal values: OTEL_BSP_MAX_EXPORT_BATCH_SIZE=512 and, the one that matters, OTEL_BSP_MAX_QUEUE_SIZE=32768. Any OpenTelemetry SDK that reads the standard OTEL_EXPORTER_OTLP_TRACES_* variables — which includes OpenLIT’s own and most LLM tracing libraries — needs nothing else. A prefix on this dependency is accepted but does nothing: the names are fixed.
Why the queue size is not optional
Section titled “Why the queue size is not optional”An OTel SDK batches spans in memory before sending them, in a queue with a fixed size; once it is full, the SDK drops new spans itself, before they ever reach the network. Most SDKs default that queue to 2048. Measured on a paid shoot (2026-09-16): a burst of 10,002 spans against a ManagedTracing with the SDK’s default queue sent only 2,688 to the store — 7,314 spans, 73%, were dropped inside the app’s own process, while the store’s own refusal counter read zero the whole time. Nothing on the receiving end saw a problem; the spans were gone before they left the pod.
That is why uses sets OTEL_BSP_MAX_QUEUE_SIZE=32768 for you automatically. Under a clean, non-bursty load the same setup lost nothing (1,002 sent, 1,002 stored). The trap is specifically a burst against the SDK default — and it is silent, because status.drops.storeRefused looks fine the entire time.
Watching for drops
Section titled “Watching for drops”status.drops carries both counters, because only one of them is real backpressure:
storeRefused— spans ClickHouse actually refused. A non-zero value here is the store under load.clientDropped— spans your app’s own SDK threw away before sending, read back over the same OTLP endpoint. This is the number the queue-size trap shows up in.
The SpansDropping condition goes True when either counter is climbing, so you do not have to poll both numbers by hand to notice.
What it costs
Section titled “What it costs”plan sizes OpenLIT; ClickHouse (mode managed) is not plan-sized — it always requests 600Mi with a 2Gi limit, because that headroom is for queries, not for a bigger instance:
| Plan | OpenLIT CPU | OpenLIT memory | ClickHouse storage |
|---|---|---|---|
xs | 100m | 256Mi | 5Gi |
s | 250m | 512Mi | 10Gi |
m | 500m | 1Gi | 20Gi |
l | 1 | 2Gi | 50Gi |
Add ClickHouse’s fixed 600Mi request to the plan’s own memory for what one instance reserves: 1,112Mi on plan s, 856Mi on xs — both measured twice, independently, on a paid shoot (2026-09-16). That is more than the platform’s own eleven operators reserve together. The portal is meant to show this number before the toggle; until it does, this page is that number.
Storage has the same kind of surprise. OpenLIT’s own data volume is a fixed 2Gi request, regardless of plan. Hetzner’s block storage never provisions below 10Gi, so that 2Gi request becomes a 10Gi volume — and on plan s, where ClickHouse’s own 10Gi request already sits exactly at that floor, one ManagedTracing instance bills 20Gi of block storage, not the 12Gi its two requests add up to. The same 10Gi-per-volume floor applies at every plan; only m and l ask ClickHouse for more than the floor itself.
Never scale it
Section titled “Never scale it”There is no replicas field on ManagedTracing, on purpose. OpenLIT’s own identity state — organisations, projects, API keys, the config that keeps tenants apart — lives in a single SQLite file on one read-write-once volume. A second pod either fails to schedule or corrupts that file. One instance is meant to serve a whole team through its own dashboards, not to be scaled like a stateless app.
Seeing it
Section titled “Seeing it”There is no public route to the OpenLIT UI yet — status.endpoint.ui is an in-cluster address on purpose. Reach it from your machine with:
kubectl -n shop port-forward svc/traces 3000:3000Where to go next
Section titled “Where to go next”- Deploy an app:
uses, and everything else anAppaccepts. - Example: a Django SaaS: the same app with tracing wired in.
- The pb CLI:
pb statuslistsManagedTracingalongsideApp,ManagedPostgresandManagedValkey.