Skip to main content
The Helm chart at deploy/helm/charts/vexa deploys the same control plane as the Compose stack, with the runtime pointed at the cluster: RUNTIME_BACKEND=k8s makes every bot and every agent dispatch its own bare Pod (restart: Never), so capacity is your cluster’s scheduler — not a bigger box.

What the chart deploys

Hardened defaults throughout: containers run non-root with all capabilities dropped, PodDisruptionBudgets on the control plane, and rolling updates that never take the last replica down.

Install

values-staging.yaml and values-test.yaml show environment overlays. Chart tests (helm lint + template tests) live under deploy/helm/tests/. Don’t have a key yet? Hosted: sign in at vexa.ai/signin with a Google account and copy your key from your account page — free credit, no card required. Self-hosted: make all prints a key when the stack comes up.

Is this install actually working? — make probe SURFACE=helm

The standing full-journey smoke probe drives spawn → schedule → boot → join → transcribe → live-view → stop through the gateway front door, then sweeps every deployment’s logs once. Each stage prints Expected / Actual / Verdict; a red stage names where the journey broke and fails the command, so you get a truthful install verdict in minutes with no meeting and no humans.
It mints its API key from the release secret’s ADMIN_API_TOKEN (or pass VEXA_API_KEY). The same journey runs on the other surfaces: make probe (compose, the fast default) and make probe SURFACE=lite.

The runtime on Kubernetes

The runtime kernel owns one runtime.v1 lifecycle (starting → running → stopping → stopped → destroyed) and delegates only the substrate question — how do I start, observe, and stop a workload? — to the backend. On k8s, a workload is a bare Pod; the same dispatch, contracts, and worker run identically to the Docker backend. See Execution.

Scheduling spawned Pods on a tainted / dedicated pool

A spawned bot or agent Pod is created with kubectl run — it is not a child of the runtime Deployment, so it does not inherit global.nodeSelector / global.tolerations. On a cluster whose nodes are all tainted (the standard pool-segregation pattern — dedicated, spot/temp, or GPU node pools), a spawned Pod with no tolerations sits Pending forever (untolerated taint(s)) and the meeting silently fails while the runtime itself — a Deployment that does apply those constraints — looks healthy. The chart closes this by handing the runtime its own scheduling constraints as env, which the spawn backend stamps onto every Pod it creates: By default a spawned Pod schedules wherever the runtime itself can — set global.tolerations / global.nodeSelector for your pool and both the runtime and everything it spawns land there, with no extra action:
Set runtime.tolerations / runtime.nodeSelector only to send spawned workloads to a different pool than the kernel (e.g. GPU bots). Both values are serialized to JSON; malformed JSON fails the spawn loudly rather than dropping the constraint (a dropped constraint is the stranded-Pod bug).
Spawned-workload commands must exist in the target image under k8s (entrypoint-replace) semantics. A runtime profile may carry a command; the k8s backend passes it as kubectl run --command -- …, which replaces the image ENTRYPOINT (it becomes the Pod’s argv[0]). So a command path that is not present in the image StartErrors every spawn on k8s — even though the Docker backend, which appends the command to the entrypoint as arguments, silently ignores the same bad path. The meeting-bot profile therefore carries no command: the bot image’s own ENTRYPOINT (/app/entrypoint.sh) boots it on every backend. At release time an image↔profile conformance gate (image-identity) asserts each profile command is either empty or an in-image executable against the published images, so a profile that diverges from its image fails the release rather than a customer’s install.

Autoscaling and bot bursts — the fresh-node reachability gate

Because every bot is its own bare Pod, a burst of meeting requests can push the cluster autoscaler to add nodes, and bots get scheduled onto them the moment they report Ready. But a brand-new node’s network can take up to a few minutes to fully converge after Ready — the CNI/NetworkPolicy programming, kube-proxy, and DNS warmup all lag node readiness. A bot scheduled into that window can find its control plane (its meeting-api callback URL and redis) unreachable on its first outbound hop. The bot handles this with a pre-join reachability gate: it makes its first joining lifecycle emit load-bearing. If the meeting-api callback is reachable (any HTTP response), it joins immediately — zero added latency. If not, it probes redis; if either channel is up it proceeds (it can still report), and only if both are down does it refuse to join, terminating fast (< a few seconds) with:
  • Pod exit code 3 — the attributable terminal signal (kubectl describe pod <bot>Last State: Terminated, Exit Code: 3). Distinct from a real join failure (exit 1).
  • a failed lifecycle event carrying failure_stage: requested and infra_fault: control_plane_unreachable on whichever channel recovers (often none — the exit code is the signal).
This converts an opaque CrashLoop / stuck-requested meeting into a fast, attributed failure an operator can read in one kubectl describe. See Troubleshooting → exit code 3. What stays cluster-side (not solved by the gate): the node-readiness window itself. Consider pre-pulling the bot image onto new nodes (a cold pull can itself take minutes and mask the network window), and treat node Ready as insufficient for scheduling network-dependent work — the gate fails fast and, on k8s, the exit lets the operator (or an autoscaler policy) reschedule onto a warmed node.

Scaling meeting-api

meetingApi.replicaCount defaults to 2, and that is safe for the segment consumer. Each replica now joins the transcription_segments consumer group under a per-pod identity (collector-<hostname>, overridable with COLLECTOR_CONSUMER_NAME), and a surviving replica periodically XAUTOCLAIMs a crashed replica’s un-acked segment batch and drains it through the normal persist path — so no transcript segments are orphaned when a pod dies mid-batch. The reclaim only fires once a batch has idled past COLLECTOR_RECLAIM_MIN_IDLE_MS (default 60000), so a live peer’s in-flight batch is never stolen. /health exposes pipeline.pending_depth (the group’s delivered-but-un-acked count) alongside pipeline.consumer_lag, and degrades to 503 once it exceeds PIPELINE_PENDING_ALARM (default 100) — a stuck batch is a reportable state, not a silent stall. A pod’s hostname changes on every recreate (a rolling deploy, a pod restart, a compose --force-recreate), so each new container joins collector_group under a new collector-<hostname> and the old name is left behind. Those abandoned consumers hold no pending entries (they read their last message and acked it), so they are harmless to correctness — but left unchecked the group fills with dead names that inflate operator XINFO CONSUMERS reads and muddy /health. The same reclaim sweep now prunes them: any consumer with pending == 0 that has idled past COLLECTOR_CONSUMER_TTL_MS (default 1800000, 30 min — well above the reclaim idle gate, so a briefly-quiet live replica is never touched) is removed with XGROUP DELCONSUMER. A consumer holding an in-flight batch (pending > 0) is never pruned, and the running replica never prunes itself; a live consumer re-registers on its next XREADGROUP, so the prune is idempotent and self-healing. Reading XINFO CONSUMERS collector_group should therefore list only the replicas that actually exist — dead collector-<oldhash> names are expected to disappear within a sweep of the TTL. For k8s you can eliminate the churn at the source by pinning a stable per-replica identity via COLLECTOR_CONSUMER_NAME (e.g. a StatefulSet ordinal) so a restart re-uses the same consumer.
Every surface now ships Valkey 8.x, which has XAUTOCLAIM — so orphan reclaim is active everywhere Vexa ships, including Vexa Lite (compose, helm, and Lite all run Valkey, the Linux Foundation BSD-3 fork of Redis 7.2.4; see the changelog and #653). Reclaim, per-pod identity, and ghost-consumer pruning all work on the shipped stack. Check yours with valkey-server --version.The degradation path below still exists as a safety net for a bring-your-own backing store older than the floor: on a Redis/Valkey without XAUTOCLAIM the reclaim disables itself and logs once at startup; everything else (normal XREADGROUP consumption, pending_depth) is unaffected. The cost is narrow but real — if a replica dies mid-batch on such a store, that batch stays pending instead of being drained by a peer, so run a single replica or move to a store with XAUTOCLAIM (Redis ≥ 6.2 / Valkey ≥ 7.2). Ghost-consumer pruning degrades the same way: if the backing store rejects XINFO CONSUMERS, the prune logs once and no-ops, leaving the consume path untouched.

In-cluster self-addressing

Under Helm the meeting-api Service is release-qualified (<release>-vexa-meeting-api), not the bare meeting-api that the Compose stack uses. So the chart sets MEETING_API_URL explicitly on the meeting-api deployment to http://<release>-vexa-meeting-api:8080 — the address a spawned bot calls back on for its lifecycle callback and recording upload. Left unset it would fall back to the compose-only default http://meeting-api:8080, which does not resolve in-cluster. The chart sets the same address as VEXA_MEETING_API_URL on the agent-api deployment. Before opening a live-transcript SSE stream (GET /api/meeting/stream), agent-api verifies the caller owns the meeting by calling GET /meetings/{id} on meeting-api, and fails closed — an unreachable meeting-api returns 403 "not authorized for this meeting" for the meeting’s own owner. Left unset it would fall back to the compose-only http://meeting-api:8080, so the owner sees a permanent “Reconnecting to live stream…” instead of streaming words. The chart likewise sets ADMIN_API_URL on the meeting-api deployment to http://<release>-vexa-admin-api:8001 (mirroring Compose’s ADMIN_API_URL=http://admin-api:8001). This is required for two background loops: calendar sync (discovering each user’s connected ICS feed through admin-api’s internal edge) and capped auto-join (fetching the per-user max-bots cap for every scheduled-meeting spawn). With ADMIN_API_URL unset, calendar sync no-ops and the auto-join sweep cannot resolve the per-user cap. The meeting-api then fails closed and refuses to spawn rather than spawn uncapped; set AUTO_JOIN_ALLOW_UNCAPPED=1 only if you deliberately want uncapped auto-join spawns on a self-host.

Exposing the cluster — ingress and TLS

The chart ships an optional Ingress, off by default. Enabling it fronts the Terminal, which is the human entry point: it proxies /ws to the gateway and /api to agent-api/admin-api server-side, so one hostname is usually enough.
tls is passed through to the Ingress spec.tls verbatim, so it is ordinary Kubernetes TLS — the chart neither issues nor terminates certificates itself. Either let cert-manager populate secretName via the annotation above, or create the Secret yourself:
annotations are rendered onto the Ingress unchanged, which is also where controller-specific settings go — body size, proxy timeouts, and the long-lived-connection settings a WebSocket upgrade needs on /ws.

Exposing the raw API as well

paths defaults to a single rule sending / to the Terminal on port 3000. Add a second rule if external clients should reach the gateway directly:
A path entry that omits service defaults to the gateway.

Tell the Terminal its external origin

Enabling the Ingress is not sufficient on its own. terminal.publicUrl is empty by default — correct for cluster-internal or port-forward access — and it is what sets NEXTAUTH_URL and TERMINAL_URL. Left empty behind a public hostname, cookie security and the OAuth callback are computed against the wrong origin:

Sign-in providers — Google and Microsoft

The Terminal’s Google and Microsoft buttons each self-gate on their own credentials: set both the client id and the client secret and that button appears; leave either empty and it stays hidden. The chart has no dedicated fields for them — pass them through terminal.extraEnv, which is rendered onto the Deployment verbatim, so valueFrom works where a literal would put a secret in your values file:
Register the redirect URI in the provider console as <publicUrl>/api/auth/callback/microsoft (or .../google) — so terminal.publicUrl above is a prerequisite, not an optional extra.
Without an Ingress the Services stay cluster-internal. helm install alone does not expose Vexa — reach it with kubectl port-forward during evaluation, or enable the Ingress above.

The deprecated dashboard (optional)

The 0.10 dashboard — the multi-user web UI that predates the Terminal — ships as an off-by-default component while it is still load-bearing: hosted production runs it against the 0.12 core, and the authenticated-session flows are only walkable through it today. Enable it explicitly:
It runs the pinned external image vexaai/dashboard on its own tag — global.imageTag does not apply to it — and talks to the gateway’s hosted-compat surface. It is deprecated: when the Terminal covers its remaining flows, the component is deleted, not ported. On Compose the same option is docker compose --profile dashboard up -d (port 13001).

Honest status

The k8s backend lifecycle and the workspace mount are implemented (the workspace store binds into each worker Pod as a PVC, scoped per-mount). The Compose stack is the path with the most production mileage today; the chart is the right starting point for a cluster evaluation — track the status page.

No-egress clusters

Everything the chart deploys pulls from images you build and host in your own registry; pair it with the self-hosted transcription unit — bundled GPU or CPU — and your own LLM endpoint (Configuration) so the audio path never calls out. See Security & compliance. One gap on this path, and it is this chart’s: TRANSCRIPTION_MODEL is not values-plumbed — set it on the meeting-api and terminal Deployments via extraEnv until it lands, which the chart’s own known boundaries state. Compose and Lite plumb it as a first-class setting. So no egress is ✅ on Compose and Lite and 🟡 here. We have not run an install on a genuinely disconnected network, so we do not claim one: what a procurement review can hold us to is self-hosted with no egress, not a certified air gap.