deploy/compose/.env for Docker Compose (seeded
from .env.example on first make all). Defaults below are what the stack falls back to when the
variable is unset — fine for local evaluation, not for anything exposed.
Transcription (STT)
The STT service is the GPU workload, deployed separately from this stack — see
Deployment → Transcription. The main stack
stays GPU-free and reaches it over the network via the two variables above.
When the backend is wrong, you learn at configure time
A URL or token that does not work is refused where you set it, not by an empty transcript hours later. The same live check runs at three surfaces, and all three name the same cause:
The verdict is cached for 60s, so after fixing the value either wait a minute or call
GET /health?force=1 to re-probe immediately.
A
404 from the transcriptions path is treated as a wrong URL, not as proof of life: a real
OpenAI-compatible endpoint answers 400 or 401 to an empty request body and never 404. Some
gateways also answer 404 for a rejected credential — the error text names both possibilities.Optional service authority
Stock self-hosted Vexa has no external service authority and continues to admit bots exactly as before. Operators who need an independently owned admission/active-service policy can opt into the sealedservice-authority.v1 boundary:
The request contains service identity, authoritative user ID, service mode, frozen transcription
provider, concurrency, and lifecycle timestamps. It deliberately contains no email, customer or
payment-provider identifier, price, balance, transcription endpoint URL, or credential. A
configured authority that is unavailable, malformed, stale, or bound to another request fails
closed before a new bot is spawned. During active service, its stop intent is committed before
runtime teardown and resumed after restart.
mode: "observe" records decisions without refusing or stopping service. It is useful for rollout,
but it is not evidence that a hard service control is enforced. A session admitted in one mode is
never reinterpreted under another mode after a deployment change.
Operator terminal callback
An operator can send terminal service facts to one deployment-owned system endpoint. This is separate from a user’s webhook configuration: the URL is frozen at boot, cannot be supplied by an API caller or meeting payload, and receives onlymeeting.completed and bot.failed.
Transient delivery failures enter a dedicated retry/dead-letter queue and retain the same
deterministic event ID. Customer webhook destinations continue through DNS/IP validation and cannot
redirect this operator callback. Leave all four values unset for the stock OSS behavior.
Secrets & identity
Bot participant name
DEFAULT_BOT_NAME on meeting-api names the bot, and compose, Lite and helm ship it as Vexa. To
rename the bot, set it in .env (compose/Lite) or meetingApi.defaultBotName (helm) and restart
meeting-api — the terminal needs no rebuild, because it omits bot_name when it sends a bot and
lets the deployment decide. Blank the variable and meeting-api falls back to VexaBot-{random}.
Two things still win over it. An explicit bot_name in a POST /bots request always takes
precedence, which is also how a calendar connection applies its own per-calendar bot name. And a
name baked into the terminal image at build time overrides the deployment: pass
NEXT_PUBLIC_DEFAULT_BOT_NAME as a Docker build arg (there is a commented-out line for it under the
terminal’s build.args in compose) and rebuild the image — a rename then costs another rebuild,
which is why the runtime variable is the better knob. Direct joinMeeting callers read
DEFAULT_BOT_NAME too and fall back to Vexa Join Layer.
Database & storage
DB_POOL_SIZE / DB_MAX_OVERFLOW are pure runtime overrides. On managed Postgres with a hard
max_connections, reconcile them with the per-service connection budget in deploy/db-budget.json,
whose accounting is Σ (replicas × (pool_size + max_overflow)) + reserved ≤ max_connections.
The defaults (5 / 10) match that budget.Object storage (recordings)
Recordings are stored over S3. Lite and Compose run the store themselves: versitygw (versity/versitygw, Apache-2.0), whose POSIX backend keeps each object as a plain file in a
local volume. The variables keep their historical MINIO_* names; no MinIO runs in Lite or
Compose.
Upgrading an install that ran MinIO: Upgrade from MinIO. Backups:
Recordings → Backups.
Object storage on Kubernetes
On Helm, meeting-api’s storage variables come from the chart’sstorage.s3 values — see
Kubernetes → Object storage. The chart refuses to render
when meetingApi.extraEnv also sets one of them.
Agent inference (bring your own)
Point the agent at your own model so no inference leaves the network.Claude subscription credentials (HOST_CLAUDE_CREDENTIALS)
Setting HOST_CLAUDE_CREDENTIALS=~/.claude/.credentials.json mounts your Claude Code
sign-in into the agent containers (read-only), so the agent runs on your subscription
instead of an API key. Whether that file stays valid depends on the host OS:
-
Linux (and Windows via WSL2 — run the stack and the
claudeCLI inside WSL): the file is Claude Code’s own store; the CLI refreshes it in place. Nothing to do. -
macOS: the CLI’s source of truth is the login Keychain — the file is a one-time
export whose token expires every ~8–12 hours. Symptom: agent chat fails with
401 Invalid authentication credentialswhileclaudeworks fine in your terminal. Install the bundled sync daemon once:It registers a launchd user agent (ai.vexa.claude-creds-sync) that re-exports the Keychain into the file every 5 minutes, write-only-on-change, preserving the inode the containers mount.install.sh uninstallremoves it. Details:deploy/bin/claude-creds-sync/README.md.
Auto-join & calendar sync
Timing knobs for the two meeting-api background sweeps (auto-join and calendar sync). Both degrade gracefully: withoutADMIN_API_URL +
INTERNAL_API_SECRET, calendar sync no-ops and auto-join spawns without per-user context — the
stack still boots.
Images & runtime
Ports
Host ports the compose stack publishes on127.0.0.1 (override any of them in .env):
The gateway (
:18056) is the one front door; the terminal web workbench is at :13000. The other host
ports above are bound to 127.0.0.1 for local inspection and aren’t needed for day-to-day use.Gateway edge protection
The gateway carries two independent abuse layers. Both are env-driven; both default sensibly for self-hosted and can be left untouched.Per-user limiter (post-auth, on by default)
A token-bucket limiter keyed by user id fires after the API key is resolved — it catches one token driving too much traffic. On by default.Edge guard (pre-auth, off by default)
An optional fastapi-guard edge layer caps requests per client IP before the API key is validated — so an IP flooding invalid keys, or rotating many keys from one IP to defeat the per-user limiter, is answered with429 at the edge and never reaches admin-api. An IP that keeps offending past a threshold
is auto-banned for a window; sustained abuse costs the abuser, not the operator. It is on by default
everywhere: the code default is GUARD_ENABLED=true, compose interpolates
GUARD_ENABLED=${GUARD_ENABLED:-true} and ships the opt-out line commented, and the Helm chart sets
gateway.guard.enabled: true. You opt OUT, not in.
GUARD_ENABLED=false is the kill switch — set it to turn the whole edge layer off. When
behind a reverse proxy you must also set GUARD_TRUSTED_PROXIES, or every request keys to the proxy
IP → one global bucket shared by all clients (one abuser throttles everyone). See
Deployment → Publishing behind a reverse proxy.With
GUARD_ENABLE_REDIS=false AND uvicorn --workers N>1, the HTTP rate-limit buckets and auto-bans are also per-process: the effective limit becomes N × GUARD_RATE_LIMIT_RPM and bans do not propagate across workers (the WS per-process ceiling above applies to the WS path independently). The same applies across multiple gateway replicas without Redis — each is its own process.Safe configurations: the shipped gateway runs a single uvicorn worker, so the default (GUARD_ENABLE_REDIS=true, one worker) enforces the limit globally. If you scale to multiple workers or replicas, keep GUARD_ENABLE_REDIS=true — Redis-backed state is what shares the buckets and bans across processes. GUARD_ENABLE_REDIS=false is only safe with a single worker and a single replica; if you must run without Redis at higher concurrency, divide GUARD_RATE_LIMIT_RPM by the process count to hold the aggregate cap (note this still will not propagate bans).fail_secure=false): a guard-check bug or a redis outage returns the request to
the app rather than taking the gateway down. Request-body WAF scanning is intentionally off — the
gateway proxies arbitrary user text (chat, meeting data, transcript shares), so signature scanning
would false-positive on legitimate content.