Base footprint
Hybrid
Minimum footprint (the installer shows the same numbers, published at charts.xpander.ai/requirements.json):
Agents are autonomous: tasks come from schedules, webhooks, the SDK and API, and other agents, not only chat (in xpander cloud production, only ~5% of executions are interactive chat). Size for the number of tasks running concurrently at peak, whatever triggers them.
Capacity figures include ~50% headroom over pod resource requests for burst load, Kubernetes system pods, and node upgrades.
Sizing calculator
Take your own numbers: peak concurrent tasks, agents, and how many of them use a workspace. The estimate follows the model described under it.How the estimate works
How the estimate works
The calculator applies the same model we use to size xpander cloud:
- Concurrent tasks is the primary driver. A task is any agent execution, from any source: chat, API/SDK calls, schedules, webhooks, MCP clients, Slack, or another agent. If you don’t know your peak concurrency yet, assume one concurrent task per agent - that covers most production organizations we measured. In practice most teams peak at just a few concurrent tasks (typically under 5), and even the busiest organization peaks around 30.
- Execution replicas (
agent-worker) = concurrent tasks ÷ 4 (each replica processes 4 tasks by default, configurable withMAX_CONCURRENT_EXECUTIONS). Each replica requests 2 vCPU / 2.5 GiB. - Agent Controllers: 1 replica up to ~8 concurrent tasks, 2 up to ~20, 3 beyond. Each requests 1 vCPU / 2 GiB.
- Active workspaces = concurrent tasks × workspace share, capped at the agent count. Budget ~0.5 vCPU / 1 GiB per active workspace (measured typical load; each can burst to its 2 vCPU / 3 GiB limit).
- Fixed base (~2 vCPU / 6 GiB): AI Gateway, MCP, Chat, Code Runner, AWS Operator, API, Redis, PostgreSQL, container registry, metrics-server at one replica each.
- Cluster capacity = requests × 1.5 for burst, DaemonSets, and safe node drains. Storage = 48 GiB of chart volumes + 5 GiB per agent (each agent that uses a workspace keeps a persistent volume) + image cache per node.
Per-service requirements
Per-service requirements
Per-replica values. Defaults are what the Helm chart applies out of the box; recommended raises the values we found low against measured production usage.
Override any of these in your Helm values (keys are camelCase):
Air-Gapped
The full prerequisite list (network, license files, registry, tooling) is on Air-Gapped: Prerequisites.
Sizing tiers
The Air-gapped deployment runs everything in-cluster: the execution tier (theagent-worker service) plus the control plane services, the web UI, the Supabase stack (its own PostgreSQL 15, auth, storage, edge functions), and MongoDB.
Storage breakdown at the production tier: Supabase PostgreSQL 50 GiB, xpander’s PostgreSQL 50 GiB, Supabase Storage 50 GiB+, MongoDB 20 GiB, Redis 8 GiB, registry ~50 GiB, mirror caches 50–100 GiB. All RWO, one StorageClass (the chart’s preflight verifies it exists).
External data tier sizing. You can bring managed stores instead (Data Tier). A reference customer install runs on these external stores:
- RDS PostgreSQL 16 on
db.t4g.smallwith 50 GiB gp3 - MemoryDB on
db.t4g.smallwith one shard and one replica - S3 as the Supabase Storage backend for blobs
agent-controller, agent-worker, api, agents, functions.
Scaling with usage
The chart ships fixed replica counts and no autoscalers: scaling policy is yours. What actually grows with usage, and how to scale each:
Node capacity: workspaces and code-execution Jobs create pods on demand, so the real ceiling on concurrent agent activity is how many nodes you have. Two workable postures:
- Static capacity: size the node group for your peak concurrent workspaces and leave it fixed. Simplest, fully offline, and the recommended posture for sealed networks.
- Cluster Autoscaler: adds nodes automatically when pods can’t schedule. It works in a no-internet VPC, but it talks to your cloud’s scaling API to do its job, so that API must be reachable from inside the network (on AWS: the
autoscalingandec2VPC endpoints).
Agent workspaces
Each agent that uses the agent workspace gets a dedicated pod, created on demand and reclaimed when idle. Size the cluster for the number of concurrently active workspaces, not total agents.
Idle workspaces consume almost nothing (~50 Mi observed); an active workspace typically uses ~0.5-1 vCPU and up to 1 GiB, bursting to its limits for heavy work. The workspace volume persists after the pod is reclaimed, so a returning agent keeps its files - plan 5 GiB of storage per workspace-using agent, not per pod.
The workspace image is ~1 GiB. The first workspace start on each node pulls it; subsequent starts on that node are fast. Budget ~2 GiB of node disk for image cache.
The fleet is always rendered by both charts (there is no enable flag) and starts warm with one executor (
fleet.replicas: 1 on Hybrid, global.fleet.replicas: 1 on Air-Gapped). The Agent Controller sends a turn to the fleet whenever a live executor is registered. Keep it warm: the first turn of any agent is then seconds away, and idle agents cost only their directory. Where a cluster cannot host the privileged executor, set replicas: 0; on a live cluster the chart also idles the executor by itself when its StorageClass is missing (see Idle guard).The executor fleet
The fleet is afleet-executor StatefulSet. Each replica is one pod with two privileged containers over one XFS volume: the executor, which runs turns, and a mounter sidecar, which binds each agent’s directory into the executor for the duration of a turn and enforces the per-agent disk quota. An agent’s directory lives on exactly one executor’s volume, and its turns always run there.
Per Executor
Per Turn
Each turn runs in its own cgroup. These are the ceilings the executor applies. An agent’s owner can raise them for that agent under Agent settings, up to xpander’s caps.The comment above the
fleet block in both charts’ values.yaml still quotes older per-turn defaults (“2Gi / 2 vCPU / 256 pids for a CLI turn”). The values in the table are what the current agent-harness image applies; the chart comment is being corrected.Concurrency Caps
Reading the runtime logs
Two logs carry one turn: the executor that ran it and the Agent Controller that dispatched it. Follow both while you send a message:fleet-executor-<n>, container executor). One turn logs these events, in this order:
Agent Controller (
deploy/agent-controller). The same turn shows as:
direct turn conversation=<id> task=<id> status=... received_to_enqueued_ms=... enqueued_to_claimed_ms=...: the message became a task and was picked up by the fleet;enqueued_to_claimed_msis where a cold executor shows. The/llm-proxylines below name the model that served the turn (<host>:<port>/<model id>for a custom provider), the model the conversation picked; the agent record’s model is only the default for new conversations.event=fleet.claimed executor=fleet-executor-0: which executor took it.POST /llm-proxy/<provider>/model/<model>/invoke-with-response-stream: one line per model call the CLI made. The path names the provider and model that served the turn.event=compute_usage.emitted cli=claude-code cpu_seconds=...: the compute record written when the turn ended.
Node Prerequisites
The installer (install.sh) and the Air-Gapped preflight hook check all of these and warn (non-blocking) when one is missing; the executor then fails to schedule and the fleet has no live executor. A cluster that fails any one of them cannot run agent turns until it is fixed, so treat the list as required.
Running the executor fleet without privileged containers
The executor fleet can run with no privileged container. Setfleet.isolation: userns on Hybrid, or global.fleet.isolation: userns on Air-Gapped. The default stays privileged, and nothing changes until you set it.
In this mode the executor is a Kubernetes user-namespace pod (hostUsers: false): root inside the pod is an unprivileged user on the node. It is a single container, with no mounter. Each agent’s work still runs in its own sandbox, as the agent’s own user, and one agent cannot see another’s files.
Agent traffic leaves the executor only through a filtering proxy. The proxy refuses the cloud metadata endpoint and private addresses. It allows the hosts the platform needs, such as the Agent Controller and your package mirrors. To let agents reach an internal host such as your git server, add it to fleet.egress.allowHosts. A corporate proxy set in HTTPS_PROXY on the executor is used for outbound traffic.
Security profile
Pick one of two pod profiles withfleet.seccompProfile.type:
The
Localhost profile must be present on every node that runs an executor. Install it in one of two ways:
- Security Profiles Operator: set
fleet.seccompProfile.operator: trueand the chart creates the profile for you. - Node file: place the chart’s profile at
/var/lib/kubelet/seccomp/xpander-fleet-pod.jsonon each node.
Requirements
The installer and the Air-Gapped preflight check these before the executor is rolled out.
Limits
Per-turn memory and process limits are enforced by the platform: a turn that goes over its ceiling is stopped on its own, and the rest of the executor keeps running. If your nodes offer a runtime class that makes the pod’s cgroup writable (containerd’scgroup_writable option), set fleet.runtimeClassName to it. The kernel then enforces the limits exactly as in the default mode.
Autoscaling
Scaling is driven by concurrent turns, never by the number of agents. Withfleet.keda.enabled: true the chart renders a KEDA ScaledObject on the StatefulSet:
Because there is exactly one executor per node,
maxReplicas can never usefully exceed the number of nodes that fit an executor. Grow the node group and raise maxReplicas together. Without KEDA, set fleet.replicas to the number of executors you want and scale it by hand; on a sealed network a fixed replica count sized for peak is the recommended posture.
Safe Rollout, Drain, and Upgrades
A roll never kills a running turn
A rolling update, a KEDA scale-in and a plain pod delete all end in the sameSIGTERM to the executor pod, and the pod treats it as a drain, not a kill:
- The executor stops claiming work, and its heartbeat says
draining. The Agent Controller routes around it: nothing new lands on it, and whatever was queued for it goes back to the shared queue. It still counts as fleet presence, so new turns wait for the replacement instead of starting anywhere else. - The turns in flight keep running for up to
fleet.executor.drainSeconds(3600 s by default). A turn still running when that window ends is stopped the way a user stop is: the CLI is interrupted, its session is kept, and the task endsstoppedwith “The executor is restarting; send the message again to continue where it left off.” Sending the message again resumes the session on the new pod. No task is left hanging as executing. - The mounter sidecar waits on the executor’s health endpoint until it reports draining with no turn in flight, and only then unbinds the agent directories. The pod grace (
fleet.executor.terminationGraceSeconds, 7500 s: the CLI run budget plus five minutes) is a ceiling, never a wait; the pod leaves as soon as both containers exit.
- Raise
drainSecondstoward the run budget if your turns run long; it must stay at least 120 s under the pod grace (the Air-Gapped chart refuses to render otherwise). - Never shorten a termination with
--grace-period=0or--forceon a pod that still has turns; that is the one way to kill a turn. - Plan image rolls for quiet hours when you can; they are safe at any time, only slower under load.
- The drain needs an
agent-harnessimage that carries it; an older executor ignores the setting and exits as before.
Idle guard when the StorageClass is missing
The executor becomes Ready only when its node can run the privileged pod and its volume binds on an XFS +prjquota StorageClass. So that an install or upgrade never waits on an executor that cannot bind, both charts guard themselves on a live cluster:
helm template and dry-runs always render the configured replicas; the guard reads the live cluster only. Create the class (or point provisioner at your cluster’s CSI driver, or bring your own class) and the next upgrade brings the executor up. Nothing else about the release changes while the executor idles: turns queue until an executor registers.
--wait, --atomic and the installer
helm upgrade --wait (and --atomic, which implies it) waits for the executor StatefulSet like any other workload. Two consequences:
- A running turn holds the roll. With
--wait, an upgrade during a long turn can take up to the drain window plus pod start. Give--timeoutat least the drain window (3600 s by default) or upgrade in a quiet period. - An idle executor never stalls an upgrade. With the idle guard above, a cluster without the storage prerequisites renders the executor at zero replicas and the upgrade completes.
install.sh handles both for you: a first install runs without --wait or --atomic (the post-install smoke hook is the readiness gate) and its --fleet checks only warn; upgrades run helm upgrade --atomic with the installer’s timeout and then verify workloads, skipping zero-replica ones.
Capacity Planning
Use the cloud reference executor (12 vCPU / 100 GiB, 10 slots, one per 16 vCPU / 128 GiB node) as the unit. It keeps every turn’s 8 GiB ceiling fully backed, and the roughly 28 GiB of node headroom covers the kubelet, DaemonSets, the mounter, the 4 GiB/dev/shm and the 20 GiB package cache on node disk.
How to read the table:
- Slots, not agents. A fleet with one executor serves hundreds of idle agents; what it cannot do is run more than 10 (or
turnsMax) turns at once. Extra turns queue and start as slots free. - Agents per volume is a ceiling, not a target. Most agent directories stay far below the 5 GiB quota, so a 500 GiB volume holds many more than 90 agents in practice; it also carries the shared read-only CLI, skills and runtime environment.
- Memory headroom. Keep the executor’s memory limit at or below about 80 % of the node so the node never evicts it, and keep
turnsMax x 8 GiBat or below that limit if you want every turn’s ceiling backed. - Smaller executors are fine. On the chart defaults (8 vCPU / 28 GiB limit, 12 slots) one executor fits on a 16 vCPU / 32 GiB node, at the price described in the warning above. Lower
turnsMaxto 3 or 4 there if your turns run heavy builds. - Workspace pods are separate. Agents that also use a workspace still add their
sb-*pod while active (0.25 vCPU / 256 Mi request, up to 2 vCPU / 3 GiB, 5 GiB volume); size those from the Hybrid page’s workspace accordion.
Data services and storage
The chart deploys single-replica Redis and PostgreSQL StatefulSets and a container registry alongside xpander:
Baseline storage: 48 GiB of persistent volumes, plus 5 GiB per workspace-using agent (see above). All volumes use your cluster’s default StorageClass unless overridden.
Production checklist:
- Node types: x86_64 only. General-purpose instances with a 1:4 vCPU:GiB ratio (
m5,m6i) fit the workload profile;t3burstable instances are fine for evaluation, not for production. - Raise the low defaults: API memory (table above) before real load.
- Enable autoscaling for the execution tier (
agent-worker) once concurrency grows beyond a couple of replicas. - Keep headroom: pod requests should stay under ~65% of cluster capacity so node upgrades and bursts don’t evict workloads.
- Storage class: SSD-backed volumes for Redis and PostgreSQL.
Image and Registry
Both fleet containers use this one image. On a fresh node the first executor pays the full pull, so budget the ~1.5 GB registry transfer per executor node and, on Air-Gapped, make sure the mirror holds it at the pinned tag before the first executor starts.
Network Egress
Executors hold no LLM provider key. The only LLM path is the Agent Controller’s/llm-proxy, reached over the cluster network, and the controller is the only component that egresses to a model provider.
The executor also blackholes the instance metadata endpoint (
169.254.169.254) at boot and refuses to register if it cannot, so a turn can never read node credentials.
Model providers
Executors hold no provider key; the Agent Controller does. Which key each provider needs on a self-hosted install, the Bedrock authentication methods, and which harness and model a defaulted agent gets are on AI vendors.Air-Gapped vs Hybrid
Recommended for both self-hosted deployments: keep the fleet warm (at least one executor always up). The first turn of any agent is then seconds, and idle agents cost only their directory. A cluster that cannot meet the prerequisites should set
replicas: 0 (or let the idle guard do it) and fix the missing capability before agents run on it, since every agent’s turns land on the fleet.
Next Steps
Hybrid deployment
The installer, PrivateLink, LLM keys and upgrades
Air-gapped deployment
Prerequisites, registry mirroring, and the install pipeline
EKS Cluster Setup
Node groups, the EBS CSI driver, and StorageClasses
Two-File Licensing
How seat and agent caps are enforced on a sealed install

