> ## Documentation Index
> Fetch the complete documentation index at: https://docs.langchain.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Scale self-hosted Sandboxes

> Plan self-hosted Sandbox capacity with host autoscaling, Kubernetes node scaling, memory headroom, and shared filesystem caching.

export const SandboxDiagram = ({name}) => {
  const diagrams = {
    architecture: {
      viewBox: "0 0 900 730",
      sources: {
        light: "/images/langsmith/sandboxes/self-host-architecture.svg",
        dark: "/images/langsmith/sandboxes/self-host-architecture-dark.svg"
      },
      title: "Self-hosted Sandbox architecture",
      description: "Apps connect to the platform backend, which manages sandbox hosts on dedicated Kubernetes nodes. PostgreSQL stores placement, while JuiceFS uses dedicated Redis metadata and object storage."
    },
    node: {
      viewBox: "0 0 820 630",
      sources: {
        light: "/images/langsmith/sandboxes/self-host-node.svg",
        dark: "/images/langsmith/sandboxes/self-host-node-dark.svg"
      },
      title: "Inside one sandbox node",
      description: "One privileged sandbox-host pod manages multiple Firecracker microVMs with separate kernels and guest agents. KVM and a disposable cache are local; JuiceFS connects to shared storage outside the node."
    },
    autoscaling: {
      viewBox: "0 0 820 470",
      sources: {
        light: "/images/langsmith/sandboxes/self-host-autoscaling.svg",
        dark: "/images/langsmith/sandboxes/self-host-autoscaling-dark.svg"
      },
      title: "Two layers of Sandbox autoscaling",
      description: "Hosts report committed CPU through Leases. An elected host scales the Deployment, and the node autoscaler provides eligible nodes. New hosts check storage before accepting placements."
    },
    "scaling-behavior": {
      viewBox: "0 0 820 400",
      sources: {
        light: "/images/langsmith/sandboxes/self-host-scaling-behavior.svg",
        dark: "/images/langsmith/sandboxes/self-host-scaling-behavior-dark.svg"
      },
      title: "Sandbox scaling behavior",
      description: "A qualitative chart shows ready hosts lagging behind rising demand. After demand falls, the replica target holds its peak for a stabilization window, then decreases gradually."
    },
    rollout: {
      viewBox: "0 0 820 430",
      sources: {
        light: "/images/langsmith/sandboxes/self-host-rollout.svg",
        dark: "/images/langsmith/sandboxes/self-host-rollout-dark.svg"
      },
      title: "One-host-at-a-time rolling upgrade",
      description: "Four hosts drain and start a new version sequentially. An otherwise healthy pool has three hosts ready during each replacement, and all four ready when the rollout finishes."
    }
  };
  if (!Object.hasOwn(diagrams, name)) {
    throw new Error("Unknown Sandbox diagram");
  }
  const diagram = diagrams[name];
  return <div className="sandbox-diagram-container">
      {["light", "dark"].map(theme => {
    const source = diagram.sources[theme];
    return <a key={theme} className="sandbox-diagram-link" data-diagram-theme={theme} href={source} target="_blank" rel="noopener noreferrer" aria-label={`${diagram.title}: open full-size diagram`}>
            <svg className="sandbox-diagram" viewBox={diagram.viewBox} role="img" aria-label={diagram.description} focusable="false">
              <use href={`${source}#diagram`} />
            </svg>
          </a>;
  })}
    </div>;
};

Self-hosted Sandboxes scale in two layers: LangSmith adjusts the host Deployment, and your node autoscaler supplies Kubernetes nodes. Plan both layers together so new hosts are ready before the existing pool runs out of capacity.

This guide builds on [Sandbox architecture](/langsmith/self-host-sandbox-architecture). For cloud-specific infrastructure setup, see [Enable Sandboxes](/langsmith/deploy-self-hosted-full-platform#enable-sandboxes).

## Understand the two scaling layers

The built-in host autoscaler runs on an elected `sandbox-host`. Hosts publish their committed vCPU and sandbox counts through Kubernetes Leases. The leader uses that information to adjust the Deployment's replica count.

<Frame caption="LangSmith requests host pods. Your infrastructure supplies nodes. A desired replica is not usable capacity until its host is ready.">
  <SandboxDiagram name="autoscaling" />
</Frame>

The host autoscaler does not use an HPA, KEDA, or a Prometheus adapter. Do not configure another controller to manage the same Deployment's replica count.

A new host pod needs an eligible node without another host pod. When none is available, it stays `Pending` until your node autoscaler adds capacity. Configure that autoscaler for the dedicated KVM-capable node pool, including compatible labels, taints, and instance requirements.

During normal scale-down, the host autoscaler retains the recent peak target for a stabilization window. It then removes at most one host per window. Pod deletion costs favor less-loaded hosts. Removing a running host [drains its sandboxes](/langsmith/self-host-sandbox-operations#follow-a-host-drain); it is not a disruption-free operation.

## Calculate a host target

Committed vCPU is the CPU requested by running sandboxes on ready hosts, not measured CPU usage. For a pool of equally sized hosts, the load-based target is:

```text theme={"theme":{"light":"catppuccin-latte","dark":"catppuccin-mocha"}}
desired hosts = ceil(committed vCPU / (host vCPU × target fraction)) + headroom hosts
```

The autoscaler clamps the result to `minReplicas` and `maxReplicas`. Without a host capacity signal, it starts from the minimum. Use a homogeneous pool: the calculation assumes a common per-host CPU capacity.

For example, with 16-vCPU hosts, a 70% target, and one headroom host:

* **Per-host target**: `16 × 0.70 = 11.2` committed vCPU.
* **Workload**: 60 sandboxes requesting 0.5 vCPU each commit 30 vCPU.
* **Pool target**: `ceil(30 / 11.2) + 1 = 4` hosts, subject to the configured bounds.

The 70% target is not a per-host capacity limit. Placement can exceed it while the pool scales or when it reaches its maximum.

### Plan for bursts

The replica target can rise before nodes and hosts become ready. Headroom covers some of that delay; it does not make node provisioning instantaneous.

<Frame caption="Illustrative scaling behavior, not a latency benchmark. Scale-up follows demand; scale-down waits through the stabilization window and proceeds gradually.">
  <SandboxDiagram name="scaling-behavior" />
</Frame>

Choose `minReplicas` and `headroomHosts` for the burst you need to absorb while a node starts. Keep at least two hosts for a pool that must continue accepting work during a one-host rollout. Measure node startup, host readiness, and sandbox creation latency under representative load.

If the pool reaches its maximum, investigate demand and infrastructure limits rather than assuming the utilization target prevents overload. With no ready hosts, requests to create a sandbox fail. Clients need bounded retries and backoff for transient capacity failures.

## Configure the host autoscaler

The Helm values below live under `sandboxes.sandboxHost.autoscaling`. Check the values shipped with your chart version; existing overrides can change the effective settings.

| Value | Chart default | Operational guidance |
| - | - | - |
| `enabled` | `false` | Enable host autoscaling when your node pool can scale with it. |
| `minReplicas` | `1` | Use at least `2` to avoid zero ready hosts during a normal rollout. |
| `maxReplicas` | `10` | Align with eligible node capacity and infrastructure quotas. |
| `targetUtilizationPercent` | `70` | Lower values leave more CPU headroom. This measures commitment, not live usage. |
| `headroomHosts` | `1` | Adds spare capacity above the load-based target. |
| `scaleDownStabilizationSeconds` | `300` | Increase to reduce repeated drains during short demand dips. |

When autoscaling is enabled, the chart leaves the Deployment's replica count to the host autoscaler. Otherwise, `sandboxes.sandboxHost.deployment.replicas` sets a fixed host count.

Deployment annotations prefixed with `sandbox-host.smith.langchain.com/autoscale-` can override the baseline settings without restarting hosts. Inspect existing overrides when Helm values and observed scaling disagree. Prefer durable Helm configuration for normal operation.

Align three separate limits:

1. Set the host replica maximum for the peak workload.
2. Allow the node pool and infrastructure quotas to supply that many eligible nodes.
3. Set workspace sandbox quotas for the expected concurrent CPU, memory, and sandbox count.

Workspace quotas control what users can request. They do not provision nodes or guarantee physical capacity.

## Size CPU and memory together

The following example applies the same 70% target and one headroom host to 500 running sandboxes. It shows the CPU-based target before `maxReplicas` clamps it, not a performance guarantee.

| vCPU per sandbox | Total committed vCPU | 16-vCPU hosts | 32-vCPU hosts | 64-vCPU hosts |
| - | - | - | - | - |
| 0.5 | 250 | 24 | 13 | 7 |
| 1 | 500 | 46 | 24 | 13 |
| 2 | 1,000 | 91 | 46 | 24 |

Check memory separately. For example, 500 sandboxes configured with 2 GiB each request 1,000 GiB of guest memory, before host and filesystem overhead.

<Warning>
  Host autoscaling uses committed CPU, not memory pressure. A CPU-sized pool can still run out of memory. Budget for guest memory, the host daemon, JuiceFS, caches, and Kubernetes overhead.
</Warning>

The host pod's memory request must reflect the workloads it contains. A small request for the daemon alone does not reserve memory for its microVMs. Review `sandboxes.sandboxHost.deployment.resources` against node allocatable memory and measured peak use.

## Balance node size and cache capacity

Larger nodes let more sandboxes share a local JuiceFS cache and amortize host overhead. However, each host failure or drain affects more sandboxes and can take longer to save their memory.

Choose nodes based on these constraints:

* **Virtualization**: Expose `/dev/kvm` on a supported Linux node configuration.
* **CPU compatibility**: Keep a consistent CPU model across the pool. See [Plan machine changes](/langsmith/self-host-sandbox-operations#plan-machine-changes).
* **Memory**: Leave capacity for host and filesystem overhead, not only the requested guest memory.
* **Local storage**: Provide cache space for the working set and monitor disk pressure.
* **Failure impact**: Keep enough hosts and spare capacity to tolerate losing or draining one.

For host-owned JuiceFS mounts, `sandboxes.juicefs.hostMount.cacheDirs` selects node-local cache paths. `sandboxes.juicefs.hostMount.mountOptions` controls cache options. The chart's default cache budget is 50 GiB across the configured directories; installing a larger disk does not automatically raise that budget.

A cache can survive a pod replacement on the same node, but not replacement of that node. Treat cache warm-up as part of capacity testing, not as durable storage recovery.

## See also

* [Sandbox architecture and lifecycle](/langsmith/self-host-sandbox-architecture)
* [Upgrade and operate Sandboxes](/langsmith/self-host-sandbox-operations)
* [Configure self-hosted resource usage](/langsmith/self-host-usage)

***

<div className="source-links">
  <Callout icon="terminal-2">
    [Connect these docs](/use-these-docs) to your agent of choice via MCP for real-time answers.
  </Callout>

  <Callout icon="edit">
    [Edit this page on GitHub](https://github.com/langchain-ai/docs/edit/main/src/langsmith/self-host-sandbox-scaling.mdx) or [file an issue](https://github.com/langchain-ai/docs/issues/new/choose).
  </Callout>
</div>
