Skip to main content
LangSmith Sandboxes run code in isolated microVMs on a dedicated pool of Kubernetes nodes. This guide explains where the components run, how requests reach a sandbox, and which state survives a host restart. For supported platforms, prerequisites, and installation instructions, see Enable Sandboxes. The architecture below describes the sandbox-host runtime. Storage mounting and configuration options can differ across Helm releases.

Scaling and capacity

Size the host pool, configure autoscaling, and plan memory and cache capacity.

Upgrades and operations

Understand host draining, failure recovery, and production monitoring.

Understand the deployment model

A sandbox is a Firecracker microVM, not a Kubernetes pod. Each sandbox-host pod manages multiple microVMs, each with its own guest kernel. Kubernetes schedules the hosts; LangSmith places sandboxes onto them. The host Deployment uses required pod anti-affinity to place at most one host pod on each node. Its node selector and tolerations target a dedicated, KVM-capable node pool. Adding a sandbox does not create another pod.

The platform backend manages sandbox placement and traffic. Hosts run microVMs; shared storage keeps persistent state independent of a node.

Understand component failure impact

Each component has a different failure scope. Distinguish temporary unavailability from loss of its stored data when planning recovery.
For fencing, recovery behavior, and storage protection, see Understand failure impact. One elected host observes pool health and runs the host autoscaler. Other hosts continue serving sandboxes without holding leadership.

Inspect one sandbox node

A host pod runs the sandbox daemon and Firecracker processes. In deployments with a host-owned JuiceFS mount, it also runs the JuiceFS client. The daemon connects to the agent inside each guest through vsock.

One Kubernetes pod contains multiple microVMs. The host needs KVM access; guest memory and node-local caches consume the node's resources.

The host uses host networking on port 19190. The platform backend connects to the host’s node address; clients use the LangSmith API or service URLs, not the host listener directly. Keep host listeners on private networks. Host pods require privileged access for virtualization, networking, and mounts. Use dedicated nodes and restrict administrative access to them. Guest isolation does not make the host pod an unprivileged workload.

Follow a request

Creating a sandbox passes through the platform backend before reaching a host. Subsequent command, file, and service requests follow the same authorization boundary. Placement considers committed CPU and prefers the sandbox’s previous host when its load is close to the least-loaded candidate. Returning to that host can reuse cached data. It is a preference, not a guarantee that a sandbox stays on one node.
The utilization target controls scaling, not admission. A busy pool can still place work on an existing host. With no ready hosts, creation fails with NoSandboxHostsAvailable rather than waiting in a capacity queue. See Plan for bursts.

Track persistent state

JuiceFS separates a filesystem’s metadata from its contents:
  • Redis metadata: Directory entries, file sizes, and the mapping from files to data chunks.
  • Object storage: Data chunks for disk images, saved memory images, and snapshots.
  • Local cache: Copies of data blocks on individual nodes. A node replacement loses that cache, not the shared filesystem.
JuiceFS Redis is durable storage, not the LangSmith cache Redis. Object storage alone cannot reconstruct the filesystem without its metadata. Protect both stores, use noeviction for metadata Redis, and test recovery. See Protect sandbox storage.

Distinguish suspend from a crash

A persistent sandbox can resume on a different compatible host because its disk and successfully saved memory live on shared storage.
  • Graceful stop: The host saves VM memory and stops the VM. A later start or wake request can restore running processes from that image.
  • Host loss: Running memory is lost. Recovery waits for the fencing window before another host can start the sandbox from persisted disk.
  • Deletion: Removes the sandbox rather than moving it to another host. A separately captured snapshot has its own lifecycle.
Restoring memory requires a complete, compatible memory image. It is not live migration and does not preserve an open client connection. Unflushed writes and in-memory work can be lost after an abrupt failure. These persistence guarantees do not apply to ephemeral sandbox storage.

See also