Skip to main content
Sandboxes provide isolated environments for running code and exposing services. Install them when you need sandbox workloads or a feature that requires them, such as Engine.
Sandboxes require an Enterprise plan.
Sandboxes are disabled by default. After installation, see LangSmith Sandboxes for user workflows in the LangSmith UI and APIs. For the infrastructure model and production planning, see Sandbox architecture, scaling and capacity, and upgrades and operations.

Supported platforms

Self-hosted Sandboxes are supported on:
  • Amazon Elastic Kubernetes Service (EKS)
  • Google Kubernetes Engine (GKE)
  • Azure Kubernetes Service (AKS)
Self-hosted Sandboxes on Azure require LangSmith Helm chart v17 (0.17.x).

Components

Enabling Sandboxes provisions the following resources:
  • Sandbox runtime pods that run sandbox workloads on KVM-capable nodes.
  • A JuiceFS metadata store backed by Redis and object storage backed by S3, GCS, or Azure Blob Storage.
  • Optional wildcard ingress for services exposed from inside Sandboxes.

Prerequisites

1

Install the base LangSmith platform

Install LangSmith on Kubernetes before enabling Sandboxes. See Self-host LangSmith on Kubernetes.Sandboxes run in the same Kubernetes cluster and namespace as the LangSmith release.If your cluster cannot pull from public registries, also mirror the sandbox runtime image. See Additional images for Sandboxes.
2

Add KVM-capable nodes

Your cluster must include dedicated nodes that can run nested workloads with Linux KVM available at /dev/kvm.These can be bare-metal machines or supported cloud instances with nested virtualization enabled. On AWS and GCP, use x86_64 Linux instances that expose /dev/kvm to the sandbox runtime.
On EKS, the VPC CNI addon must be v1.21 or later. v1.20.0 crashes on 8th-generation Intel instances (for example m8i): aws-node enters CrashLoopBackOff, the node reports cni plugin not initialized, and the managed node group eventually fails with NodeCreationFailure: Unhealthy nodes in the kubernetes cluster.
The default Helm scheduling values expect these nodes to have the following label and taint:
If your nodes use different labels or taints, override sandboxes.sandboxHost.deployment.nodeSelector and sandboxes.sandboxHost.deployment.tolerations.
3

Configure JuiceFS storage

Sandboxes require JuiceFS-backed shared storage. You must provide:
  • A Redis-compatible metadata store.
  • An object storage bucket or bucket root.
  • A JuiceFS configuration Secret, or enough Helm values for the chart to create one.
Use these object storage backends for sandboxes.juicefs.storage and sandboxes.juicefs.bucket:Do not use object-store subpaths in sandboxes.juicefs.name. Use a flat name, such as sandbox-juicefs. JuiceFS stores objects under that name inside the configured bucket.
For the Redis metadata store, we recommend setting maxmemory-policy to noeviction. This avoids evicting JuiceFS metadata under memory pressure. Monitor Redis capacity and scale it before it reaches memory limits.With noeviction, Redis writes can fail when the instance reaches max memory, so keep enough memory headroom for sandbox metadata growth.
4

Configure sandbox secrets

Sandboxes need additional secret material for service-to-service authentication and callback signing.
The callback signing value must be an Ed25519 private JWK. Keep it stable across upgrades.
5

Choose a proxy CA mode

The sandbox egress authentication proxy uses this CA for TLS interception and credential injection.The chart supports two proxy CA modes:In GitOps workflows that render manifests without live cluster access, prefer existingSecret. The generatedSecret mode uses Helm’s live lookup behavior to reuse the generated Secret on upgrades; pure render workflows cannot read the live Secret and may produce new cert material on each render.

Enable with Helm

Enable Sandboxes by setting Helm values directly, as shown in this section. On AWS and GCP, you can instead enable Sandboxes with Terraform, which provisions the infrastructure and generates the Helm values. Add the following values to your langsmith_config.yaml, along with the sandbox secret values described in the Prerequisites. Replace placeholders with your deployment-specific values.
Apply the updated chart:

Enable with Terraform

The LangSmith Terraform modules can provision the required AWS and GCP infrastructure and generate the corresponding Helm values.
In modules/aws/infra/terraform.tfvars, enable Sandboxes and configure the sandbox node capacity:
AWS Sandboxes require redis_source = "external". The Terraform module:
  • Creates a dedicated ElastiCache Redis instance for JuiceFS sandbox metadata.
  • Configures that dedicated instance with the recommended noeviction policy.
  • Reuses the LangSmith S3 bucket for sandbox object storage.
  • Creates the JuiceFS configuration Secret.
  • Adds the expected node label and taint.
The AWS setup script generates the sandbox service-auth secret, callback signing JWK, and dedicated JuiceFS Redis auth token through the normal SSM-backed setup flow. Run the infra setup script before applying Terraform if those values do not exist yet.If you deploy the Helm release with the Terraform app module, set the sandbox app values in modules/aws/app/terraform.tfvars as well:
When enable_sandboxes = true, the Terraform app module requires an explicit LangSmith Helm chart v17 release and a sandbox runtime image tag.Run the normal AWS flow:

Optional: enable service URLs

Set sandboxes.serviceUrlBaseUrl when users need browser or programmatic access to HTTP services running inside Sandboxes.
This requires wildcard DNS and TLS for *.sandbox-services.example.com. When ingress.enabled is true, the chart also adds a wildcard ingress rule that routes these service URLs to the LangSmith platform backend. Service URLs also require an Ed25519 private JWKS with a nonempty kid on its first key. Set config.signingJwks, or add langsmith_signing_jwks to your existing LangSmith app Secret. This is separate from sandboxes.callbackSigningJwk. For key generation instructions, see Configure a signing JWKS. Service URLs are required to build and edit custom apps with chat, even though they are optional for other sandbox workflows.

Verify the installation

After the upgrade completes, verify that the sandbox runtime pods and JuiceFS volumes are ready:
Then run a sandbox smoke test:
  1. Create a sandbox from a public image, such as a Python image.
  2. Start a Python HTTP server inside the sandbox.
  3. Snapshot the sandbox with memory enabled.
  4. Create a new sandbox from the snapshot.
  5. Verify that the HTTP server is still running in the restored sandbox.

Upgrade notes

Sandbox runtime image changes roll out through the sandbox-host Kubernetes Deployment. The chart uses a no-surge rolling update strategy by default, so hosts are replaced one at a time. During a normal Helm upgrade, a terminating host stops accepting new Sandboxes, attempts to save each running Sandbox’s VM memory to JuiceFS, and then stops those VMs before the pod exits. This shutdown is bounded by the sandbox-host pod termination grace period, which defaults to 300 seconds. This is not live migration: Sandboxes on that host are interrupted during the restart. Sandboxes are not proactively restarted. They start again when a user or API action starts the Sandbox, or when a request path wakes it. LangSmith then places the Sandbox on an available host and restores from the saved memory image if the shutdown capture completed. If the memory image is absent or incomplete, the Sandbox starts from the saved root filesystem.