Understand a rolling upgrade
Thesandbox-host image packages the host daemon and runtime artifacts, including Firecracker, guest artifacts, and the JuiceFS client. Set images.sandboxHostImage.tag to the runtime image for your LangSmith release. Follow release-specific instructions when upgrading older storage layouts.
The host Deployment uses maxSurge: 0 and maxUnavailable: 1. Kubernetes terminates a host before creating its replacement, then waits for the replacement to become ready. Required one-host-per-node placement and the shared host port prevent two host pods from occupying the same node.
A healthy four-host pool has three ready hosts during each replacement. The timeline is illustrative; drain and startup durations depend on the workload and storage.
maxUnavailable is not a guarantee against unrelated outages.
For host-owned JuiceFS mounts, startup includes mounting the filesystem and checking storage access before opening the host listener. A replacement that cannot access storage blocks progress instead of accepting sandbox traffic.
Follow a host drain
A graceful host shutdown attempts to suspend persistent sandboxes before the pod exits:1
Stop new placement
The host enters a draining state and notifies the platform backend. New sandbox placements use other ready hosts.
2
Save running sandboxes
The host stops routing to each sandbox as it suspends that VM. It attempts to save the VM’s memory to JuiceFS and reports the sandbox stopped.
3
Release the host
The host keeps renewing its Lease while stopping workloads. Once the VMs have stopped, it releases the Lease and tears down its storage mount.
4
Resume on demand
A later start or wake request places the sandbox on an available, compatible host. A complete saved memory image allows processes to resume; otherwise recovery uses the persisted filesystem.
Budget enough time to drain
sandboxes.sandboxHost.deployment.terminationGracePeriodSeconds defaults to 300. Kubernetes can kill the host when that grace period expires, even if the host has not finished saving every sandbox’s memory.
Measure a drain with representative sandbox density and memory use. Check the shutdown summary for the drain duration and how many sandboxes suspended successfully or failed to suspend. Increase the grace period when the host cannot save every sandbox’s memory in time, and keep node-maintenance drain timeouts longer than that period.
An incomplete drain can lose running memory and require crash recovery for affected sandboxes. Increasing the timeout does not solve storage unavailability or insufficient capacity.
Coordinate node maintenance
Node replacement and Kubernetes upgrades also terminate host pods. Configure your node maintenance process to respect graceful shutdown and avoid draining several sandbox hosts simultaneously. Enablesandboxes.sandboxHost.pdb.enabled to use the chart’s PodDisruptionBudget for voluntary evictions. Its default budget permits one unavailable host. A PDB does not prevent node failure, direct pod deletion, or every provider-specific maintenance action.
Keep node autoscaling consistent with the host autoscaler. Where your node autoscaler supports eviction protection, prevent it from independently evicting active host pods during routine consolidation. Let host scale-down drain workloads before removing empty nodes. Check your autoscaler’s behavior rather than assuming one annotation works for every provider.
Do not accelerate a rollout by deleting multiple host pods or increasing disruption beyond the capacity you have tested.
Diagnose a stalled rollout
Compare desired, ready, and updated host counts. Then inspect pod events and node provisioning:- Hosts stay
Pending: Check node pool limits, infrastructure quota, regional capacity, node selectors, taints, and resource requests. - Hosts start but do not become ready: Check KVM availability, runtime logs, and access to JuiceFS metadata and object storage.
- Hosts remain terminating: Inspect drain progress and storage latency before changing the grace period.
sandbox-host.smith.langchain.com/autoscale-override Deployment annotation can temporarily pin the host count when autoscaling is enabled. Use it only after evaluating remaining capacity, and remove the override afterward. A lower target does not fix the underlying capacity shortage.
Plan machine changes
Memory snapshots include CPU and virtual machine state. Do not assume that a saved VM can resume across different CPU models or incompatible host kernels. Firecracker’s snapshot compatibility guidance requires matching hardware and software configurations, with limited, explicitly caveated exceptions. Keep a consistent CPU model in the sandbox pool. Treat changes to the CPU vendor or model, host kernel, or sandbox runtime as transitions that require restore testing. A newer CPU or kernel does not automatically make a destination compatible. Compatibility in one direction does not establish compatibility in reverse.Example transitions, not a support matrix
These are documentation-based examples, not a tested LangSmith compatibility matrix or an expansion of supported platforms. Every lower-risk candidate still requires validation on your exact deployment. Keep the exposed CPU model and features, host kernel, and sandbox runtime consistent. Confirm that the destination meets the Sandbox platform and KVM requirements.
A matching SKU or family is not sufficient evidence. Provider placement and available processors can vary. For example, Google’s CPU platform documentation describes machine types with multiple CPU platforms and how platform selection works.
Validate a transition before rollout
To validate the source-to-target transition:- Confirm that the target exposes
/dev/kvmand meets the supported platform requirements. - Record the exposed CPU model and features, host kernel, and
sandbox-hostimage version on both source and target nodes. - Use isolated test capacity matching the target configuration. Resume a persistent sandbox suspended on a source node, and separately create a sandbox from a memory snapshot captured there. Confirm both run on a target node, not the original host.
- Verify that in-memory state survives, then exercise representative workload paths. A successful VM boot alone does not establish compatibility.
- Test the reverse direction if the rollout requires rollback. Do not assume that a snapshot captured on the target can resume on an older source configuration.
Understand failure impact
Protect sandbox storage
The JuiceFS metadata store and object store form one persistent filesystem. Protect them separately from the platform’s cache and queue Redis.- Dedicated metadata Redis: Use
noeviction, high availability, persistence, and backups. Do not flush or recreate it as part of cache maintenance. - Memory headroom: Alert before Redis reaches its memory limit. With
noeviction, writes fail when memory is exhausted rather than evicting filesystem metadata. - Object retention: Do not apply lifecycle rules that delete live JuiceFS objects or move them to an inaccessible storage tier.
- Recovery testing: Test restoring filesystem metadata and object data together. A replica improves availability but does not replace backups.
- Private access: Keep storage and host endpoints private. Use the installation guide’s identity and secret configuration for your environment.
Monitor the pool
The host exposes Prometheus metrics on its internal/metrics endpoint. Scrape it through your private monitoring infrastructure, not public ingress. See Export LangSmith telemetry for platform monitoring setup.
The elected host exports these pool gauges:
Also monitor host pod readiness,
Pending pods, node memory and disk pressure, storage latency, Redis memory and evictions, sandbox startup failures, and drain duration. A persistent gap between desired and ready hosts calls for checking node provisioning and host startup, not only workload demand.
Check production readiness
Before running production workloads:- Confirm the dedicated nodes expose KVM and match the host selector and tolerations.
- Reserve memory for guests, the host runtime, JuiceFS, and Kubernetes overhead.
- Align host limits, node limits, infrastructure quotas, and workspace quotas.
- Test scaling from the minimum pool size with a cold cache.
- Test one-host loss and a one-host-at-a-time upgrade under load.
- Verify application behavior when execution streams or service connections disconnect.
- Measure drain duration and set compatible pod and node shutdown timeouts.
- Confirm metadata Redis uses
noevictionand has tested backups and restore procedures. - Monitor object-store access, metadata capacity, and the gap between desired and ready hosts.
See also
- Sandbox architecture and lifecycle
- Scale self-hosted Sandboxes
- Enable Sandboxes
- Upgrade self-hosted LangSmith
- Plan platform disaster recovery
Connect these docs to your agent of choice via MCP for real-time answers.

