Skip to main content
A sandbox gives a managed deep agent an isolated filesystem and shell for working with files, running code, and executing commands.
Managed Deep Agents is in public beta and available on LangSmith Cloud in the US region only.
Put the sandbox declaration under sandbox/. Add sandbox/setup.sh only if you want to provision a snapshot:
For the full project layout, see Project structure.

Configure a sandbox

Use a sandbox when the agent needs to write files, run code, or execute shell commands. mda init scaffolds a sandbox declaration. Managed Deep Agents enables the sandbox only while the sandbox/ directory is present. mda init does not create setup.sh. Add that file yourself if the snapshot should install packages, clone a tree, or otherwise change the image. Managed Deep Agents uses LangSmith Sandboxes for this backend. Reuse is always one sandbox per durable thread. Declare the sandbox with define_sandbox:
sandbox/__init__.py

Configure the sandbox proxy

The sandbox proxy injects headers into matching outbound requests and controls which destinations the sandbox can reach. The proxy runs outside the sandbox, so sandbox code can call authenticated APIs without handling the credentials. For example, to call the OpenAI API from the sandbox, store OPENAI_API_KEY in your LangSmith workspace secrets and configure this proxy rule:
sandbox/__init__.py
For configuration options and network restrictions, see Sandbox auth proxy.

Use connections in proxy headers

Use Connections in sandboxes to authenticate CLI commands and API requests. For example, to call the GitHub API from the sandbox as the current user, create the github connection first, then configure the proxy:
sandbox/__init__.py
Use a connection reference as the header value, or format it with bearer(ref) or basic(username, ref). Omit the header’s type for connection values. Managed Deep Agents sets it to opaque.

Provision a snapshot

If sandbox/setup.sh exists, mda deploy and mda dev run the script once and save the resulting environment as a snapshot. Modifications from that run, such as cloned repositories and installed packages, persist in the snapshot. New threads clone that snapshot instead of running setup.sh. The snapshot is reused until setup.sh changes, at which point it is rebuilt. The script runs with bash -e. A non-zero exit fails the snapshot and the deploy or mda dev session. LangSmith does not update the live deployment to the failed snapshot. Any previously successful snapshot continues to serve.
sandbox/setup.sh
Project .env values that deploy forwards are available as environment variables when setup.sh runs, for example a token used to clone a private repo. Thread sandboxes that clone the snapshot do not inherit those variables. Do not write secrets onto the filesystem while setup.sh runs; anything on disk is part of every thread’s image. Editing setup.sh and redeploying does not wipe /workspace on live threads. Those boxes keep the files they already have. A new thread clones the new snapshot.

Choose a bake base

With no bake base, LangSmith’s default sandbox template is the starting point. To start from something else, set exactly one of these:
sandbox/__init__.py
For a private image, pass the image and a registry. Managed Deep Agents creates or updates a deployment-owned Host registry at bake time. Only the variable name is compiled; the credential value does not enter the build or the snapshot. Name the password in password_env:
sandbox/__init__.py
Put GHCR_TOKEN in the project .env or the process environment. After bake, Managed Deep Agents does not forward that value to the running Agent Server.

How the agent uses the sandbox

The agent uses built-in filesystem tools such as ls, read_file, write_file, edit_file, delete, glob, and grep, and runs shell commands with execute. Use instructions to specify where the agent should work and what it must not modify.

Read and write sandbox files from code

Authored tools and middleware reach the sandbox filesystem through runtime.backend. Use it when your own code needs a file, rather than prompting the agent to fetch one for you.
runtime.backend requires managed-deepagents>=0.8.0.
Annotate the runtime parameter to receive the typed surface:
tools/report.py
Each operation binds to the sandbox of the thread handling the current run, so two threads reading /workspace/report.txt see their own copy. The backend resolves lazily, and a tool that never touches it never provisions a sandbox.

Available operations

Every method has an async counterpart prefixed with a, such as aread, awrite, and adownload_files. Arguments and return types come from the Deep Agents backend contract. See Backends.

Transfer binary files

upload_files and download_files move raw bytes, so they suit images, archives, and any other file that text operations would corrupt. Download returns the bytes for each requested path:
tools/checksum.py
Upload takes path and content pairs, one per file:
Each result carries path and error, and a download also carries content. On failure, error is one of file_not_found, permission_denied, is_directory, or invalid_path, and the downloaded content is empty. Check error rather than assuming the transfer succeeded.

Limits

runtime.backend covers the sandbox only. It has no route to Context Hub, so skills, instructions, and memory are not reachable through it. Without a sandbox, runtime.backend is None. Guard on it before every call, because a project can remove sandbox/ after the tool ships. delete, upload_files, and download_files depend on the installed backend and raise when it does not implement them. A root glob returns an error instead of provisioning a sandbox.

Disable the sandbox

Delete the sandbox/ directory to opt out, such as for an agent that only needs its prompt, memory, and tools. For existing deployments, deleting the deployment with mda delete also deletes the managed sandboxes associated with it, the {deployment}--setup-* recipe snapshots, and the deployment-owned registry when one exists.

Deployment

Managed Deep Agents owns sandbox naming, recipe bake, reuse, recovery, and cleanup. Each durable thread gets its own sandbox, cloned from the current recipe snapshot. For platform-level lifecycle details, see Sandboxes.

When to use a sandbox

For more information, see Project structure.