Skip to content

E2B Sandbox

Run your agent’s commands and file edits in an isolated E2B cloud sandbox instead of on your machine.

Source

While Pydantic AI Harness is on 0.x releases, the API may change between minor releases; when it does, deprecation warnings and release-note migration guidance tell you (or your agent) exactly how to upgrade. See the version policy.

Install

Terminal
pip install "pydantic-ai-harness[e2b,anthropic]"

The anthropic extra is there because the examples use an Anthropic model; swap it for your model provider’s extra. Then set E2B_API_KEY to your E2B API key, and ANTHROPIC_API_KEY for the examples.

Quick start

from pydantic_ai import Agent
from pydantic_ai_harness.coder import Coder
from pydantic_ai_harness.e2b_sandbox import E2BSandbox

agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(), Coder()])
result = agent.run_sync('Clone https://github.com/pydantic/pydantic-ai and summarize how capabilities work.')

Coder’s shell and file tools now run in the sandbox, not on your machine. With Coder, RepoContext creates the sandbox when the run starts, even without a tool call; use Coder(repo_context=False) for lazy creation. It keeps running, and billing, after the run ends; see Clean up.

A sandbox lives for 1 hour by default. When that runs out it pauses, and the next run resumes it. On E2B’s Pro plan you can pass up to E2BSandbox(sandbox_timeout=86400). A plan-limit hint is added only when E2B rejects the requested lifetime; other create errors, including network timeouts, retain the upstream error.

Commands start in the home directory, /home/user, which holds dotfiles and the pip cache. Pass a project directory, such as E2BSandbox(working_dir='/home/user/project'), so the project is not the home directory; it is created for you on a new sandbox.

E2B’s default template does not include ripgrep (rg). Coder can search without it, but installing rg makes searches faster. Build a reusable template once, outside an agent run (building can take a minute):

from e2b import AsyncTemplate, Template
await AsyncTemplate.build(Template().from_base_image().apt_install(['ripgrep']), 'my-rg-template')
agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(template='my-rg-template'), Coder()])

The build snippet is illustrative and not part of the runnable agent examples below.

The default template has Python and git but not your project’s dependencies, not even pytest. Install them in a custom template, or with await workspace.run(['pip', 'install', ...]) before the run, as in Prepare a sandbox before the run.

Default templateValue
Useruser
HOME and default working directory/home/user
rgNot installed
pytestNot installed

Continue in the same sandbox

from pydantic_ai import Agent
from pydantic_ai_harness.coder import Coder
from pydantic_ai_harness.e2b_sandbox import E2BSandbox

agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(), Coder()])
result = agent.run_sync('Clone https://github.com/pydantic/pydantic-ai and summarize how capabilities work.')

followup = agent.run_sync(
    'Which capability would you add next, and where would it live?',
    message_history=result.all_messages(),
)

The follow-up run finds the sandbox in the message history and works in it, so the clone is still there. Without the history, a run starts a new sandbox. A paused sandbox resumes. If the sandbox has been killed, the run raises WorkspaceUnavailableError instead of starting over in an empty one.

Choose the tools

For a narrower agent, use Shell and FileSystem instead of Coder, or write your own tool:

from pydantic_ai import Agent, RunContext
from pydantic_ai_harness.e2b_sandbox import E2BSandbox
from pydantic_ai_harness.filesystem import FileSystem
from pydantic_ai_harness.shell import Shell

agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(), Shell(), FileSystem()])


@agent.tool
async def run_python(ctx: RunContext, code: str) -> str:
    """Run a Python snippet in the sandbox."""
    result = await ctx.workspace.run(['python', '-c', code], timeout=10)
    return result.stdout + result.stderr

Shell commands run as the user account, and files the backend creates are owned by it. E2B’s file service still reads and writes with its own root privileges, so Unix permissions do not restrict file operations; do not rely on them to keep file tools out of a path.

E2B file reads reject FIFOs rather than waiting for a writer. An ordinary read performs a shell FIFO check before downloading the file. E2B’s file API omits looping symlinks from directory listings.

A background child that inherits stdout or stderr can keep run() waiting for the SDK stream to close after the main command exits. Redirect background output to a file when starting a long-running job; Shell.start_command manages its own output log.

What a timeout stops

On a command deadline or cancellation, the backend sends TERM then KILL to the command’s foreground process group, including children that have not detached into a separate session. Custom templates and attached sandboxes are probed once for setsid (util-linux); without it, only the command leader is stopped, so child processes may continue. Custom templates also need /bin/bash and /bin/sh. Intentionally detached background jobs are not stopped. It does not destroy the shared sandbox. Stop is best effort: if E2B cannot be reached, or a start won registration but does not respond to the stop request, work may still be running. A cancelled start that arrives after stop is fenced before user code runs. Retain the sandbox ref to inspect or kill it explicitly.

A command whose combined output passes 10 MiB is stopped the same way and raises WorkspaceOutputLimitError, with the first 65,536 characters of each stream in stdout and stderr. Redirect large output to a file instead.

A command killed by a signal may return exit_code=-1 in E2B; the SDK does not identify the signal, so this is not converted to 128+signal.

Concurrent writes to the same path are not atomic on E2B: uploads may interleave, and a reader can see a partial file. Coordinate writers or write to separate paths when using FileSystem or Coder.

See Workspaces for more.

Reattach later

To come back to the sandbox without the message history, store the run’s workspace ref and pass it back as workspace=:

from pydantic_ai import Agent
from pydantic_ai_harness.coder import Coder
from pydantic_ai_harness.e2b_sandbox import E2BSandbox

agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(), Coder()])

result = agent.run_sync('Clone https://github.com/pydantic/pydantic-ai and summarize how capabilities work.')
ref = result.workspace.ref  # store this, e.g. in your database

later = agent.run_sync('Which capability would you add next, and where would it live?', workspace=ref)

The ref holds no credentials, so the process that reattaches needs E2B_API_KEY too. Pass workspace='new' to start a fresh sandbox even when the message history names one.

template and allow_internet_access only shape a new sandbox. sandbox_timeout also applies when you reattach, and working_dir and env apply to every command.

Already have an e2b.AsyncSandbox? Pass workspace=E2BSandboxBackend(sandbox=sandbox) to a run, with E2BSandboxBackend from pydantic_ai_harness.e2b_sandbox. E2BSandbox’s settings don’t apply to it; pass working_dir= and env= to the backend.

Prepare a sandbox before the run

To seed files or install packages before the agent starts, create the backend yourself, work in it through Workspace, and pass it to the run:

import asyncio

from pydantic_ai import Agent
from pydantic_ai.workspaces import Workspace
from pydantic_ai_harness.coder import Coder
from pydantic_ai_harness.e2b_sandbox import E2BSandbox, E2BSandboxBackend

agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(), Coder()])


async def main() -> None:
    backend = E2BSandboxBackend(working_dir='/home/user/project')
    try:
        workspace = Workspace(backend)
        await workspace.write_text('calc.py', 'def add(a, b):\n    return a - b\n')
        install = await workspace.run(['pip', 'install', 'pytest'], timeout=300)
        if install.exit_code != 0:
            raise RuntimeError(f'pip install failed: {install.stderr}')
        result = await agent.run('Fix the bug in calc.py.', workspace=backend)
        print(result.output)
    finally:
        ref = backend.ref  # set once the sandbox exists
        if ref is not None:
            await E2BSandbox().destroy(ref)


if __name__ == '__main__':
    asyncio.run(main())

The finally kills the sandbox even when setup or the run fails. To keep it instead, store backend.ref and pass it as workspace= to reattach later.

Preview a dev server

With Shell, ask the agent to use start_command for npm run dev -- --host 0.0.0.0 --port 3000, then poll check_command and curl http://localhost:3000/health until ready. Save the returned command ID. Given the workspace ref, connect with sandbox = await e2b.AsyncSandbox.connect(ref.id) and use sandbox.get_host(3000) for the public hostname (prefix with https:// for the preview URL). When done, call stop_command with the ID while the workspace is attached, then delete the sandbox with await E2BSandbox().destroy(ref). Do not leave a public preview running longer than necessary.

Clean up

The sandbox keeps running, and billing, after the run ends. Pydantic AI never kills it. If acquisition is cancelled while creation is in flight, the backend finishes recording the ref when E2B responds; a lost response may still leave a sandbox without a ref. Kill it with the ref you stored:

from pydantic_ai.workspaces import WorkspaceRef
from pydantic_ai_harness.e2b_sandbox import E2BSandbox


async def kill_sandbox(ref: WorkspaceRef) -> None:
    await E2BSandbox().destroy(ref)

E2BSandbox.backend(ref) constructs an attached backend without I/O; destroy(ref) kills by ID without attaching. This kills a paused sandbox too, without resuming it. A sandbox you don’t kill is paused when its sandbox_timeout runs out. See E2B’s sandbox lifecycle.

A failed run returns no result, so there is no ref to store. To terminate its sandbox, clean up in an on_run_error hook; after_run doesn’t run when a run fails:

from typing import Any

from pydantic_ai import Agent, RunContext
from pydantic_ai.capabilities import Hooks
from pydantic_ai.run import AgentRunResult
from pydantic_ai_harness.coder import Coder
from pydantic_ai_harness.e2b_sandbox import E2BSandbox

hooks = Hooks()


@hooks.on.run_error
async def terminate_failed_run(ctx: RunContext[None], *, error: BaseException) -> AgentRunResult[Any]:
    if ctx.workspace.ref is not None:
        await E2BSandbox().destroy(ctx.workspace.ref)
    raise error


agent = Agent('anthropic:claude-opus-5-5', capabilities=[E2BSandbox(), Coder(), hooks])

Configuration

OptionWhat it does
templateE2B template name or ID for a new sandbox. Default: None, E2B’s base. An unknown template fails on first use.
allow_internet_accessWhether a new sandbox can reach the internet. Default: True.
sandbox_timeoutSeconds the sandbox lives before E2B pauses it, set on create and on reattach. Default: 3_600 (1 hour, the most E2B’s Hobby plan allows).
working_dirAbsolute directory commands start in and relative paths resolve against. Default: None, the sandbox’s own (/home/user on the default template, where commands run as user); prefer relative paths or set working_dir for portable code. Created on a new sandbox; on an attached or caller-supplied sandbox it must already exist.
envEnvironment variables every command gets. Default: None. Output is decoded as UTF-8 either way; the default template’s locale is POSIX, so pass env={'LC_ALL': 'C.UTF-8'} if a tool such as wc -m should count characters rather than bytes. Nothing from your machine’s environment reaches the sandbox.

Commands run on the asyncio event loop only: the E2B SDK reads their output with asyncio tasks, so under Trio run raises UserError.

Durable execution

Install the temporal extra too:

Terminal
pip install "pydantic-ai-harness[e2b,anthropic,temporal]"

Run a Temporal dev server on localhost:7233 first. The agent and workflow must be defined at module level for activity registration.

import asyncio
import uuid

from pydantic_ai import Agent
from pydantic_ai.durable_exec.temporal import PydanticAIPlugin, PydanticAIWorkflow, TemporalDurability
from pydantic_ai_harness.coder import Coder
from pydantic_ai_harness.e2b_sandbox import E2BSandbox
from temporalio import workflow
from temporalio.client import Client
from temporalio.worker import Worker

agent = Agent(
    'anthropic:claude-opus-5-5',
    name='e2b_coder',
    capabilities=[E2BSandbox(), Coder(), TemporalDurability()],
)


@workflow.defn
class SandboxWorkflow(PydanticAIWorkflow):
    __pydantic_ai_agents__ = [agent]

    @workflow.run
    async def run(self, prompt: str) -> str:
        return (await agent.run(prompt)).output


async def main() -> None:
    client = await Client.connect('localhost:7233', plugins=[PydanticAIPlugin()])
    async with Worker(client, task_queue='sandbox', workflows=[SandboxWorkflow]):
        print(
            await client.execute_workflow(
                SandboxWorkflow.run, 'Use the shell tool to run pwd.',
                id=f'sandbox-{uuid.uuid4()}', task_queue='sandbox',
            )
        )


if __name__ == '__main__':
    asyncio.run(main())

Removing a capability while workflows using it are still running changes their replay history. Drain those workflows or use Temporal worker versioning before deploying the change.

Telemetry

E2BSandbox emits no spans of its own. Core’s instrumentation records the sandbox on the agent run span as pydantic_ai.workspace.provider and pydantic_ai.workspace.id, and each command or file operation a tool makes runs inside that tool call’s span; operations at run start, such as RepoContext loading repo instructions, run in the agent run span. Creating a sandbox also logs Created E2B sandbox <id> at INFO on the pydantic_ai_harness.e2b_sandbox._backend logger, so the id is on record even if the run ends before it is stored.

API reference

E2BSandbox

Bases: AbstractCapability[AgentDepsT]

Supply an isolated E2B sandbox as the run’s workspace.

A run with no reference creates a fresh sandbox. Pass a WorkspaceRef supplied by the application to attach to an environment managed elsewhere.

This capability supplies execution only. Compose it with tools or capabilities that use the workspace, such as Coder, Shell, or FileSystem. Shell commands run under sh -c in the sandbox’s login shell environment.

Attributes

template

E2B template name or ID for a newly created sandbox; E2B’s default when None.

An unknown template raises WorkspaceUnavailableError on first use.

Type: str | None Default: None

allow_internet_access

Whether a newly created workspace may reach the internet.

Type: bool Default: True

sandbox_timeout

Total lifetime of the sandbox in seconds, applied on create and on attach.

When it runs out, E2B pauses the sandbox and attaching resumes it. The default, 1 hour, is the most E2B’s Hobby plan allows; Pro plans allow up to 86400.

Type: int Default: DEFAULT_SANDBOX_TIMEOUT

working_dir

Absolute directory commands start in and relative paths resolve against; None uses the sandbox’s own.

Type: str | None Default: None

env

Environment variables every command gets, also on an attached workspace; nothing is read from the host.

Type: Mapping[str, str] | None Default: field(default=None, repr=False)

Methods

backend
def backend(ref: WorkspaceRef) -> E2BSandboxBackend

Attach lazily to an existing E2B sandbox.

Returns

E2BSandboxBackend

destroy

@async

def destroy(ref: WorkspaceRef) -> None

Kill a sandbox by ID, including paused sandboxes, without attaching.

Returns

None

get_workspace
def get_workspace(
    ctx: RunContext[AgentDepsT],
    *,
    ref: WorkspaceRef | None,
) -> WorkspaceBackend | None

Build the backend for this run. No I/O here: it attaches or creates on first use.

Returns

WorkspaceBackend | None

E2BSandboxBackend

Bases: WorkspaceBackend, SupportsCommands, SupportsFilesystem

An E2B sandbox as a Pydantic AI WorkspaceBackend.

Commands and file operations run inside an E2B microVM, so the host is never exposed.

Building one does no I/O. The first operation creates or attaches to a workspace, and the typed e2b.AsyncSandbox is available through get_sandbox(). The backend does not kill the sandbox; killing it is the application’s job.

Commands run as one-shot operations, with complete output returned after they finish.

Every command runs through /bin/bash -l -c, so an argv sequence is quoted into a single shell word string first and login startup files run before the command does; a shell=True string runs under /bin/sh -c inside that login shell. E2B’s own command timeout abandons the output stream and leaves the command running, so the deadline is enforced client-side instead. On timeout or cancellation (including while starting), a separate command signals the process group when setsid is available. Custom templates without setsid fall back to stopping the command leader only.

Constructor Parameters

sandbox : e2b.AsyncSandbox | None Default: None

A live e2b.AsyncSandbox you already have. Whoever created it owns killing it.

ref : WorkspaceRef | None Default: None

Identity of an existing sandbox to attach to on first use.

template : str | None Default: None

E2B template name or ID a newly created sandbox runs; E2B’s default when None. An unknown template raises WorkspaceUnavailableError on first use. Custom templates need /bin/bash and /bin/sh; without setsid, stop targets only the leader.

allow_internet_access : bool Default: True

Whether a newly created sandbox may reach the internet.

sandbox_timeout : int Default: DEFAULT_SANDBOX_TIMEOUT

Total lifetime of the sandbox in seconds, applied when it is created and again when attaching to it. When it runs out, E2B pauses the sandbox rather than killing it, and attaching resumes it. The default, 3600, is the most E2B’s Hobby plan allows; Pro plans allow up to 86400.

working_dir : str | None Default: None

Absolute directory commands start in and relative paths resolve against. E2B has no create-time working directory, so this is applied per command, including on an attached sandbox; None uses the sandbox’s own default, discovered with pwd -P on first use.

env : Mapping[str, str] | None Default: None

Environment variables every command gets, on a created or an attached sandbox; per-command env is layered on top. Nothing is read from the host environment.

Attributes

ref

Identity of the sandbox, or None before one has been created.

Type: WorkspaceRef | None

Methods

get_sandbox

@async

def get_sandbox() -> e2b.AsyncSandbox

Return the typed e2b.AsyncSandbox, for E2B features the workspace API does not cover.

On a backend with no sandbox yet, this creates one (which E2B bills) or attaches to the one ref names, just like the first operation; attaching resumes a paused sandbox. Attaching by ref to a sandbox that no longer exists raises WorkspaceUnavailableError; it does not create a replacement. After that it returns the cached handle without checking that the sandbox is still running: one killed elsewhere surfaces on the next operation. Calling this does not make you responsible for killing the sandbox; whoever holds the ref decides, as before.

Returns

e2b.AsyncSandbox

working_dir

@async

def working_dir() -> str

The canonical absolute directory commands start in.

Returns

str

run

@async

def run(
    command: WorkspaceCommand,
    *,
    shell: bool = False,
    env: Mapping[str, str] | None = None,
    timeout: float | None = None,
) -> CommandResult

Run a command, killing it on timeout, cancellation, or a failed result read.

Returns

CommandResult