Sandboxes

A sandbox is an isolated execution environment where an agent's actions run: a container or lightweight virtual machine with its own filesystem, processes, and network policy, created for one rollout and discarded after it. In agent RL it is the part of the environment that model-generated commands touch, and its guarantees set which scores are trustworthy and how many rollouts fit in a training step.

For the agent the sandbox is the world. For the operator it is also a security boundary and a unit of scheduling, and training stresses all three roles at once, because it runs thousands of untrusted programs in parallel and reinforces whatever raises the reward.

Sandbox lifetime

A coding rollout contains dozens of turns. The repository keeps edits from one command to the next, background processes keep running, and temporary files become inputs to later actions. The sandbox therefore lives for one rollout:

prepare image or snapshot
        ↓
launch sandbox
        ↓
agent turns and tool execution
        ↓
score final state
        ↓
discard

Initialization must be reproducible. A task names an image, a repository revision, setup commands, and resource limits. When setup installs an unpinned package or calls a mutable network service, two rollouts of the same task start from different states, and the group comparison that GRPO relies on mixes two tasks.

Reset must be complete. Reusing a warm sandbox saves startup time only when the state is fully restored. Leftover files, caches, ports, or processes disclose information from another rollout and make its score incomparable.

The scoring boundary

Optimization reinforces any action that raises the measured outcome. When the agent is able to read hidden tests, edit the verification script, or write to the score file, those actions are part of the effective task. No adversarial instruction is needed. Repeated selection makes an accidental shortcut more common each time it pays (reward hacking).

The boundary around scoring has to be enforced by the system, not requested in the prompt:

  • Private tests are mounted only after the agent stops, or scoring runs in a separate sandbox that receives only the agent's output.
  • The verifier sees a read-only copy of the candidate.
  • Network access is denied by default, or allowed only to named hosts.
  • Credentials are scoped to one service and one rollout.

Every channel counts, including the ones the harness itself needs. A sandbox with no general network access still reaches the model, and any service on that path that fetches content on request is a route to the outside.

Resource limits

Time, memory, disk, network, and process limits determine which solutions exist. A program that passes with unlimited memory fails under the intended constraint. An agent with open internet access is solving a different task from one that must reason from the repository.

Each limit also needs a recorded meaning. A timeout can be a task failure, a hung tool, or an overloaded host, and only the first says anything about the policy, so the trace records which limit fired and why (rollouts and traces).

Sandboxes and throughput

Agent rollouts alternate inference with execution. A short command returns in milliseconds, while a build holds a sandbox for minutes while the model waits. A training system needs enough concurrent sandboxes to keep generation busy without launching more state than it manages.

Each rollout pays for image transfer and startup, repository checkout and dependency installation, execution within turns, and teardown and scoring. A batch of 256 rollouts with a two-minute setup each spends 8.5 sandbox-hours on setup per training step. Prebuilt images and cached snapshots move that work out of the critical path, and warm pools cut startup latency. Both are safe only while the initial state stays identical.

Remote sandboxes add a network boundary. The harness runs near the inference server while commands execute on a sandbox service, so the trace depends on delivery and retries. Retrying a read is harmless. Retrying a state-changing command applies it twice.

Choosing an isolation level

The same task can run in a local subprocess, a container, a user-space kernel, or a microVM. The agent sees the same filesystem and shell in each, and the runtimes differ in what the agent's code shares with the host:

RuntimeShared with the hostSuited to
SubprocessKernel, filesystem, networkDebugging a taskset on trusted inputs
ContainerKernel (isolated by namespaces and cgroups)Tasks from trusted sources on a dedicated host
User-space kernel (gVisor)Host kernel, reached only through a narrow, filtered set of syscallsUntrusted code that avoids syscalls gVisor does not implement
MicroVMHardware, through a minimal virtual machine monitor; each sandbox boots its own guest kernelUntrusted code, and tasks that run Docker or system services inside

Stronger isolation costs startup time and density per host. A container escape through a kernel bug exposes every other rollout on the machine, and a policy under optimization tries actions its authors did not anticipate. Training at scale therefore sits at the bottom two rows, and a subprocess runtime is kept for tasks with no code execution, such as single-turn text tasks.