RL environments
An RL environment is the executable definition of a task for an agent: where each rollout starts, what the agent can do, how its actions change the world, when the rollout ends, and how the result is scored. For agent RL it bounds what training can teach: the policy improves only at what the environment rewards, and only on the states the environment makes reachable.
The formal model is the MDP or POMDP. An environment is the software that makes one concrete. The same environment code evaluates a model, generates training rollouts, or produces data for distillation, because it describes an interaction and leaves the choice of policy to the caller.
A coding environment
Suppose the task is to repair a parser. The agent receives the issue description and a failing test's output as its first observation, along with a working repository. The repository is external state. Reading a file reveals part of it, editing a file changes it, and running the tests starts a process and returns another observation.
A task specification of this kind needs four parts:
- Starting state: a container image with the repository at a pinned commit
- Interaction surface: a shell and a file editor
- Stopping rule: a final message, or a budget of 50 turns
- Scoring procedure: a hidden test suite run on the final repository
These parts have different lifetimes. The image is built once and reused by every rollout of the task. The shell lives for one rollout. The scorer runs once, after the agent stops, on a state the agent can no longer change.
Taskset, runtime, and harness
The taskset states what should be accomplished. It loads the tasks (prompts, files, reference answers, resource needs), defines tools specific to the task, and scores the result.
The runtime is where actions execute: a local process, a container, or a remote sandbox. It sits outside the task's definition, so the same taskset runs locally for debugging and in an isolated fleet for training.
The harness determines how the agent acts: a free-form shell loop, a single-response patch format, or a commercial coding agent. In the agent/environment split it belongs on the agent's side, together with the model (harnesses). It sets the prompt layout, the tool names and schemas the model sees, and what survives when the context fills. Environment libraries often ship a default harness with a taskset for convenience, which blurs the boundary without moving it.
Holding two of the three fixed isolates the third: a score change after swapping only the harness measures the harness.
Visible feedback and private scoring
An agent needs feedback while it works. A failing command's error message tells it what to try next, and a passing local test suggests checking a neighboring case. This feedback is part of the rollout, because it changes what the agent does.
The final score evaluates the completed state, and the agent must not be able to see or modify it while working. A coding environment shows a few tests during the interaction and keeps a larger suite for scoring. The hidden suite still encodes one meaning of success. It misses behavior it does not test and rejects valid implementations when the specification is narrower than intended, and RL pushes the policy toward whatever the suite accepts (verifiable rewards, reward hacking).
Scoring the final state, rather than the agent's last message, leaves room for many strategies. Two agents that take different paths and write different patches both receive full credit when the tests pass.
State within and between rollouts
The repository connects the turns. An edit must still be there when the agent runs the next command. Each model request sees a rendered history, but the shell, filesystem, and processes persist outside the context window.
State also needs a known beginning. Two rollouts of the same task must not inherit each other's patches or background processes: in evaluation a leaked file turns a hard task into a lookup, and in training leaked state correlates rollouts that the advantage computation treats as independent (sandboxes).
Simulated users
Conversational tool agents, such as a support agent that changes a booking, act on behalf of a user who is part of the task. In training, a language model plays that user. The simulator reads a scenario the agent never sees ("you want to move your flight to Friday, and you will accept a fee under $50"), answers the agent's questions, and ends the conversation. Its replies are observations, so the simulator becomes part of the transition function.
τ²-bench makes the environment dual-control: in its telecom domain the user also holds tools and must perform actions on a shared device state when the agent asks. Moving from a setup where the agent operates every tool itself to one where it must guide the user lowers pass^1 by 18 percentage points for GPT-4.1 and 25 for o4-mini. MUA-RL puts a GPT-4o simulated user inside the RL loop for multi-turn tool use, and its 32B model reaches 67.3 on τ²-bench Retail.
A model in the transition function brings three problems that a deterministic environment does not have:
- Reliability. The simulator makes mistakes of its own. τ²-bench's annotators found simulator errors in 16% of reviewed telecom conversations and 40–47% of retail and airline ones, with 6–13% severe enough to prevent task completion. Reward earned or lost that way is noise in the advantage.
- Seeding. A group of rollouts is compared on the premise that they share start conditions (rollout generation). The scenario and the simulator's model stay fixed across the group, and the simulator's own sampling adds reward variance the policy did not cause.
- Exploitability. The policy trains against the simulator and learns its weaknesses. When the reward reads the simulator's verdict, pressing the simulator to declare the goal met pays. Scoring the final database state instead of the simulator's closing message removes that route.
Other interaction surfaces
A browser, a desktop, or a phone screen is an environment too, with screenshots as observations and clicks and keystrokes as actions. UI-TARS-2 trains a GUI agent with multi-turn RL against a managed cluster of several thousand virtual machines.
Each of these environments needs a supply of tasks with trustworthy scorers. Tasks are mined from repositories, synthesized, or generated procedurally, and each one is checked before training (building environments at scale).