Turns and tool calls
A turn is one model generation: the model receives a token context and samples until it returns a response, emits a tool call, or hits another stopping condition. A tool call is a structured action encoded in those tokens, such as a shell command or a search query, that the harness parses and executes outside the model. The tool's result is an observation. It is not part of the turn, and it enters the context of the next turn.
The definition fixes where the policy stops and the environment starts. A person might call a sequence of searching, reading and answering one turn of a conversation, but at the inference boundary it contains several model generations, each one a turn, with the environment acting between them.
A tool call ends the turn
From tokens to an action
Suppose a turn ends with this text:
{"name": "wiki_search_pages", "arguments": {"query": "policy gradient variance"}}The model generated a token sequence. The harness parses it against the declared tool schema and invokes the search, so the environment sees a search action. The trainer sees the sampled tokens that encoded it, and computes their probability under the policy.
The two views do not always line up. Several serializations can represent the same arguments. A parser may reject malformed JSON, normalize whitespace or repair a field, and a tool API may fill in defaults that never appeared in the tokens. The environment transition follows the parsed call, while the probability used for training follows the tokens the model sampled.
Parsing strictness has training consequences. A strict parser turns near-valid output into an error observation, and a model trained against it learns the exact format. A permissive parser lets more outputs succeed, and the model may never learn to produce well-formed calls, which then fail under a stricter deployment parser.
Tool results as observations
A tool result may be a few bytes or several megabytes. It may contain a fact, a transient error or text written by another system. The harness chooses how to show it to the model: in full, truncated, summarized, or stored elsewhere with a pointer. The trace should record both what the tool produced and what the model was shown.
A minimal agent loop, written as generic pseudocode:
messages = [system, task_prompt]
for _ in range(max_turns):
reply = model.sample(messages) # one turn
messages.append(reply)
call = parse_tool_call(reply)
if call is None: # plain response: the agent is done
break
try: # environment step, outside the turn
result = tools[call.name](**call.arguments)
except ToolError as err:
result = f"error: {err}" # recoverable: show the model
messages.append({"role": "tool", "content": truncate(result)})Each pass through the loop is one turn followed by one environment step. Production harnesses add parallel tool calls, retries, compaction and budgets around the same structure.
Tools and inference timing
Generation and execution alternate. A file read takes milliseconds, a test suite takes seconds, and a build or browser task can take minutes. While a tool runs, the model has nothing to sample, because the next observation does not exist yet.
The inference server therefore sees short bursts of generation separated by environment waits. Keeping a request's KV cache resident through every wait holds scarce accelerator memory, and evicting it forces the next turn to recompute the whole prefix (KV cache).
Tools also hold state with different lifetimes. A search index over a fixed corpus is expensive to build and safe to share across every rollout. A repository or database that the agent can modify must be isolated per rollout, or one rollout's edits leak into another (sandboxes).
Training tool use
Tool output sits in the same token sequence as model output, but the model did not generate it. Search-R1 (Jin et al.) masks retrieved tokens out of the loss for stable training, so the gradient covers only the model's queries and reasoning. With that mask and an outcome-only reward, it improved over retrieval-augmented baselines by 41% on Qwen2.5-7B and 20% on Qwen2.5-3B. The retrieved text still conditions every later token (loss masks).
A tool call has value only through its effect on the task, so the reward usually scores the finished task and charges per call only when execution has a price or a rate limit. ToolRL (Qian et al.) compared reward types, scales, granularity and schedules for tool selection and argument filling, Its best design adds a format check to a correctness score that matches each call's tool name, parameter names and parameter values against a reference, and it gained 17% over base models and 15% over SFT models.
Long tasks need long rollouts. ASearcher (Gao et al.) argues that turn limits of ten or fewer, common in earlier search-agent RL, cap the strategies a policy can discover. With fully asynchronous training, its agents exceeded 100 tool calls and 400k output tokens in a single rollout during training.
Training can also drive tool use toward zero. If early rollouts that skip tools score slightly better, the policy stops calling them and never discovers when they help (entropy collapse). Two rollouts with the same reward can follow different tool strategies, and only their traces show which.