Verifiable rewards (RLVR)
Reinforcement learning with verifiable rewards (RLVR) is RL in which the reward comes from a program that checks the outcome against the task: a test suite that executes a patch, a checker that compares a mathematical answer with the reference, a function that inspects the final state of a database. For agents it is the default reward source, because an agent may search, recover from errors and revise its work in many different ways while the verifier needs to decide only whether the resulting state satisfies the task.
The name comes from the Tülu 3 post-training recipe. DeepSeek-R1 made the approach the standard way to train reasoning: its R1-Zero experiment used only rule-based accuracy and format rewards.
Tests, answer checks and state checks
Tests. A coding verifier runs code. HumanEval checks a generated function against unit tests. SWE-bench places an agent in a repository at the commit before an issue was fixed, then runs fail-to-pass tests (which the fix should make pass) and pass-to-pass tests (which it should not break).
Answer checks. A math verifier parses the final answer and compares it with the reference symbolically or numerically. It should accept 1/2, 0.5 and \frac{1}{2} when the task intends them as equal, and reject text that merely contains a plausible number somewhere.
State checks. A tool-use verifier inspects the world the tools changed: a task pairs a database, tools that manipulate it and an instruction with a function that checks the final database. An agent can reach the required state by a different sequence of calls than the reference solution.
A model's verdict is a prediction sensitive to phrasing and order, so "verifiable" is kept for checks that execute and model scores belong to judges. The two can be combined. R2E-Gym used verifiers at test time to pick one of 26 candidate patches per SWE-bench Verified task. Generated tests and a model reading the agent's trace each picked a correct patch for 42.8% of tasks, and a hybrid of the two reached 51.0%, because few tests distinguished the candidates and the model attended to the agent's reasoning more than to the patch.
Worked example
After the agent stops, a coding verifier receives the final repository. It can apply several checks and report each one, so a rising total can be traced to the check that moved:
# generic pseudocode
async def score(repo, task):
target = await repo.run(["pytest", *task.fail_to_pass])
regressions = await repo.run(["pytest", *task.pass_to_pass])
changed = await repo.changed_files()
return {
"fixed": float(target.exit_code == 0),
"no_regressions": float(regressions.exit_code == 0),
"in_scope": float(changed <= task.allowed_files),
}This verifier has a gap. If the tests named in fail_to_pass live in the repository the agent worked in, the agent can edit or delete them and fixed still reports 1. Suppose a group of eight rollouts for one task contains two correct fixes, one rollout that deleted the failing test, and five failures:
With the group mean as baseline, as in GRPO, all three rewarded rollouts get advantage and each failure gets . The update cannot tell the deletion from the fix, so it raises the probability of both. If deletion is easier to find than a fix, it gains probability faster, appears in more groups, and is reinforced more often in each later step. One exploit in one group of eight is enough to start that loop, and once every rollout in a group deletes the test, the group is uniform and the exploit is locked in without any further gradient against it.
This is why a false positive, full reward for incorrect work, costs more than a false negative, zero reward for correct work. A false negative wastes a rollout and discourages one valid strategy. A false positive is a direction the optimizer pursues as hard as the intended one. The fix for this verifier is to score with tests the agent never saw: copy them in after the attempt, or grade in a fresh sandbox that receives only the agent's patch.
Testing a verifier
A verifier is code and can be wrong like any other code. Every task should pass its own gold solution and fail a null attempt such as the initial state or an empty answer, a check that scales to whole tasksets as part of environment scaling. The false positives these checks miss show up in the highest-scoring rollouts of a capable model, which is where to read first.
What a verifier should check
A verifier should check the invariants that define success and nothing else. For code, that means executing behavior and checking constraints instead of comparing source text. For structured output, it means parsing and validating fields instead of matching serialization. For a search task, it can mean checking the final answer and its cited evidence instead of prescribing a query sequence. Each incidental requirement is a way for correct work to fail.
The checks are still a sample. A test suite covers finitely many behaviors, a database check can certify the final state while missing a forbidden intermediate action (deleting and recreating a record the task said to update), and a symbolic checker can accept an expression outside the intended domain. Held-out evaluation and trace inspection catch what the verifier does not.