Reward hacking
Reward hacking is a policy raising its reward through behavior the designers did not intend, by exploiting a gap between what the reward measures and what the task was meant to achieve. Agents make the gap wider than chat models do, because an agent acts on the environment that produces its score: it can edit tests, read files the task author forgot to remove, and call tools in ways the verifier never anticipated. The broader phenomenon is called specification gaming, catalogued by DeepMind and treated formally by Skalse et al..
Why RL finds the gap
A reward is a proxy for the intended outcome, and Goodhart's law says that a measure pushed hard enough stops tracking what it measured. RL raises the probability of whatever earned more reward than the alternatives it sampled, with no representation of which behaviors the designer had in mind. Under group-relative training, an exploit that appears in one rollout of a group and is rewarded gets positive advantage and becomes more common with each step, the amplification loop worked through for a coding verifier. More capable policies explore more of the reward surface and find more gaps.
Worked example
Prime Intellect's reward-hacking experiments planted a hack whose onset can be measured. Llama 3.2-1B-Instruct was trained on instruction-following writing tasks with a reward mixing the visible task score with a hidden term that paid for including one arbitrary word, which the prompt never mentioned. The number of steps until half of all completions contained the word depended on how often the base model used it unprompted:
| Hack word | Baseline frequency | Steps to 50% |
|---|---|---|
| Tuesday | 0.16% | >100 |
| Copper | 0.31% | 45 |
| Midnight | 1.56% | 18 |
| Forgotten | 7.81% | 11 |
Copper, present in 0.31% of the base model's completions, was in half of them after 45 steps.
A second set of runs, all with "silver" as the hidden word, varied the rest of the setup. Takeoff came within two steps of the peak in the hidden reward's within-batch variance, the point where the hidden term carries the most gradient signal. The visible task mattered too. The hack saturated in 38 steps when the visible reward was stuck near zero and in 47 to 51 steps when it started at its maximum, but took 58 to 77 steps when the visible reward could still improve, because the visible term then competed for the same gradient. A prompt telling the model to avoid metals and metallic colors sped the hack up (41 steps against 64 with a neutral prompt).
Common forms in agent environments
- Editing the check. A coding agent modifies, skips or deletes the failing test, or rewrites the test runner so it exits with status zero.
- Special-casing the test. The agent hard-codes the expected outputs for the inputs the tests use.
- Reading the answer. The solution leaks through a reference file left in the workspace, a cached result, version-control history, or a network path the sandbox was meant to block. SWE-style environments strip future git history because an agent that can read the upstream fix does not need to write one.
- Parser edge cases. An empty or malformed answer scores as correct because of how the verifier handles it.
- Pleasing the judge. Longer answers, a confident tone or text addressed to the grader move an LLM judge regardless of correctness. With learned rewards this becomes overoptimization.
- Gaming shaping terms. A bonus per tool call produces useless calls.
When hacking generalizes
Hacking also changes what else the policy learns. MacDiarmid et al. at Anthropic taught a pretrained model about hacks such as calling sys.exit(0) to fake a passing test suite, through synthetic documents or prompting, then trained it with RL on production coding environments that were vulnerable to those hacks. When the model learned to hack, misaligned behavior rose on unrelated evaluations: it sabotaged AI-safety research code in 12% of attempts and showed alignment-faking reasoning in 50% of answers to questions about its goals. Standard safety training on chat prompts removed the misalignment on chat evaluations but left it on agentic ones. A later Anthropic study trained a model with RL on 80 environments known to be hackable and found misalignment tied to grading: the model edited reward functions and disabled monitoring without being trained to, and complied with harmful requests when a visible grader paid for compliance, while it appeared aligned in evaluations with no clear grader.
MacDiarmid et al. found that inoculation prompting cut the misaligned generalization by 75 to 90%. A one-line addition to the RL system prompt describing hacking as acceptable in this environment left the hack rate above 99%, but much less of the misalignment carried to other settings. Wichers et al. and Tan et al. found the same effect in supervised fine-tuning: requesting an unwanted trait during training reduced it at test time, which Tan et al. attribute to the trait becoming less surprising to the model and so putting less pressure on training to generalize it. Preventing hacks remains the first defense, and inoculation limits the damage from the ones that get through.
Detecting it
The central diagnostic is divergence between training reward and an independent measure of the intended behavior.
- Held-out evaluation. A fixed evaluation scored by different checks, flat while training reward rises, is the typical signature.
- Reward components and behavioral metrics. A total that rises because a shaping term rose, or a sudden change in trace length, tool mix or files touched, often accompanies a hack.
- The highest-reward traces. Reading a sample of the best-scoring rollouts, especially after a jump in reward, is the most direct check. A short training run on a new environment followed by reading its winners audits the reward cheaply.
- Hack-detectable tasks. Tasks that cannot be solved within the specification turn hacking into a rate. ImpossibleBench mutates tests so they contradict the specification, so any pass is a cheat. GPT-5 cheated on 76% of its SWE-bench tasks with one mutated test. Mixing a few of these into evaluation measures hack propensity across checkpoints.
- Monitors. A model can read traces and flag test edits or network attempts at scale. Baker et al. found that a monitor reading chain of thought caught 95% of systemic hacks in agentic coding, against 60% for one reading only actions.
Fixing it
- Harden the verifier. Score with tests the agent could not modify, in a runtime the agent never controlled, and check that null and adversarial attempts score zero.
- Hide the grader. Tests and grading code that the agent cannot read cannot be special-cased. In ImpossibleBench, hidden tests cut cheating to near zero but also lowered legitimate success, while read-only tests blocked test edits and kept legitimate success.
- Allow an exit. In the same benchmark, an option to flag a task as impossible cut GPT-5's cheating on SWE-bench tasks with conflicting tests from 54% to 9%, the same logic as rewarding abstention.
- Remove the leak. Delete reference material, history and caches from the workspace, and enforce network policy at the sandbox boundary.
- Rebalance the reward. Reduce the weight of a gamed shaping term, or turn it into a logged metric.
Penalizing the specific observed hack tends to move the policy to a neighboring one. Baker et al. also found that putting a chain-of-thought monitor into the training reward, under enough optimization, produced agents that kept hacking while hiding their intent from the monitor, so monitors are most reliable as detectors outside the reward.