LoRA and parameter-efficient RL
LoRA (low-rank adaptation) fine-tunes a model by freezing its weights and training a small low-rank correction added to selected weight matrices. In agent RL it leaves rollouts, rewards and the loss unchanged and shrinks what the trainer stores, optimizes and ships to inference after every step.
How a LoRA adapter works
LoRA (Hu et al.) replaces a frozen weight matrix in the forward pass with
Only and are trained. The rank is small, and is a scaling constant, so sets the effective step size of the adapter. starts at zero, which means the adapted model begins identical to the base model.
Worked example
The wiki-search example trains Qwen3-4B-Instruct-2507 as a multi-turn search agent with rank-8 adapters on every attention and MLP projection. The model has 36 layers with hidden size 2560, and a rank- adapter on a matrix has parameters.
| projection | full matrix | rank-8 adapter |
|---|---|---|
q_proj, 2560 → 4096 | 10,485,760 | 53,248 |
k_proj, 2560 → 1024 | 2,621,440 | 28,672 |
gate_proj, 2560 → 9728 | 24,903,680 | 98,304 |
The seven projections in a layer (v_proj matches k_proj, o_proj matches q_proj, and up_proj and down_proj match gate_proj) give 458,752 adapter parameters per layer, and the 36 layers give about 16.5 million, 0.4% of the model's 4 billion. The optimizer state shrinks by the same factor, and each policy update ships about 1/240 of the parameters a full-weight update would. The frozen base still has to sit in memory for the forward and backward passes, which QLoRA reduces further by quantizing it.
Why low rank suffices for RL
The amount a model has to learn depends on how much information the training signal carries. A supervised example supplies a target for every token. A policy-gradient rollout supplies one scalar reward, turned into an advantage and spread across all of its tokens. Thinking Machines' LoRA Without Regret argues that policy gradients absorb on the order of one bit per rollout, and reports that LoRA matches full fine-tuning for RL on math reasoning even at rank 1.
Other results point the same way. Tina (Wang et al.) trained LoRA adapters with RL on a 1.5B reasoning model and reached 43.3% pass@1 on AIME24 for $9 of post-training and evaluation compute. Mukherjee et al. found that full-parameter RL changes only 5–30% of parameters, but that the resulting weight updates are nearly full-rank. The case for LoRA in RL therefore rests on how little information the reward carries, not on RL's full-parameter update being low-rank.
The same Thinking Machines study gives three practical findings:
- Apply LoRA to all layers, especially the MLP and MoE layers. Attention-only LoRA underperformed MLP-only LoRA at matched parameter counts.
- Raise the learning rate. The optimal LoRA learning rate was about ten times the full fine-tuning rate, in both SFT and RL.
- Capacity limits SFT first. Large supervised datasets carrying many bits per example can exceed what a low-rank adapter stores. RL rarely approaches that limit.
What LoRA changes in the training system
In an asynchronous RL system the trainer ships new weights to inference after every step (weight broadcast). With LoRA it ships only the adapter, and checkpoints shrink the same way.
The frozen base is also a free reference model. Turning the adapter off recovers the starting policy, so a KL penalty against the initial model needs no second copy of the weights.
Two costs come with it. Kernel fusions and optimized layouts built for the base weights may not apply to adapted layers, and the inference engine computes the adapter path separately from the base matmul. Any numerical difference between how the trainer and the inference engine apply the adapter adds to trainer–inference mismatch, so the mismatch metrics need rechecking after LoRA is enabled.
Serving many adapters on one base model
Every adapter shares the same base weights, so one inference server can host many adapters and choose one per request. S-LoRA and Punica batch requests for different adapters together: the base matmul runs once over the whole batch, and a segmented kernel applies each request's own low-rank update. Adapters move in and out of GPU memory as demand changes.
For RL this lets several runs, tasks or checkpoints share one inference fleet, each addressing its own adapter by name. The output of a training run becomes a file of tens of megabytes that deploys next to others on a shared base.