Advantage Reweighting
About 320 wordsAbout 1 min
2026-07-03
Source: Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs (arXiv:2505.12929).
Motivation
In policy-gradient RL, a token's gradient magnitude scales inversely with its probability. Low-probability tokens therefore produce outsized gradients and can dominate the update, destabilizing training. Advantage Reweighting (AR) damps them with a per-token weight that grows with the token's probability.
Rule
For each token t with policy probability π_θ(t):
wt=α⋅πθ(t)+(1−α)
- Low-prob tokens (
π→0) get weight near1−α(damped). - High-prob tokens (
π→1) get weight near1. - Weights are mean-normalized over valid tokens to preserve the effective LR.
α ∈ [0, 1] controls damping strength; α=0 recovers the unweighted baseline.
In DataFlex-RL
AR is token-granularity: it uses the token_prob scorer (per-token π_θ = exp(old_log_prob), forward-pass-free) and the advantage_reweight reweighter, which returns a (bs, L) weight matrix directly.
trainer: {v1: {trainer_mode: dataflex_sync}}
dataflex:
mechanism: reweight
scorer: {name: token_prob}
actuator: {name: advantage_reweight, params: {alpha: 0.5}}
warmup_step: 0CLI:
+dataflex.mechanism=reweight \
+dataflex.scorer.name=token_prob \
+dataflex.actuator.name=advantage_reweight \
+dataflex.actuator.params.alpha=0.5Why it's a strong pick for small models
Among the reweighting methods surveyed, AR has direct small-model evidence in its originating paper: that paper reports gains at 3B and 7B (for example, a +46% relative change on Knights-and-Knaves and positive math results). These are prior-work numbers, not a claim that every DataFlex-RL run will show the same improvement. The implementation needs only old_log_probs—no extra forward pass or reward model—and can be useful when studying later training stages.
Signal / rule / granularity
| Property | Value |
|---|---|
| Signal | per-token π_θ (from old_log_probs) |
| Rule | w = α·π + (1−α), mean-normalized |
| Granularity | token |
| Needs groups | no (any advantage estimator) |
| Extra forward pass | no |