Introduction
About 415 wordsAbout 1 min
2026-07-03
Data scheduling (Select · Mix · Reweight) for RLVR and GRPO — a zero-fork plugin for verl.
DataFlex-RL brings a data-centric view to reinforcement learning. It plugs into verl's open registries without modifying verl source: install both packages, add a small YAML block, and run verl's normal entrypoint.
pip install verl
pip install dataflex_verlWhy data scheduling for RL?
In RL fine-tuning, not every rollout is equally useful. A prompt whose responses are all correct (or all wrong) carries no gradient signal; a low-probability token can dominate the update; one data domain may already be mastered while another lags. DataFlex-RL lets a training run use the signals it already produces to decide which data to emphasize, drop, or sample more of — without a second reward model or extra forward passes.
Those signals include reward, advantage, per-token log-probabilities, and group structure (uid), giving a policy several natural ways to define “useful data.”
Three mechanisms
| Mechanism | What it does | Mount point |
|---|---|---|
| Reweight | per-token / per-sample loss weights | trainer, after advantage |
| Select | drop samples (zero their gradient) | trainer, after advantage |
| Mix | dynamic domain sampling proportions | custom replay buffer, pre-rollout |
All three are expressed through one design: a shared Scorer (signal → score) feeding a mechanism-specific Actuator (score → action). See Framework Design for the rationale.
Relationship to DataFlex
DataFlex is the original data-centric training system for SFT, built on LLaMA-Factory. DataFlex-RL is its RL counterpart: the same Select / Mix / Reweight vocabulary and the same Scorer/Actuator decomposition, re-grounded on verl's RL training loop and signals. The two share design DNA but are independent packages (different host frameworks).
Included examples
The repository includes short end-to-end GRPO smoke tests for Qwen2.5-0.5B and GSM8K, covering all three mechanisms. They print DataFlex metrics such as retention, per-token weight statistics, and domain proportions. The examples are intended to verify the integration; scale them up for a full training campaign.
The separate evaluation companion contains the current controlled RLVR comparison study and its released analysis records. This site documents the runtime package and its extension points.
Next
- Framework Design — the signal/action architecture and verl mount points
- Installation — install, sanity check, and compatibility notes
- Then dive into a mechanism: Reweighter · Selector · Mixer