HF RL Explorer

dusersad12/verl-deepscaler-curated

verl-deepscaler-curated: RL dataset on Hugging Face. A cleaned, de-duplicated and evaluation-safe training split derived from the DeepScaleR-Preview-Dataset, reformatted for rule-based-reward RL post-training with verl. Total examples: 38,783 (from 41,705 raw records read across three source…

dusersad12/verl-deepscaler-curated on the Hugging Face Hub