dusersad12/verl-deepscaler-curated
verl-deepscaler-curated: RL dataset on Hugging Face. A cleaned, de-duplicated and evaluation-safe training split derived from the DeepScaleR-Preview-Dataset, reformatted for rule-based-reward RL post-training with verl. Total examples: 38,783 (from 41,705 raw records read across three source…