HF RL Explorer

dusersad12/verl-deepscaler-cleaned

verl DeepScaleR - cleaned & split: RL dataset on Hugging Face. A cleaned version of the DeepScaleR-Preview-Dataset (40,315 math problem/answer pairs used for R1-style "aha moment" reproductions) converted into the parquet layout that verl expects for rule-based math RL (GRPO / PPO).

dusersad12/verl-deepscaler-cleaned on the Hugging Face Hub