HF RL Explorer

dusersad12/verl-deepscaler-reconciled

verl-deepscaler-reconciled: RL dataset on Hugging Face. A training-ready math reasoning dataset formatted for GRPO training with the Verl framework, rebuilt from a raw export of the DeepScaleR math data.

dusersad12/verl-deepscaler-reconciled on the Hugging Face Hub