dusersad12/verl-deepscaler-reconciled
verl-deepscaler-reconciled: RL dataset on Hugging Face. A training-ready math reasoning dataset formatted for GRPO training with the Verl framework, rebuilt from a raw export of the DeepScaleR math data.
dusersad12/verl-deepscaler-reconciled on the Hugging Face Hub