dusersad12/DeepScaleR-Verl-Clean
DeepScaleR-Verl-Clean: RL dataset on Hugging Face. A single cleaned, deduplicated, verl-ready parquet built from four partial dumps of the DeepScaleR math dataset (agentica-org/DeepScaleR-Preview-Dataset, MIT licensed). It is meant for rule-based-reward GRPO/RL training with the verl framework.