dusersad12/deepscaler-curated
DeepScaleR Curated (8,000 problems): RL dataset on Hugging Face. A quality-screened, 8,000-problem subset of the DeepScaleR math reasoning corpus, prepared as a drop-in training set for a Verl-based GRPO run reproducing the DeepSeek-R1 "aha moment" experiment on a small (1.5B) pretrained model.