HF RL Explorer

Ringo1110/VeriEvol-RL

VeriEvol-RL: RL dataset on Hugging Face. RL training data for VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct. This is the reinforcement-learning stage dataset used for GRPO-style training on top of the SFT-initialized policy (see the companion SFT set…

Ringo1110/VeriEvol-RL on the Hugging Face Hub