HF RL Explorer

caiyuchen/OPV-Math

OPV Math: original experiment data: RL dataset on Hugging Face. Prepared for Learning to Steer, Steering to See. train is filtered DeepMath (distillation); teacher train is the original DAPO teacher corpus. validation is AIME2024. These corpora are not interchangeable.

caiyuchen/OPV-Math on the Hugging Face Hub