HF RL Explorer

Add a picklable, length-shaped cosine-scaled correctness reward for RLVR trainers, reusing math-verification…

Add a picklable, length-shaped cosine-scaled correctness reward for RLVR trainers, reusing math-verification…: a task in HF ML Bench v0 (Harbor dataset). Extend the built-in rule-based reward library with a length-shaped correctness reward for use with reinforcement-learning-from-verifier-rewards…

The task

Extend the built-in rule-based reward library with a **length-shaped correctness reward** for use with reinforcement-learning-from-verifier-rewards trainers (e.g. GRPO / RLOO). The reward combines math-verification correctness with a cosine schedule over completion length, following the recipe in Appendix C.1 of…

Part of AdithyaSK/HF_ML_Bench_v0.