HF RL Explorer

OctoReasoner/mercury_verl

Mercury (verl efficiency eval set): RL dataset on Hugging Face. The eval split of Elfsong/Mercury (arXiv 2402.07844; 256 LeetCode-style tasks; the train split ships no test cases and is not gradable), converted to the verl rule-reward schema by verl/scripts/data/mercury.py. Source license…

OctoReasoner/mercury_verl on the Hugging Face Hub