HF RL Explorer

amphora/MathConstructOptimize-Envs

MathConstructOptimize-Envs: Harbor dataset on Hugging Face. 3,577 RL tasks in 142 families where the model must construct a mathematical object, graded by a deterministic checker. There is no answer matching and no LLM judge in the reward path. Most math RL data asks for a final number. Here the…

amphora/MathConstructOptimize-Envs on the Hugging Face Hub