jasonkena/guru-RL-66k-privileged
guru-RL-66k-privileged: Guru-RL — Privileged (math + code): RL dataset on Hugging Face. A privileged variant of the math and code training splits of LLM360/guru-RL-92k, built for privileged-information RL: every row carries a worked reference solution z that the policy never sees, available to a…