RL-Forgetting-Experiments-3/mbpp-code-rl
mbpp-code-rl: MBPP for code RL (deduplicated against MBPP+): RL dataset on Hugging Face. MBPP prepared for RLVR training in verl, with two independent hold-outs so both MBPP+ and MBPP's own canonical test split stay reportable after training on this data.
RL-Forgetting-Experiments-3/mbpp-code-rl on the Hugging Face Hub