HF RL Explorer

OctoReasoner/FinalMix3

FinalMix3: RL dataset on Hugging Face. A multi-task code reinforcement-learning mixture in the verl RL prompt format: one code-generation split plus a suite of auxiliary code-understanding tasks, so the same corpus drives three training regimes. It is OctoReasoner/FinalMix2 with three changes…

OctoReasoner/FinalMix3 on the Hugging Face Hub