HF RL Explorer

nvidia/Nemotron-RL-Super-Training-Blends

Nemotron-RL-Super-Training-Blends: NeMo Gym dataset on Hugging Face. Nemotron-3-Super-RL-Training-Blends contains the dataset blends used to train the Nemotron-3-Super-120B-A12B model. RL training for the Nemotron-3-Super-120B-A12B model is done in 6 stages: RLVR 1, RLVR 2, RLVR 3, SWE 1, SWE 2…

nvidia/Nemotron-RL-Super-Training-Blends on the Hugging Face Hub