HF RL Explorer

nvidia/Nemotron-RL-Ultra-Training-Blends

Nemotron-RL-Ultra-Training-Blends: NeMo Gym dataset on Hugging Face. This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training…

nvidia/Nemotron-RL-Ultra-Training-Blends on the Hugging Face Hub