HF RL Explorer

nvidia/Nemotron-RL-Safety-v1

Nemotron-RL-Safety-v1: NeMo Gym dataset on Hugging Face. The Nemotron-RL-Safety-v1 data is designed to provide labeled comparisons necessary to train Reward Models to distinguish between safe, helpful responses and undesired, non-compliant outputs. This dataset is a collection of:

nvidia/Nemotron-RL-Safety-v1 on the Hugging Face Hub