HF RL Explorer

jayshah5696/humanize-rl-tasks

humanize-rl-tasks: Verifiers environment on Hugging Face. Single-turn RL task dataset for the humanize-rl project. Each row is one writing task. A model receives the prompt (instruction + source text), produces a completion, and the environment scores it with the 50/50 reward formula: reward =…

jayshah5696/humanize-rl-tasks on the Hugging Face Hub