jayshah5696/humanize-rl-tasks
humanize-rl-tasks: Verifiers environment on Hugging Face. Single-turn RL task dataset for the humanize-rl project. Each row is one writing task. A model receives the prompt (instruction + source text), produces a completion, and the environment scores it with the 50/50 reward formula: reward =…