Correct reward-margin reporting in the experimental KTO trainer ( trl/experimental/kto/kto trainer.py ). A…
Correct reward-margin reporting in the experimental KTO trainer ( trl/experimental/kto/kto trainer.py ). A…: a task in HF ML Tasksmith (Harbor dataset). During training and evaluation, each batch containing both chosen and rejected examples contributes one rewards/margins value. That value is the…
The task
During training and evaluation, each batch containing both chosen and rejected examples contributes one `rewards/margins` value. That value is the mean chosen reward minus the mean rejected reward for the current batch, using the same aggregated example counts and rewards as the other reward metrics. This must work…
Part of FineEnvs/HF_ML_Tasksmith.