Add an experimental Geometric-Mean Policy Optimization trainer, available as GMPOConfig and GMPOTrainer from…
Add an experimental Geometric-Mean Policy Optimization trainer, available as GMPOConfig and GMPOTrainer from…: a task in HF ML Tasksmith (Harbor dataset). GMPO changes how policy importance weights are aggregated. For each completion, compute token importance ratios between current and old policy…
Part of FineEnvs/HF_ML_Tasksmith.