HF RL Explorer

Add an experimental Geometric-Mean Policy Optimization trainer, available as GMPOConfig and GMPOTrainer from…

Add an experimental Geometric-Mean Policy Optimization trainer, available as GMPOConfig and GMPOTrainer from…: a task in HF ML Tasksmith (Harbor dataset). GMPO changes how policy importance weights are aggregated. For each completion, compute token importance ratios between current and old policy…

Part of FineEnvs/HF_ML_Tasksmith.