Fix three bugs in OnlineDPOTrainer. generate vllm server so that it produces the same block-layout batch…
Fix three bugs in OnlineDPOTrainer. generate vllm server so that it produces the same block-layout batch…: a task in HF ML Tasksmith (Harbor dataset). Background. Online DPO generates exactly 2 completions per prompt to form preference pairs. Everything downstream ( rewards.split(batch size) )…
Part of FineEnvs/HF_ML_Tasksmith.