The GRPO algorithm currently offers no mechanism to save and resume training runs. Users who train GRPO…
The GRPO algorithm currently offers no mechanism to save and resume training runs. Users who train GRPO…: a task in LegoFlow-SWE (Harbor dataset). Observed behavior: GRPO training proceeds normally, but if the process is stopped (e.g., due to a crash, preemption, or manual interruption), all…
The task
**Observed behavior:** GRPO training proceeds normally, but if the process is stopped (e.g., due to a crash, preemption, or manual interruption), all learned parameters, optimizer history, scheduler state, and the current training step are lost. When the training script is restarted, it begins from the initial state,…
Part of Lego-X/LegoFlow-SWE.