The following issues were observed when working with the verl training framework and need to be resolved.
The following issues were observed when working with the verl training framework and need to be resolved.: a task in LegoFlow-SWE (Harbor dataset). When loading a Megatron model onto GPU with gradients requested ( load grad=True ), the internal logic attempts to resize the gradient storage by…
The task
When loading a Megatron model onto GPU with gradients requested (`load_grad=True`), the internal logic attempts to resize the gradient storage by accessing a `grad_data_size` attribute on gradient buffer objects. However, in some Megatron versions these buffer objects do not expose a `grad_data_size` attribute,…
Part of Lego-X/LegoFlow-SWE.