Feature Request: Custom reward function support in GRPOTrainer
Feature Request: Custom reward function support in GRPOTrainer: a task in LegoFlow-SWE (Harbor dataset). Currently, GRPOTrainer only supports reward models (sequence classification models) for computing rewards during training. It is not possible to pass a plain Python callable — for example, a…
The task
Currently, `GRPOTrainer` only supports reward models (sequence classification models) for computing rewards during training. It is not possible to pass a plain Python callable — for example, a function that checks formatting, computes accuracy, or implements any other domain-specific reward signal — directly to the…
Part of Lego-X/LegoFlow-SWE.