HF RL Explorer

Add chunked LM head for memory-efficient log-prob computation for AsyncGRPOTrainer

Add chunked LM head for memory-efficient log-prob computation for AsyncGRPOTrainer: a task in HF ML Bench v0 (Harbor dataset). The asynchronous GRPO trainer computes, per token, the log-probability of the sampled token and the entropy of the model's next-token distribution. The naive…

The task

The asynchronous GRPO trainer computes, per token, the log-probability of the sampled token and the entropy of the model's next-token distribution. The naive implementation runs the LM head over the full hidden state, materializing a logits tensor of shape `[N, V]` (positions × vocabulary size). For long sequences…

Part of AdithyaSK/HF_ML_Bench_v0.