Add a memory-efficient chunked cross-entropy loss mode to TRL's SFT trainer that avoids materializing full…
Add a memory-efficient chunked cross-entropy loss mode to TRL's SFT trainer that avoids materializing full…: a task in HF ML Bench v0 (Harbor dataset). Add a new loss mode to TRL's supervised fine-tuning trainer that computes the language-modeling cross-entropy loss WITHOUT ever materializing the…
The task
Add a new loss mode to TRL's supervised fine-tuning trainer that computes the language-modeling cross-entropy loss WITHOUT ever materializing the full `[batch × seq × vocab]` logits tensor. The point is peak-memory reduction: on long sequences with large vocabularies, the standard path spends most of its activation…
Part of AdithyaSK/HF_ML_Bench_v0.