Feature Request: Add debug mode to collators for inspecting tokenization
Feature Request: Add debug mode to collators for inspecting tokenization: a task in LegoFlow-SWE (Harbor dataset). Problem When training with Oumi’s data collators, there is currently no built-in way to inspect how raw dataset examples are being processed into model-ready tokens. This makes…
The task
**Problem** When training with Oumi’s data collators, there is currently no built-in way to inspect how raw dataset examples are being processed into model-ready tokens. This makes debugging tokenization issues—such as prompt template formatting, token ID mapping, and label masking—extremely difficult.
Part of Lego-X/LegoFlow-SWE.