Make offline evaluation usable and enforce label consistency
Make offline evaluation usable and enforce label consistency: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). The library exposes a public evaluation helper evaluate, importable from libreco.evaluation. It is meant to score an already-fitted model on a chunk of data and…
The task
The library exposes a public evaluation helper evaluate, importable from libreco.evaluation. It is meant to score an already-fitted model on a chunk of data and return the requested metrics. Two problems: the entry point is currently broken (importing/using it blows up), and…
Part of XiaomiMiMo/MiMo-V2.6-RL-oss.