Add a "consecutive irrelevant" stopping rule
Add a "consecutive irrelevant" stopping rule: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). The active learning loop decides when to stop screening through a stopping mechanism: a small object with a stop(results, data) method that returns True when the review should halt…
The task
The active learning loop decides when to stop screening through a stopping mechanism: a small object with a stop(results, data) method that returns True when the review should halt. The project already ships a handful of these (for example StoppingDefault, StoppingN…
Part of XiaomiMiMo/MiMo-V2.6-RL-oss.