The training workflow already constructs linear, gradient-boosted-tree, and random-forest contextual…
The training workflow already constructs linear, gradient-boosted-tree, and random-forest contextual…: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). Make the public BanditPredictor workflow genuinely backend-agnostic so users can train, serve, persist, and reload any of…
The task
Make the public BanditPredictor workflow genuinely backend-agnostic so users can train, serve, persist, and reload any of the documented model families. A predictor must preprocess one request exactly as today: expand it into the experiment's ordered decision candidates, apply…
Part of XiaomiMiMo/MiMo-V2.6-RL-oss.