HF RL Explorer

The training workflow already constructs linear, gradient-boosted-tree, and random-forest contextual…

The training workflow already constructs linear, gradient-boosted-tree, and random-forest contextual…: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). Make the public BanditPredictor workflow genuinely backend-agnostic so users can train, serve, persist, and reload any of…

The task

Make the public BanditPredictor workflow genuinely backend-agnostic so users can train, serve, persist, and reload any of the documented model families. A predictor must preprocess one request exactly as today: expand it into the experiment's ordered decision candidates, apply…

Part of XiaomiMiMo/MiMo-V2.6-RL-oss.