HF RL Explorer

Fix double-backprop through shared feature networks

Fix double-backprop through shared feature networks: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). In our actor-critic setup, a single FeatureNetwork produces a feature representation that is shared by several downstream heads (e.g. a value function and a policy). Each…

The task

In our actor-critic setup, a single FeatureNetwork produces a feature representation that is shared by several downstream heads (e.g. a value function and a policy). Each head computes its own loss and calls .backward(). Today, because the features the network hands out are…

Part of XiaomiMiMo/MiMo-V2.6-RL-oss.