Fix double-backprop through shared feature networks
Fix double-backprop through shared feature networks: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). In our actor-critic setup, a single FeatureNetwork produces a feature representation that is shared by several downstream heads (e.g. a value function and a policy). Each…
The task
In our actor-critic setup, a single FeatureNetwork produces a feature representation that is shared by several downstream heads (e.g. a value function and a policy). Each head computes its own loss and calls .backward(). Today, because the features the network hands out are…
Part of XiaomiMiMo/MiMo-V2.6-RL-oss.