HF RL Explorer

I'm seeing weird samples when I train DQN/DDPG with batched environments and replay: n-step sequences and…

I'm seeing weird samples when I train DQN/DDPG with batched environments and replay: n-step sequences and…: a task in MiMo-V2.6-RL-oss: Agentic RL Environments (MiMo RL release). It also looks like when one batch item hits done/reset, the replay state for the other still-running envs gets…

The task

It also looks like when one batch item hits done/reset, the replay state for the other still-running envs gets disturbed. This shows up with interleaved env steps where sampled trajectories have state/action markers from multiple envs. Expected outcomes Sampled n-step replay…

Part of XiaomiMiMo/MiMo-V2.6-RL-oss.