HF RL Explorer

Repair persisted MLP incremental training across the estimator and stochastic optimizer source modules.

Repair persisted MLP incremental training across the estimator and stochastic optimizer source modules.: a task in MiMo-V2.6-RL-harbor-terminal: MiMo-V2.6-RL Terminal (Harbor) (Harbor dataset). The frozen scikit-learn source slice under /app/vendor/scikit-learn has a regression in a stateful CPU…

Part of FineEnvs/MiMo-V2.6-RL-harbor-terminal.