I want python algos/main.py to support a TRPO training mode selected with --algo-name TRPO , and to use…
I want python algos/main.py to support a TRPO training mode selected with --algo-name TRPO , and to use…: a task in MiMo-V2.6-RL-harbor-code: MiMo-V2.6-RL Code (Harbor) (Harbor dataset). In TRPO mode, argument parsing should provide the trust-region settings --max-kl and --damping with defaults of…
The task
In TRPO mode, argument parsing should provide the trust-region settings `--max-kl` and `--damping` with defaults of `1e-2`, while the existing `--l2-reg` value regularizes value-function fitting. The batch update should fit the value network to the computed returns with L-BFGS, compute the policy surrogate from old…
Part of FineEnvs/MiMo-V2.6-RL-harbor-code.