HF RL Explorer

The training workflow already constructs linear, gradient-boosted-tree, and random-forest contextual…

The training workflow already constructs linear, gradient-boosted-tree, and random-forest contextual…: a task in MiMo-V2.6-RL-harbor-code: MiMo-V2.6-RL Code (Harbor) (Harbor dataset). A predictor must preprocess one request exactly as today: expand it into the experiment's ordered decision…

The task

A predictor must preprocess one request exactly as today: expand it into the experiment's ordered decision candidates, apply the configured numeric/categorical/dense-product transformations, and preserve candidate IDs in experiment order. For neural models, retain the existing behavior. For sklearn-style estimators…

Part of FineEnvs/MiMo-V2.6-RL-harbor-code.