FineEnvs/multilingual-multimodal-rl-runs
multilingual-multimodal-rl-runs: RL dataset on Hugging Face. The evidence behind the two Kannada GRPO runs in the Multilingual Multimodal Envs collection: what the model wrote for every held-out item at every checkpoint, the curves, the training logs, and the exact scripts that ran. If a number on…
FineEnvs/multilingual-multimodal-rl-runs on the Hugging Face Hub