sungyub/toolrl-4k-verl
toolrl-4k-verl: ToolRL Dataset - GPT OSS 120B Format: RL dataset on Hugging Face. A preprocessed tool-learning dataset in GPT OSS 120B native format for reinforcement learning training with GRPO/PPO algorithms.
toolrl-4k-verl: ToolRL Dataset - GPT OSS 120B Format: RL dataset on Hugging Face. A preprocessed tool-learning dataset in GPT OSS 120B native format for reinforcement learning training with GRPO/PPO algorithms.