RL environments on the Hugging Face Hub
Explore reinforcement learning environments and tasks on the Hugging Face Hub. Browse OpenEnv, Harbor, MiMo, NeMo Gym and Verifiers, inspect rewards, and run supported agent rollouts.
- Agentic RL Environments (XiaomiMiMo/MiMo-V2.6-RL-oss, MiMo RL release)
- š SmolDataEnvs (FineEnvs/SmolDataEnvs, RL dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 (nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1, NeMo Gym dataset) We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as aā¦
- HF ML Tasksmith (FineEnvs/HF_ML_Tasksmith, Harbor dataset) Fifty PR-derived Harbor tasks from Accelerate, Diffusers, PEFT, Transformers and TRL, including CPU and GPU tasks. Contains 50 Harbor tasks generated with theā¦
- Terminal-Bench 2.1 (Harbor git-repos dataset) (harborframework/terminal-bench-2.1, Harbor dataset)
- Long-Horizon Terminal-Bench (LHTB) (IntelligenceLab/Long-Horizon-Terminal-Bench, Harbor dataset) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizonā¦
- SWE-bench Science (OpenMOSS-Team/SWE-bench-Science, Harbor dataset) SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20ā¦
- nvidia/Nemotron-RL-instruction_following (nvidia/Nemotron-RL-instruction_following, NeMo Gym dataset) The Nemotron-RL-instruction following is a dataset created by combining prompts from the WildChat-1M dataset (made available under the ODC Attributionā¦
- nvidia/Nemotron-3-Nano-RL-Training-Blend (nvidia/Nemotron-3-Nano-RL-Training-Blend, NeMo Gym dataset) Nemotron-3-Nano-RL-Training-Blend is a curated dataset blend used to train the Nemotron-3-Nano-30B-A3B model. The blend consists of the following componentā¦
- SkillNet-Gym (zjunlp/SkillNet-Gym, Harbor dataset) SkillNet-Gym is a benchmark of 159 terminal-based tasks for evaluating agents in skill-rich, reproducible environments. This repository uses the native Harborā¦
- general-agent (PrimeIntellect/general-agent, Verifiers environment) Requires verifiers harbor =0.3.1. Multi-turn tool-use tasks from the self-growing general-agent toolbench, where each task ships its own tool world and a goldā¦
- Terminal-Bench (harborframework/terminal-bench, Harbor dataset)
- Spreadsheet-RL Dataset (Spreadsheet-RL/Spreadsheet-RL, RL dataset) This dataset contains the training and evaluation data used by Spreadsheet-RL, a reinforcement learning framework for spreadsheet agents that edit Excelā¦
- š SmolDataEnvs: Harbor (train) (FineEnvs/SmolDataEnvs-harbor-train, Harbor dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- MiMo-V2.6-RL Code (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-code, Harbor dataset) Fix a real issue in a real repository. 2,698 Harbor tasks from the Code domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained onā¦
- SWE-Lego-Real-Data-Verified (PrimeIntellect/SWE-Lego-Real-Data-Verified, Verifiers environment) Gold-patch-validated subset of PrimeIntellect/SWE-Lego-Real-Data (itself a fixed fork of SWE-Lego's real-data split). The resolved split contains 4,323 /ā¦
- SWE-rebench-V2-Filtered-Verified (PrimeIntellect/SWE-rebench-V2-Filtered-Verified, Verifiers environment) Filtered and gold-patch-verified subset of Nebius's SWE-rebench-V2 (paper): 6,272 / 32,079 freshly-mined GitHub PR tasks across 17 languages. Default datasetā¦
- Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 (nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1, NeMo Gym dataset) Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injectionā¦
- nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1 (nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1, NeMo Gym dataset) This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as aā¦
- Procedural Pile (reasoning-core/procedural-pile, Verifiers environment) Verifiable reasoning problems, generated by code, formatted as supervised training data. Procedural Pile contains 24,806,238 problems from 55 task familiesā¦
- Multilingual Multimodal RL runs (FineEnvs/multilingual-multimodal-rl-runs, RL dataset) The evidence behind the two Kannada GRPO runs in the Multilingual Multimodal Envs collection: what the model wrote for every held-out item at everyā¦
- TaskCompendium spike examples (open-athena/taskcompendium-spike, Harbor dataset) Schema 0.9 contains 63 semantic TaskSpec records and 164 Harbor lowerings. A TaskSpec defines the problem, semantic requirements, provenance, and privateā¦
- Calendar-Scheduling-Dataset (nvidia/Nemotron-RL-agent-calendar_scheduling, NeMo Gym dataset) The Calendar-Scheduling-Dataset is a multi-turn conversation dataset that can understand natural language scheduling constraints, follow instructions acrossā¦
- nvidia/Nemotron-RL-ReasoningGym-v1 (nvidia/Nemotron-RL-ReasoningGym-v1, NeMo Gym dataset) The Nemotron-RL-ReasoningGym-v1 dataset is designed to improve reasoning capabilities across a broad range of domains, including algebra, arithmeticā¦
- SWE-rebench-V2-Filtered-Easy-Verified (PrimeIntellect/SWE-rebench-V2-Filtered-Easy-Verified, Verifiers environment) Easy slice of PrimeIntellect/SWE-rebench-V2-Filtered-Verified: rows whose upstream LLM-judge difficulty is easy (implementation-time estimate < 15 min)ā¦
- nvidia/Nemotron-RL-instruction_following-structured_outputs (nvidia/Nemotron-RL-instruction_following-structured_outputs, NeMo Gym dataset) The Nemotron-RL-instruction following-structured outputs dataset tests the ability of the model to follow output formatting instructions under schemaā¦
- nvidia/Nemotron-RL-litmus-bench-v0.1 (nvidia/Nemotron-RL-litmus-bench-v0.1, NeMo Gym dataset) Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 testā¦
- š Paper (arXiv:2609.29050) ⢠š» Code ⢠š¤ Collection (YanZhanPKU/SLCA-GRPO-Datasets, RL dataset) This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced inā¦
- small-ow-agent-bench (catalog) (junjun77/small-ow-agent-bench, Harbor dataset) Frozen compact-shell system reliability for compact open-weight coding agents ā not coding IQ.
- FLEURS Multilingual ASR (FineEnvs/fleurs-asr-env, OpenEnv Space) OpenEnv speech recognition over 102 FLEURS languages
- ShinChan Life Simulator (Gladiator-codes/sinchan-env, OpenEnv Space) Shin-chan OpenEnv RL with TRL GRPO training and Gradio UI.
- Terminal-Bench 3.0 (harborframework/terminal-bench-3.0, Harbor dataset)
- Terminal-Bench-Science (harborframework/terminal-bench-science, Harbor dataset)
- harborframework/terminal-bench-2.0 (harborframework/terminal-bench-2.0, Harbor dataset)
- PrimeIntellect/Terminal-Lego-15k (PrimeIntellect/Terminal-Lego-15k, Harbor dataset)
- Multi-SWE-bench (PrimeIntellect/Multi-SWE-bench, Verifiers environment) Re-upload of ByteDance's Multi-SWE-bench evaluation benchmark: 2,132 issue-resolving tasks across the seven Multi-SWE languages. This is the held-out evalā¦
- Repo2RLEnv SWE-Flow (FineEnvs/repo2rlenv-swe-flow, Harbor dataset) Trace passing repository tests, select exercised source functions, remove their implementations and generate a reconstruction instruction. Restore theā¦
- Reverse-Text-RL (PrimeIntellect/Reverse-Text-RL, Verifiers environment) A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the generalā¦
- TaskTrove (open-thoughts/TaskTrove, Harbor dataset) v5.1 (current) ā independent-review source retirement ā moves 15 sources with majority or unanimous REJECT verdicts out of the default config and intoā¦
- š§Ŗ Data Agent ā Harbor (eval) (FineEnvs/data-agent-harbor-eval, Harbor dataset) A small, difficulty-balanced validation split ā 144 tasks ā perfect for quick checkpoints while you train. Same idea as the rest of the family: your agentā¦
- skilltrainbench training tasks (armin-aptura/skilltrainbench-public, Harbor dataset) The training half of the skilltrainbench benchmark suite: for each of the four datasets, the dev task names of its pinned train/test split, in Harbor taskā¦
- R2E-Gym-Subset-Verified (PrimeIntellect/R2E-Gym-Subset-Verified, Verifiers environment) Gold-patch-validated subset of R2E-Gym/R2E-Gym-Subset (paper). The train split contains 4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply theā¦
- WildClawBench-Harbor (internlm/WildClawBench-Harbor, Harbor dataset) This repository is the WildClawBench benchmark converted to the Harbor task format, so that all 60 tasks can be run directly with harbor run against anyā¦
- LHTB Leaderboard ā Long-Horizon Terminal-Bench (IntelligenceLab/LHTB-leaderboard, Harbor dataset) This repository hosts submitted runs for Long-Horizon Terminal-Bench (LHTB), a 46-task benchmark measuring how well LLM agents sustain useful work in aā¦
- Multi-SWE-RL-Verified (PrimeIntellect/Multi-SWE-RL-Verified, Verifiers environment) Gold-patch-validated subset of PrimeIntellect/Multi-SWE-RL-Reupload (ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across C, Go, Java, JavaScript, Rust, andā¦
- Human Behavior Atlas (HumanBehaviorAtlas/human_behavior_atlas, RL dataset) A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screeningā¦
- š Data Agent ā Harbor (train) (FineEnvs/data-agent-harbor-train, Harbor dataset) Teach an agent to actually do data science. This is a suite of 5,000 hands-on data-analysis tasks: each one drops your agent into a sandbox with a realā¦
- Scale-SWE-Verified (PrimeIntellect/Scale-SWE-Verified, Verifiers environment) Gold-patch-validated fork of AweAI-Team/Scale-SWE (paper): 17,202 / 20,181 Python issue-resolving tasks that produce a clean reward signal end-to-end. Defaultā¦
- PatchRecoveryGym for Laguna (poolside-laguna-hackathon/patchrecoverygym-laguna, Verifiers environment) Submitted by: Kannappan Sirchabesan (@kannappans) Ā· Poolside Research Hackathon (Foundations track) A reproducible eval + RL environment that tests whether aā¦
- Repo2RLEnv CLI-Gym (FineEnvs/repo2rlenv-cli-gym, Harbor dataset) Start with a healthy repository, synthesize a disruption and recovery, verify that damage breaks the original tests and recovery restores them, then export anā¦
- dabstep-harbor (AdithyaSK/dabstep-harbor, Harbor dataset) A /workdir-native Harbor edition of DABstep. 450 financial-analytics factoid tasks over a single shared Adyen-payments corpus ā every task's workingā¦
- Lego-X/LegoFlow-SWE (Lego-X/LegoFlow-SWE, Harbor dataset) LegoFlow-SWE Ā· 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases
- nvidia/Nemotron-RL-agent-workplace_assistant (nvidia/Nemotron-RL-agent-workplace_assistant, NeMo Gym dataset) The Nemotron-RL-agent-workplace assistant is a tool use - multi step agentic environment that tests the agentās ability to execute tasks in a workplaceā¦
- AgentTrove (open-thoughts/AgentTrove, Harbor dataset) AgentTrove is the largest open-source collection of agentic interaction traces to date, released by the OpenThoughts-Agent team. It contains 1,696,847 rowsā¦
- SWE-Lego-RL-2699 (Lego-X/Lego-RL-2699, Harbor dataset) 2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two parallel views of the same instances:
- MiMo-V2.6-RL Terminal (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-terminal, Harbor dataset) Terminal-Bench style tasks. 64 Harbor tasks from the Terminal domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted soā¦
- TerminalWorld Seeds, oracle-validated (andylizf/TerminalWorld-Seeds-Clean, Harbor dataset) The current TW selection retains 405 tasks, excluding any original TW task excluded by either the preserved local quality curation or the upstream v3 trainingā¦
- Repo2RLEnv SWE-smith (FineEnvs/repo2rlenv-swe-smith, Harbor dataset) Apply owned procedural mutations to real repository functions, require an executable test contrast, and author a natural issue from bounded evidence. Containsā¦
- šÆ Data Agent ā Harbor (test) (FineEnvs/data-agent-harbor-test, Harbor dataset) A held-out benchmark for data-analysis agents: 250 tasks, deliberately balanced across difficulty and leaning toward the harder end so it actually separatesā¦
- MiMo-V2.6-RL General (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-general, Harbor dataset) Work in a simulated company through its MCP systems. 925 Harbor tasks from the General domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 wasā¦
- Repo2RLEnv DataArc terminal synthesis (Envs-FORGE-linked code) (FineEnvs/repo2rlenv-dataarc, Harbor dataset) Read complete Harbor parent tasks and enumerate few-shot, self-instruct and evolution transformations. Author a complete child environment, instructionā¦
- Terminal-Bench Hard (Zhongzhi1228/Terminal-Bench-Hard, Harbor dataset) Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, systemā¦
- Repo2RLEnv Endless Terminals (FineEnvs/repo2rlenv-endless-terminals, Harbor dataset) Sample native terminal categories, complexity and scenarios; author instructions and private expected behavior, initial-state tests, final tests, fixtures andā¦
- Repo2RLEnv R2E (FineEnvs/repo2rlenv-r2e, Harbor dataset) Author equivalence tests around repository functions, measure their execution and coverage remotely, and export function-reconstruction tasks with theā¦
- FACET-Terminal-Tasks-6k (FACET-Terminal/FACET-Terminal-Tasks-6k, Harbor dataset) 6,020 execution-grounded tasks for terminal agents, coding agents, and executable workflow research
- Repo2RLEnv SWE-gen (FineEnvs/repo2rlenv-swe-gen, Harbor dataset)
- MiMo-V2.6-RL Cyber (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-cyber, Harbor dataset) Reproduce a real memory-safety crash (ARVO). 1,000 Harbor tasks from the Cyber domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trainedā¦
- Repo2RLEnv R2E-Gym / SWEGEN (FineEnvs/repo2rlenv-r2e-gym, Harbor dataset) Mine supported bug-fix commits, measure failing-to-passing and regression tests across revisions, and export the buggy starting state and deterministicā¦
- Repo2RLEnv TerminalWorld (EuniAI) (FineEnvs/repo2rlenv-terminalworld, Harbor dataset) Screen original public terminal recordings, score feasibility and value, extract and refine a solution, write its instruction, reconstruct an offlineā¦
- š SmolDataEnvs: Harbor (test) (FineEnvs/SmolDataEnvs-harbor-test, Harbor dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- repo2rlenv-commit-runtime-v2 (FineEnvs/repo2rlenv-commit-runtime, Harbor dataset) Generated by Repo2RLEnv ā turning real GitHub repositories into verifiable RL environments.
- MiMo-V2.6-RL Music (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-music, Harbor dataset) Compose a piece in ABC notation. 1,000 Harbor tasks from the Music domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained onā¦
- repo2rlenv-pr-runtime (FineEnvs/repo2rlenv-pr-runtime, Harbor dataset) Generated by Repo2RLEnv ā turning real GitHub repositories into verifiable RL environments.
- Repo2RLEnv SCALER (FineEnvs/repo2rlenv-scaler, Harbor dataset) Generate instances from owned reasoning families with seeded procedural samplers, deterministic references and family-specific answer verification; generationā¦
- Repo2RLEnv TMax (FineEnvs/repo2rlenv-tmax, Harbor dataset) Sample the native TMax legacy taxonomy, create a template, initial-state tests and final completion tests in distinct stages, then build fixtures and aā¦
- TaskTrove Clean (open-athena/task-trove, Harbor dataset) The current train split contains 930,918 tasks. The original conversion report below describes the original population. Native additions are documented inā¦
- š SmolDataEnvs: Harbor (eval) (FineEnvs/SmolDataEnvs-harbor-eval, Harbor dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- nvidia/Nemotron-RL-knowledge-mcqa (nvidia/Nemotron-RL-knowledge-mcqa, NeMo Gym dataset)
- Repo2RLEnv SETA Seed2Synth (FineEnvs/repo2rlenv-seta-seed2synth, Harbor dataset) Use attributed Unix Stack Exchange questions and accepted answers to design terminal tasks, then author self-contained offline environments, instructionsā¦
- Repo2RLEnv SWE-Next (FineEnvs/repo2rlenv-swe-next, Harbor dataset) Mine real repository history for supported code changes and associated tests, preserve regression behavior, and export the historical implementation change asā¦
- MiMo-V2.6-RL Webdev (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-webdev, Harbor dataset) Build a website from a design brief. 2,093 Harbor tasks from the Webdev domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained onā¦
- Terminal-Bench 2.1 (offline Apptainer, v1) (laion/terminal-bench-2-1-apptainer-v1, Harbor dataset) A validated subset of Terminal-Bench 2.1 (revision 7131e4375048a0e408a8fb404b5f499d726b695b, Apache-2.0) for running on HPC clusters without Docker andā¦
- AdithyaSK/data_agent_rl_environment_train_multireward (AdithyaSK/data_agent_rl_environment_train_multireward, Harbor dataset) Multi-reward variant of AdithyaSK/data agent rl environment train (2238 tasks, identical data/instructions). The only change: each task's verifier now emits aā¦
- Keep the failed attempts. Check the artifact. (glayguo/noteflow-research-pilots, Harbor dataset) Versioned public development evidence from Robot Reel Ć Skills Anywhere Ć EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harborā¦
- nvidia/Nemotron-RL-Ultra-Training-Blends (nvidia/Nemotron-RL-Ultra-Training-Blends, NeMo Gym dataset) This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultraā¦
- Watercolour reference pool (FineEnvs/watercolour-reference-pool, RL dataset) The reference paintings that define the reward in the watercolour RL environment: an agent writes a p5.brush sketch, the sketch is rendered, and a visionā¦
- repo2rlenv-pr-diff (FineEnvs/repo2rlenv-pr-diff, Harbor dataset) Generated by Repo2RLEnv ā turning real GitHub repositories into verifiable RL environments.
- AACR-Bench Harbor (osolmaz/aacr-bench-harbor, Harbor dataset)
- Repo2RLEnv SETA Evol (FineEnvs/repo2rlenv-seta-evol, Harbor dataset) Start with complete Harbor parents, apply one of six named evolution strategies, then author a complete child task with its own instruction, environmentā¦
- repo2rlenv-equivalence-tests (FineEnvs/repo2rlenv-equivalence-tests, Harbor dataset) Generated by Repo2RLEnv ā turning real GitHub repositories into verifiable RL environments.
- ProgramDistill-Bench (microsoft/ProgramDistill, Harbor dataset) ProgramDistill is a dataset and executable benchmark for evaluating whether coding agents can restore missing behavior in interactive web applications. Eachā¦
- repo2rlenv-code-instruct (FineEnvs/repo2rlenv-code-instruct, Harbor dataset) Generated by Repo2RLEnv ā turning real GitHub repositories into verifiable RL environments.
- Watercolour rollouts, judge-led run (FineEnvs/watercolour-rollouts-judge-led, RL dataset) Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
- LoLBench (lolbench26/LoLBench, Harbor dataset) Software development often begins with user intent and high-level design rather than an implementation-ready specification. Developers must ground those goalsā¦
- Monthly-SWEBench 2026-05 (UnipatAI/Monthly-SWEBench-2026-05, Harbor dataset) This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRsā¦
- nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1 (nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1, NeMo Gym dataset) The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. Theā¦
- Monthly-SWEBench-2026-04 (UnipatAI/Monthly-SWEBench-2026-04, Harbor dataset) 90 tasks ā 43 bugfix + 47 non-bugfix Tasks span diverse open-source repositories Each task includes a runnable environment, test suite, and reference solution
- Monthly-SWEBench-2026-03 (UnipatAI/Monthly-SWEBench-2026-03, Harbor dataset) 112 tasks ā 68 bugfix + 44 non-bugfix Tasks span diverse open-source repositories Each task includes a runnable environment, test suite, and reference solution
- TerminalWorld Seeds, packaged in the RST release layout (andylizf/TerminalWorld-Seeds, Harbor dataset) 1,530 validated terminal tasks from EuniAI/TerminalWorld, repackaged in the release layout of Zhongzhi1228/Recursive-Task-Synthesis (RST, arXiv:2608.05466)ā¦
- SkillFlow Test Tasks (zhang-ziao/SkillFlow-Task, Harbor dataset) SkillFlow Test Tasks is the task repository used in the SkillFlow benchmark for evaluating lifelong skill discovery, skill revision, and cross-task proceduralā¦
- nvidia/Nemotron-RL-coding-competitive_coding (nvidia/Nemotron-RL-coding-competitive_coding, NeMo Gym dataset) The Nemotron-RL-coding-competitive coding dataset is a python-only, reasoning-based, synthetic dataset. It contains competitive coding style problems andā¦
- jupyter-agent-eval-v1 ā Harbor task suite (AdithyaSK/jupyter-agent-eval-v1-harbor, Harbor dataset) 100 Harbor task(s) for the Jupyter data-analysis agent. Each task is one (question, gold answer) pair against a real Kaggle dataset.
- Human Behavior Atlas (DennisDengHUst/human_behavior_atlas, RL dataset) A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screeningā¦
- repo2rlenv-cve-patches (FineEnvs/repo2rlenv-cve-patches, Harbor dataset) Generated by Repo2RLEnv ā turning real GitHub repositories into verifiable RL environments.
- HubBench 1.4.0 (SamuelChien821/hubbench, Harbor dataset) One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependentā¦
- Nemotron Terminal easy clean-1K: rootless Apptainer release (laion/nemotron-terminal-easy-clean1k-rootless-apptainer, Harbor dataset) This derived release contains exactly 1,000 balanced easy tasks from four Nemotron Terminal domains. It is pinned toā¦
- PDBThink Coordinate Tasks (open-athena/pdbthink-coordinate-tasks, Harbor dataset) 100,000 new coordinate-interpretation tasks across all 19 active PDBThink families, from 2,671 experimental PDB entries in 1,867 source groups. Version 1.3.0ā¦
- rounakbende/rh-swe-bench (rounakbende/rh-swe-bench, Harbor dataset) A Red Hat-specific benchmark for evaluating coding agents, packaged in Harbor format. Contains 337 tasks across 24 Red Hat ecosystem repositories. Theā¦
- nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 (nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2, NeMo Gym dataset) Split 1: Direct Generation tests the modelās ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity andā¦
- AdithyaSK/data_agent_rl_environment_train_difficulty_ranked (AdithyaSK/data_agent_rl_environment_train_difficulty_ranked, Harbor dataset) A Harbor task suite of 2238 data-agent tasks, ordered easy - hard by empirical difficulty measured from a pass@4 rollout sweep (Qwen3.5-4B + 2B, bash harness).
- data_agent_rl_environment_train (AdithyaSK/data_agent_rl_environment_train, Harbor dataset) The official verified training suite for the data-agent RL pipeline. 2238 Harbor-format data-analysis tasks, each with:
- RTS-1K: offline rootless Apptainer release (laion/rts-1k-rootless-apptainer, Harbor dataset) This derived release contains 984 tasks from Recursive Task Synthesis. RTS-1K is the stable project label; the exact size is 984. The tasks are the strictā¦
- DeepScaleR-Preview VERL (sungyub/deepscaler-preview-verl, RL dataset) This dataset contains 35,789 mathematical reasoning problems in VERL format, processed from agentica-org/DeepScaleR-Preview-Dataset. Key Features:
- Endless Terminals 1K: rootless Apptainer release (laion/endless-terminals-rootless-apptainer, Harbor dataset) This release contains exactly 1,000 tasks from Endless Terminals, pinned to obiwan96/endless-terminals@26ecf78458e7f756e4d06780d5fbf3dd78e91815 and validatedā¦
- Harbor SWE-Smith å¼ŗåå¦ä¹ ę°ę®äŗ§ē© (keryszhan/harbor-swesmith-rl-artifacts, Harbor dataset) ę¬ę°ę®éęÆ Harbor Qwen å·„å ·č°ēØä»£ē ęŗč½ä½å¼ŗåå¦ä¹ 锹ē®ä½æēØēå»ē»ä»»å”éļ¼ęå”äŗ GRPOćåēä»·å¼ęØ”å/GAE PPOćč®ē»čæēØčÆęåē»äøåč®®čÆęµć 锹ē®å·²äŗ 2026 幓 8 ę 30 ę„å®ę P0ā¦
- nvidia/Nemotron-RL-Instruction-Following-Calendar-v2 (nvidia/Nemotron-RL-Instruction-Following-Calendar-v2, NeMo Gym dataset) The Calendar-Scheduling-Dataset is a multi-turn conversation dataset that can understand natural language scheduling constraints, follow instructions acrossā¦
- introvoyz041/terminal-bench-2.0 (introvoyz041/terminal-bench-2.0, Harbor dataset)
- Agent trajectory study (pranay5255/agent-trajectory-study-20260930, Harbor dataset) This corpus contains 1,451 recorded episodes: 570 OLMo, 510 parseable Qwen, and 371 forestOfAudits episodes. It retains companion views and rejected OLMoā¦
- FlowBench (jwu323/FlowBench, Harbor dataset) Dataset ID: jwu323/FlowBench FlowBench is a tool-use benchmark for deterministic business operations workflows. Each task asks an agent to compose Pythonā¦
- MegaTerminal (di-zhang-fdu/MegaTerminal, Harbor dataset) MegaTerminal is a Harbor-style terminal task-folder dataset assembled from the terminal task sources collected for task matching experiments. The task foldersā¦