Explore RL environments on Hugging Face
411 RL environment datasets and 6,819 environment Spaces on Hugging Face, across OpenEnv, Harbor, Verifiers, NeMo Gym and more: see what each task asks and how it's graded, then run an agent on it.
- Agentic RL Environments (XiaomiMiMo/MiMo-V2.6-RL-oss, MiMo RL release)
- 📈 SmolDataEnvs (FineEnvs/SmolDataEnvs, RL dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 (nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1, NeMo Gym dataset) We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a…
- nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1 (nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1, NeMo Gym dataset) This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a…
- nvidia/Nemotron-RL-instruction_following (nvidia/Nemotron-RL-instruction_following, NeMo Gym dataset) The Nemotron-RL-instruction following is a dataset created by combining prompts from the WildChat-1M dataset (made available under the ODC Attribution…
- Calendar-Scheduling-Dataset (nvidia/Nemotron-RL-agent-calendar_scheduling, NeMo Gym dataset) The Calendar-Scheduling-Dataset is a multi-turn conversation dataset that can understand natural language scheduling constraints, follow instructions across…
- 📊 SmolDataEnvs: Harbor (train) (FineEnvs/SmolDataEnvs-harbor-train, Harbor dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- MiMo-V2.6-RL Code (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-code, Harbor dataset) Fix a real issue in a real repository. 2,698 Harbor tasks from the Code domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on…
- SWE-bench Science (OpenMOSS-Team/SWE-bench-Science, Harbor dataset) SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20…
- ProgramDistill-Bench (microsoft/ProgramDistill, Harbor dataset) ProgramDistill is a dataset and executable benchmark for evaluating whether coding agents can restore missing behavior in interactive web applications. Each…
- nvidia/Nemotron-RL-ReasoningGym-v1 (nvidia/Nemotron-RL-ReasoningGym-v1, NeMo Gym dataset) The Nemotron-RL-ReasoningGym-v1 dataset is designed to improve reasoning capabilities across a broad range of domains, including algebra, arithmetic…
- Terminal-Bench 2.1 (Harbor git-repos dataset) (harborframework/terminal-bench-2.1, Harbor dataset)
- PrimeIntellect/Terminal-Lego-15k (PrimeIntellect/Terminal-Lego-15k, Harbor dataset)
- HF ML Tasksmith (FineEnvs/HF_ML_Tasksmith, Harbor dataset) Fifty PR-derived Harbor tasks from Accelerate, Diffusers, PEFT, Transformers and TRL, including CPU and GPU tasks. Contains 50 Harbor tasks generated with the…
- Long-Horizon Terminal-Bench (LHTB) (IntelligenceLab/Long-Horizon-Terminal-Bench, Harbor dataset) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon…
- R2E-Gym-Subset-Verified (PrimeIntellect/R2E-Gym-Subset-Verified, Verifiers environment) Gold-patch-validated subset of R2E-Gym/R2E-Gym-Subset (paper). The train split contains 4,522 / 4,578 rows (98.78%) verified scoreable end-to-end: apply the…
- SWE-Lego-Real-Data-Verified (PrimeIntellect/SWE-Lego-Real-Data-Verified, Verifiers environment) Gold-patch-validated subset of PrimeIntellect/SWE-Lego-Real-Data (itself a fixed fork of SWE-Lego's real-data split). The resolved split contains 4,323 /…
- SWE-rebench-V2-Filtered-Verified (PrimeIntellect/SWE-rebench-V2-Filtered-Verified, Verifiers environment) Filtered and gold-patch-verified subset of Nebius's SWE-rebench-V2 (paper): 6,272 / 32,079 freshly-mined GitHub PR tasks across 17 languages. Default dataset…
- nvidia/Nemotron-RL-knowledge-mcqa (nvidia/Nemotron-RL-knowledge-mcqa, NeMo Gym dataset)
- nvidia/Nemotron-RL-Ultra-Training-Blends (nvidia/Nemotron-RL-Ultra-Training-Blends, NeMo Gym dataset) This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra…
- Procedural Pile (reasoning-core/procedural-pile, Verifiers environment) Verifiable reasoning problems, generated by code, formatted as supervised training data. Procedural Pile contains 24,806,238 problems from 55 task families…
- MiMo-V2.6-RL Cyber (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-cyber, Harbor dataset) Reproduce a real memory-safety crash (ARVO). 1,000 Harbor tasks from the Cyber domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained…
- Terminal-Bench 2.1 (offline Apptainer, v1) (laion/terminal-bench-2-1-apptainer-v1, Harbor dataset) A validated subset of Terminal-Bench 2.1 (revision 7131e4375048a0e408a8fb404b5f499d726b695b, Apache-2.0) for running on HPC clusters without Docker and…
- Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 (nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1, NeMo Gym dataset) Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection…
- nvidia/Nemotron-RL-coding-competitive_coding (nvidia/Nemotron-RL-coding-competitive_coding, NeMo Gym dataset) The Nemotron-RL-coding-competitive coding dataset is a python-only, reasoning-based, synthetic dataset. It contains competitive coding style problems and…
- TaskTrove Clean (open-athena/task-trove, Harbor dataset) TaskTrove Clean is a normalized release of open-thoughts/TaskTrove at revision 0292300. It contains 861,848 retained Harbor tasks from 1,739,326 input rows…
- nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 (nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2, NeMo Gym dataset) Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and…
- TaskCompendium spike examples (open-athena/taskcompendium-spike, Harbor dataset) Schema 0.9 contains 63 semantic TaskSpec records and 164 Harbor lowerings. A TaskSpec defines the problem, semantic requirements, provenance, and private…
- nvidia/Nemotron-RL-Science-v1 (nvidia/Nemotron-RL-Science-v1, NeMo Gym dataset) Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable…
- nvidia/Nemotron-RL-Super-Training-Blends (nvidia/Nemotron-RL-Super-Training-Blends, NeMo Gym dataset) Nemotron-3-Super-RL-Training-Blends contains the dataset blends used to train the Nemotron-3-Super-120B-A12B model. RL training for the…
- nvidia/Nemotron-RL-instruction_following-structured_outputs (nvidia/Nemotron-RL-instruction_following-structured_outputs, NeMo Gym dataset) The Nemotron-RL-instruction following-structured outputs dataset tests the ability of the model to follow output formatting instructions under schema…
- SWE-rebench-V2-Filtered-Easy-Verified (PrimeIntellect/SWE-rebench-V2-Filtered-Easy-Verified, Verifiers environment) Easy slice of PrimeIntellect/SWE-rebench-V2-Filtered-Verified: rows whose upstream LLM-judge difficulty is easy (implementation-time estimate < 15 min)…
- nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1 (nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1, NeMo Gym dataset) The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting…
- nvidia/Nemotron-3-Nano-RL-Training-Blend (nvidia/Nemotron-3-Nano-RL-Training-Blend, NeMo Gym dataset) Nemotron-3-Nano-RL-Training-Blend is a curated dataset blend used to train the Nemotron-3-Nano-30B-A3B model. The blend consists of the following component…
- nvidia/Nemotron-RL-math-OpenMathReasoning (nvidia/Nemotron-RL-math-OpenMathReasoning, NeMo Gym dataset) The Nemotron-RL-math-OpenMathReasoning dataset contains mathematical problems and solutions sourced from the AoPS forums. These problems and solutions were…
- 📄 Paper (arXiv:2609.29050) • 💻 Code • 🤗 Collection (YanZhanPKU/SLCA-GRPO-Datasets, RL dataset) This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in…
- nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1 (nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1, NeMo Gym dataset) The inverseIF dataset focuses on adversarial prompts designed to explicitly conflict with an AI model’s standard training instincts—such as writing code…
- nvidia/Nemotron-RL-Instruction-Following-Calendar-v2 (nvidia/Nemotron-RL-Instruction-Following-Calendar-v2, NeMo Gym dataset) The Calendar-Scheduling-Dataset is a multi-turn conversation dataset that can understand natural language scheduling constraints, follow instructions across…
- nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1 (nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1, NeMo Gym dataset) Teaches the model to follow arbitrary text formatting instructions (bullet styles, numbering, delimiters, heading formats, inline emphasis, web-answer…
- nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1 (nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1, NeMo Gym dataset) Teaches the model to cite specific document parts using reference markers like ref:1 , ref:3, etc. Supports single-reference, multi-reference, and inline…
- Unified Code VERL Dataset (sungyub/code-verl-unified, RL dataset) This dataset aggregates seven code-reasoning collections into a single VERL-formatted repository containing approximately 958,539 unique problems. The…
- nvidia/Nemotron-RL-Identity-Following-v1 (nvidia/Nemotron-RL-Identity-Following-v1, NeMo Gym dataset) This synthetic dataset uses human created seed data to create a bank of prompts which probes a model's identity information: which company created it, and its…
- nvidia/Nemotron-RL-math-stack_overflow (nvidia/Nemotron-RL-math-stack_overflow, NeMo Gym dataset) The Nemotron-RL-math-stack overflow dataset contains mathematical problems and solutions sourced from the Stack Overflow forums. The method of extracting…
- nvidia/Nemotron-RL-math-advanced_calculations (nvidia/Nemotron-RL-math-advanced_calculations, NeMo Gym dataset) The Nemotron-RL-math-advanced calculations is a dataset designed to test a model's ability to solve complex, multi-step math problems in a multi-step agentic…
- small-ow-agent-bench (catalog) (junjun77/small-ow-agent-bench, Harbor dataset) Frozen compact-shell system reliability for compact open-weight coding agents — not coding IQ.
- general-agent (PrimeIntellect/general-agent, Verifiers environment) Requires verifiers harbor =0.3.1. Multi-turn tool-use tasks from the self-growing general-agent toolbench, where each task ships its own tool world and a gold…
- SkillNet-Gym (zjunlp/SkillNet-Gym, Harbor dataset) SkillNet-Gym is a benchmark of 159 terminal-based tasks for evaluating agents in skill-rich, reproducible environments. This repository uses the native Harbor…
- Multilingual Multimodal RL runs (FineEnvs/multilingual-multimodal-rl-runs, RL dataset) The evidence behind the two Kannada GRPO runs in the Multilingual Multimodal Envs collection: what the model wrote for every held-out item at every…
- Pharma Cold-Chain Inventory Environment (BishnuGanguly/pharma_coldstorage_env, OpenEnv Space)
- Form Filling OpenEnv (tgvarun/form-filling-env, OpenEnv Space)
- Survival Island Game (suraj291/Survival_Island_Game, OpenEnv Space)
- KaggleRL ENV (jiyajahnavi/Kaggle-RL, OpenEnv Space)
- ShinChan Life Simulator (Gladiator-codes/sinchan-env, OpenEnv Space) Shin-chan OpenEnv RL with TRL GRPO training and Gradio UI.
- Terminal-Bench 3.0 (harborframework/terminal-bench-3.0, Harbor dataset)
- Terminal-Bench (harborframework/terminal-bench, Harbor dataset)
- harborframework/terminal-bench-2.0 (harborframework/terminal-bench-2.0, Harbor dataset)
- Terminal-Bench-Science (harborframework/terminal-bench-science, Harbor dataset)
- Spreadsheet-RL Dataset (Spreadsheet-RL/Spreadsheet-RL, RL dataset) This dataset contains the training and evaluation data used by Spreadsheet-RL, a reinforcement learning framework for spreadsheet agents that edit Excel…
- Multi-SWE-bench (PrimeIntellect/Multi-SWE-bench, Verifiers environment) Re-upload of ByteDance's Multi-SWE-bench evaluation benchmark: 2,132 issue-resolving tasks across the seven Multi-SWE languages. This is the held-out eval…
- 🧪 Data Agent — Harbor (eval) (FineEnvs/data-agent-harbor-eval, Harbor dataset) A small, difficulty-balanced validation split — 144 tasks — perfect for quick checkpoints while you train. Same idea as the rest of the family: your agent…
- TaskTrove (open-thoughts/TaskTrove, Harbor dataset) v5.1 (current) — independent-review source retirement — moves 15 sources with majority or unanimous REJECT verdicts out of the default config and into…
- Repo2RLEnv SWE-Flow (FineEnvs/repo2rlenv-swe-flow, Harbor dataset) Trace passing repository tests, select exercised source functions, remove their implementations and generate a reconstruction instruction. Restore the…
- skilltrainbench training tasks (armin-aptura/skilltrainbench-public, Harbor dataset) The training half of the skilltrainbench benchmark suite: for each of the four datasets, the dev task names of its pinned train/test split, in Harbor task…
- Reverse-Text-RL (PrimeIntellect/Reverse-Text-RL, Verifiers environment) A small, scrappy RL dataset used in prime-rl's CI to debug RL training asking a model to reverse small sentences character-by-character. Follows the general…
- Multi-SWE-RL-Verified (PrimeIntellect/Multi-SWE-RL-Verified, Verifiers environment) Gold-patch-validated subset of PrimeIntellect/Multi-SWE-RL-Reupload (ByteDance's Multi-SWE-RL): 2,232 / 4,703 rows across C, Go, Java, JavaScript, Rust, and…
- LHTB Leaderboard — Long-Horizon Terminal-Bench (IntelligenceLab/LHTB-leaderboard, Harbor dataset) This repository hosts submitted runs for Long-Horizon Terminal-Bench (LHTB), a 46-task benchmark measuring how well LLM agents sustain useful work in a…
- 📊 Data Agent — Harbor (train) (FineEnvs/data-agent-harbor-train, Harbor dataset) Teach an agent to actually do data science. This is a suite of 5,000 hands-on data-analysis tasks: each one drops your agent into a sandbox with a real…
- Human Behavior Atlas (HumanBehaviorAtlas/human_behavior_atlas, RL dataset) A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening…
- WildClawBench-Harbor (internlm/WildClawBench-Harbor, Harbor dataset) This repository is the WildClawBench benchmark converted to the Harbor task format, so that all 60 tasks can be run directly with harbor run against any…
- Scale-SWE-Verified (PrimeIntellect/Scale-SWE-Verified, Verifiers environment) Gold-patch-validated fork of AweAI-Team/Scale-SWE (paper): 17,202 / 20,181 Python issue-resolving tasks that produce a clean reward signal end-to-end. Default…
- dabstep-harbor (AdithyaSK/dabstep-harbor, Harbor dataset) A /workdir-native Harbor edition of DABstep. 450 financial-analytics factoid tasks over a single shared Adyen-payments corpus — every task's working…
- AgentTrove (open-thoughts/AgentTrove, Harbor dataset) AgentTrove is the largest open-source collection of agentic interaction traces to date, released by the OpenThoughts-Agent team. It contains 1,696,847 rows…
- Repo2RLEnv CLI-Gym (FineEnvs/repo2rlenv-cli-gym, Harbor dataset) Start with a healthy repository, synthesize a disruption and recovery, verify that damage breaks the original tests and recovery restores them, then export an…
- TerminalWorld Seeds, oracle-validated (andylizf/TerminalWorld-Seeds-Clean, Harbor dataset) The current TW selection retains 405 tasks, excluding any original TW task excluded by either the preserved local quality curation or the upstream v3 training…
- nvidia/Nemotron-RL-agent-workplace_assistant (nvidia/Nemotron-RL-agent-workplace_assistant, NeMo Gym dataset) The Nemotron-RL-agent-workplace assistant is a tool use - multi step agentic environment that tests the agent’s ability to execute tasks in a workplace…
- SWE-Lego-RL-2699 (Lego-X/Lego-RL-2699, Harbor dataset) 2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two parallel views of the same instances:
- Lego-X/LegoFlow-SWE (Lego-X/LegoFlow-SWE, Harbor dataset) LegoFlow-SWE · 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases
- Terminal-Bench Hard (Zhongzhi1228/Terminal-Bench-Hard, Harbor dataset) Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, system…
- FACET-Terminal-Tasks-6k (FACET-Terminal/FACET-Terminal-Tasks-6k, Harbor dataset) 6,020 execution-grounded tasks for terminal agents, coding agents, and executable workflow research
- PatchRecoveryGym for Laguna (poolside-laguna-hackathon/patchrecoverygym-laguna, Verifiers environment) Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track) A reproducible eval + RL environment that tests whether a…
- Repo2RLEnv SWE-smith (FineEnvs/repo2rlenv-swe-smith, Harbor dataset) Apply owned procedural mutations to real repository functions, require an executable test contrast, and author a natural issue from bounded evidence. Contains…
- 🎯 Data Agent — Harbor (test) (FineEnvs/data-agent-harbor-test, Harbor dataset) A held-out benchmark for data-analysis agents: 250 tasks, deliberately balanced across difficulty and leaning toward the harder end so it actually separates…
- AdithyaSK/data_agent_rl_environment_train_multireward (AdithyaSK/data_agent_rl_environment_train_multireward, Harbor dataset) Multi-reward variant of AdithyaSK/data agent rl environment train (2238 tasks, identical data/instructions). The only change: each task's verifier now emits a…
- Repo2RLEnv DataArc terminal synthesis (Envs-FORGE-linked code) (FineEnvs/repo2rlenv-dataarc, Harbor dataset) Read complete Harbor parent tasks and enumerate few-shot, self-instruct and evolution transformations. Author a complete child environment, instruction…
- Watercolour reference pool (FineEnvs/watercolour-reference-pool, RL dataset) The reference paintings that define the reward in the watercolour RL environment: an agent writes a p5.brush sketch, the sketch is rendered, and a vision…
- MiMo-V2.6-RL Terminal (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-terminal, Harbor dataset) Terminal-Bench style tasks. 64 Harbor tasks from the Terminal domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so…
- Repo2RLEnv R2E (FineEnvs/repo2rlenv-r2e, Harbor dataset) Author equivalence tests around repository functions, measure their execution and coverage remotely, and export function-reconstruction tasks with the…
- Keep the failed attempts. Check the artifact. (glayguo/noteflow-research-pilots, Harbor dataset) Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor…
- TerminalWorld Seeds, packaged in the RST release layout (andylizf/TerminalWorld-Seeds, Harbor dataset) 1,530 validated terminal tasks from EuniAI/TerminalWorld, repackaged in the release layout of Zhongzhi1228/Recursive-Task-Synthesis (RST, arXiv:2608.05466)…
- Repo2RLEnv Endless Terminals (FineEnvs/repo2rlenv-endless-terminals, Harbor dataset) Sample native terminal categories, complexity and scenarios; author instructions and private expected behavior, initial-state tests, final tests, fixtures and…
- Repo2RLEnv SWE-gen (FineEnvs/repo2rlenv-swe-gen, Harbor dataset)
- Repo2RLEnv R2E-Gym / SWEGEN (FineEnvs/repo2rlenv-r2e-gym, Harbor dataset) Mine supported bug-fix commits, measure failing-to-passing and regression tests across revisions, and export the buggy starting state and deterministic…
- AACR-Bench Harbor (osolmaz/aacr-bench-harbor, Harbor dataset)
- MiMo-V2.6-RL General (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-general, Harbor dataset) Work in a simulated company through its MCP systems. 925 Harbor tasks from the General domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was…
- MiMo-V2.6-RL Music (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-music, Harbor dataset) Compose a piece in ABC notation. 1,000 Harbor tasks from the Music domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on…
- Watercolour rollouts, judge-led run (FineEnvs/watercolour-rollouts-judge-led, RL dataset) Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
- Repo2RLEnv TerminalWorld (EuniAI) (FineEnvs/repo2rlenv-terminalworld, Harbor dataset) Screen original public terminal recordings, score feasibility and value, extract and refine a solution, write its instruction, reconstruct an offline…
- SkillFlow Test Tasks (zhang-ziao/SkillFlow-Task, Harbor dataset) SkillFlow Test Tasks is the task repository used in the SkillFlow benchmark for evaluating lifelong skill discovery, skill revision, and cross-task procedural…
- nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1 (nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1, NeMo Gym dataset) The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The…
- 📊 SmolDataEnvs: Harbor (test) (FineEnvs/SmolDataEnvs-harbor-test, Harbor dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- repo2rlenv-commit-runtime-v2 (FineEnvs/repo2rlenv-commit-runtime, Harbor dataset) Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments.
- rounakbende/rh-swe-bench (rounakbende/rh-swe-bench, Harbor dataset) A Red Hat-specific benchmark for evaluating coding agents, packaged in Harbor format. Contains 357 quality-verified tasks across 25 Red Hat ecosystem…
- Human Behavior Atlas (DennisDengHUst/human_behavior_atlas, RL dataset) A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening…
- jupyter-agent-eval-v1 — Harbor task suite (AdithyaSK/jupyter-agent-eval-v1-harbor, Harbor dataset) 100 Harbor task(s) for the Jupyter data-analysis agent. Each task is one (question, gold answer) pair against a real Kaggle dataset.
- repo2rlenv-pr-runtime (FineEnvs/repo2rlenv-pr-runtime, Harbor dataset) Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments.
- Monthly-SWEBench 2026-05 (UnipatAI/Monthly-SWEBench-2026-05, Harbor dataset) This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRs…
- 📊 SmolDataEnvs: Harbor (eval) (FineEnvs/SmolDataEnvs-harbor-eval, Harbor dataset) 5.5K+ RL tasks for hill-climbing small models in code and data science.
- Repo2RLEnv TMax (FineEnvs/repo2rlenv-tmax, Harbor dataset) Sample the native TMax legacy taxonomy, create a template, initial-state tests and final completion tests in distinct stages, then build fixtures and a…
- Repo2RLEnv SETA Seed2Synth (FineEnvs/repo2rlenv-seta-seed2synth, Harbor dataset) Use attributed Unix Stack Exchange questions and accepted answers to design terminal tasks, then author self-contained offline environments, instructions…
- Monthly-SWEBench-2026-04 (UnipatAI/Monthly-SWEBench-2026-04, Harbor dataset) 90 tasks — 43 bugfix + 47 non-bugfix Tasks span diverse open-source repositories Each task includes a runnable environment, test suite, and reference solution
- Repo2RLEnv SCALER (FineEnvs/repo2rlenv-scaler, Harbor dataset) Generate instances from owned reasoning families with seeded procedural samplers, deterministic references and family-specific answer verification; generation…
- Repo2RLEnv SWE-Next (FineEnvs/repo2rlenv-swe-next, Harbor dataset) Mine real repository history for supported code changes and associated tests, preserve regression behavior, and export the historical implementation change as…
- DeepScaleR-Preview VERL (sungyub/deepscaler-preview-verl, RL dataset) This dataset contains 35,789 mathematical reasoning problems in VERL format, processed from agentica-org/DeepScaleR-Preview-Dataset. Key Features:
- Monthly-SWEBench-2026-03 (UnipatAI/Monthly-SWEBench-2026-03, Harbor dataset) 112 tasks — 68 bugfix + 44 non-bugfix Tasks span diverse open-source repositories Each task includes a runnable environment, test suite, and reference solution
- LoLBench (lolbench26/LoLBench, Harbor dataset) Software development often begins with user intent and high-level design rather than an implementation-ready specification. Developers must ground those goals…
- MiMo-V2.6-RL Webdev (Harbor) (FineEnvs/MiMo-V2.6-RL-harbor-webdev, Harbor dataset) Build a website from a design brief. 2,093 Harbor tasks from the Webdev domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on…
- introvoyz041/terminal-bench-2.0 (introvoyz041/terminal-bench-2.0, Harbor dataset)
- repo2rlenv-equivalence-tests (FineEnvs/repo2rlenv-equivalence-tests, Harbor dataset) Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments.
- RTS-1K: offline rootless Apptainer release (laion/rts-1k-rootless-apptainer, Harbor dataset) This derived release contains 984 tasks from Recursive Task Synthesis. RTS-1K is the stable project label; the exact size is 984. The tasks are the strict…
- Nemotron Terminal easy clean-1K: rootless Apptainer release (laion/nemotron-terminal-easy-clean1k-rootless-apptainer, Harbor dataset) This derived release contains exactly 1,000 balanced easy tasks from four Nemotron Terminal domains. It is pinned to…