HF RL Explorer

IntelligenceLab/Long-Horizon-Terminal-Bench

Long-Horizon Terminal-Bench (LHTB): Harbor dataset on Hugging Face with 46 tasks. LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and…

Tasks

IntelligenceLab/Long-Horizon-Terminal-Bench on the Hugging Face Hub