IntelligenceLab/Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB): Harbor dataset on Hugging Face with 46 tasks. LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and…
Tasks
- Play the game of 2048 turn-by-turn through a terminal game server, merging tiles to build the largest tile possible. Scored in bands by…
- Reproduce a calibrated ALP floating-point compression experiment from the paper method and sample columns.
- Act as the analyst on an investment bank's deal team advising a private-equity sponsor on taking the EdTech company KSchool private, and…
- Act as the analyst on an investment bank's deal team advising Advent International on taking Planet Fitness (PLNT) private, and work the…
- Act as outside counsel for a law firm juggling four client matters and work them through 70 dependent stages over three realistic…
- Act as the analyst on a management-consulting team across two client matters, worked through 33 dependent stages built on two APEX worlds…
- Close RTL-to-GDSII signoff on the lowRISC ibex core CPU for the SkyWater sky130hd PDK using OpenROAD-flow-scripts, ported from the…
- Repair a multimodal audio-video synchronization audit pipeline so it detects visual flashes/motion events, audio onsets, dropped-frame…
- Play a full game of chess from the standard opening, turn-by-turn through a terminal game server, as White against a bundled superhuman…
- Repair a climate-science extreme-event audit pipeline so it correctly computes gridded heatwave, precipitation, drought, anomaly, and…
- Commit0-style TEST-DRIVEN long-horizon coding task: reimplement THREE whole Python libraries (tinydb 4.8.2, sortedcontainers 2.4.0…
- Audit synthetic DICOM CT series geometry, HU conversion, windowing, masks, and lesion measurements.
- Repair a multimodal scanned-document table reconstruction pipeline so it detects table grids/cells from PNG page images, assigns noisy OCR…
- Iteratively patch DuckDB's query optimizer to improve TPC-H performance while preserving correctness.
- Audit EPA SWMM stormwater hydraulic and hydrologic simulations from official solver outputs.
- Fit and control a multi-city age-structured SEIR-H-ICU epidemic model.
- Reproduce a calibrated Foldseek-style structural search experiment from the Nature Biotechnology paper.
- Reproduce GDAL/PROJ raster reprojection, coordinate transform, point sampling, and zonal-statistics regression checks.
- Develop a bot for a Generals.io-style grid strategy game by iteratively playing matches and refining strategy.
- Iteratively develop a grammar-guided fuzzer to maximize code coverage of target parsers.
- Build a Great Expectations based multi-table data quality and reconciliation audit pipeline.
- Migrate a real legacy LangChain support RAG + router app across a full API break with hard-fail compatibility gates.
- Repair a materials-science phase-diagram audit pipeline so it correctly reconstructs alloy phase boundaries, invariant reactions…
- Reproduce MATPOWER/PYPOWER AC/DC power flow and OPF regression checks with detailed grid constraint diagnostics.
- Repair a multimodal microscopy cell-count QC audit so it segments synthetic nuclei/cell PNG imagery, detects focus/artifact failures, and…
- Audit USGS MODFLOW 6 groundwater simulations from official solver outputs.
- Iteratively optimize an N-body gravitational simulation from O(N^2) to O(N log N) with correctness verification at multiple scales.
- Audit renewable generation and storage dispatch using NREL PySAM official PVWatts outputs.
- Audit seismic structural dynamics using official OpenSeesPy finite-element solver outputs.
- Generate proof-of-concept inputs triggering specific vulnerability classes using sanitizer feedback.
- Bring up and debug the twosigma/frost RV32GCB SystemVerilog CPU (an out-of-order, Tomasulo-style core): the RTL under hw/rtl has been…
- Repair a robotics SLAM benchmark audit pipeline so it correctly composes SE(2) poses, evaluates loop closures, reconstructs landmark maps…
- Solve a four-stage 6x6 Rush Hour campaign by manual board play, submitting legal move routes under length caps.
- Repair a multimodal satellite flood change-detection audit so it compares before/after imagery, masks clouds/no-data, segments new water…
- Repair a multimodal scientific figure digitization pipeline so it reconstructs plotted series data from figure images, captions, and…
- Play a deterministic hard Snake variant through a terminal game server: eat seeded apples, survive a growing sequence of seeded obstacles…
- Solve a ramp of Sokoban puzzles turn-by-turn through a terminal game server, advancing through as many levels as possible. Scored in bands…
- Reproduce NASA NAIF SPICE ephemeris, angle, and ground-station visibility regression checks from pinned kernels.
- Implement and iteratively refine a cloud spot-instance scheduling policy to minimize cost while meeting hard deadlines.
- Reproduce a hardened SU2 official aerodynamic regression matrix from official TestCase references and solver history artifacts.
- Recover 7 progressively corrupted, genuinely hard variant-Sudoku puzzles (50-54 near-minimal blank cells) turn-by-turn through a terminal…
- Play Super Mario Bros turn-by-turn through a terminal game server, advancing through as many levels as possible. Scored in bands (0-3) by…
- Sparse inverse-problem marathon with covariate shift on the underlying linear operator. Recover a 200-dim sparse signal X from 50-dim…
- Reproduce a calibrated UNISON fat-tree MTP experiment from the paper method and settings.
- Recover the TRUE semantics of an undocumented, legacy-friendly job-config format and repair a buggy config-normalisation engine so it…
- Build an approximate nearest-neighbor vector search service from scratch, iteratively optimizing recall and throughput.
IntelligenceLab/Long-Horizon-Terminal-Bench on the Hugging Face Hub