HF RL Explorer

PinchBench Evaluation Harness Scores Far Below Expected Due to Multiple Infrastructure Bugs

PinchBench Evaluation Harness Scores Far Below Expected Due to Multiple Infrastructure Bugs: a task in LegoFlow-SWE (Harbor dataset). When running the PinchBench evaluation harness, we observe scores in the 5–26% range, whereas expected performance should be 80–90%. Tracked down to several issues…

The task

When running the PinchBench evaluation harness, we observe scores in the 5–26% range, whereas expected performance should be 80–90%. Tracked down to several issues in the evaluation pipeline that prevent automated graders from correctly inspecting agent behavior and workspace artifacts.

Part of Lego-X/LegoFlow-SWE.