PinchBench Evaluation Harness Scores Far Below Expected Due to Multiple Infrastructure Bugs
PinchBench Evaluation Harness Scores Far Below Expected Due to Multiple Infrastructure Bugs: a task in LegoFlow-SWE (Harbor dataset). When running the PinchBench evaluation harness, we observe scores in the 5–26% range, whereas expected performance should be 80–90%. Tracked down to several issues…
The task
When running the PinchBench evaluation harness, we observe scores in the 5–26% range, whereas expected performance should be 80–90%. Tracked down to several issues in the evaluation pipeline that prevent automated graders from correctly inspecting agent behavior and workspace artifacts.
Part of Lego-X/LegoFlow-SWE.