Issue: Benchmark results are not versioned, causing misleading comparisons after benchmark definition changes
Issue: Benchmark results are not versioned, causing misleading comparisons after benchmark definition changes: a task in LegoFlow-SWE (Harbor dataset). When a benchmark's source code is modified (e.g., the benchmark function itself is changed), old results stored in the results database become…
The task
When a benchmark's source code is modified (e.g., the benchmark function itself is changed), old results stored in the results database become incompatible with the new definition. Currently, the system still compares these old results against new ones, producing performance comparisons that do not reflect the same…
Part of Lego-X/LegoFlow-SWE.