evaluation mode num entries avg tokens gen seconds accuracy
evaluation mode num entries avg tokens gen seconds accuracy: a task in LegoFlow-SWE (Harbor dataset). Problem: Our evaluation framework does not yet support the ComputeEval benchmark, a suite of 406+ CUDA code generation challenges that assess functional correctness through compilation and test…
The task
**Problem:** Our evaluation framework does not yet support the ComputeEval benchmark, a suite of 406+ CUDA code generation challenges that assess functional correctness through compilation and test execution. Users need the ability to evaluate models on CUDA programming tasks and obtain pass@k metrics.
Part of Lego-X/LegoFlow-SWE.