HF RL Explorer

evaluation mode num entries avg tokens gen seconds accuracy

evaluation mode num entries avg tokens gen seconds accuracy: a task in LegoFlow-SWE (Harbor dataset). Problem: Our evaluation framework does not yet support the ComputeEval benchmark, a suite of 406+ CUDA code generation challenges that assess functional correctness through compilation and test…

The task

**Problem:** Our evaluation framework does not yet support the ComputeEval benchmark, a suite of 406+ CUDA code generation challenges that assess functional correctness through compilation and test execution. Users need the ability to evaluate models on CUDA programming tasks and obtain pass@k metrics.

Part of Lego-X/LegoFlow-SWE.