HF RL Explorer

PrimeIntellect/Multi-SWE-bench

Multi-SWE-bench: Verifiers environment on Hugging Face. Re-upload of ByteDance's Multi-SWE-bench evaluation benchmark: 2,132 issue-resolving tasks across the seven Multi-SWE languages. This is the held-out eval benchmark; for RL training data use PrimeIntellect/Multi-SWE-RL-Verified.

PrimeIntellect/Multi-SWE-bench on the Hugging Face Hub