HF RL Explorer

internlm/WildClawBench-Harbor

WildClawBench-Harbor: Harbor dataset on Hugging Face with 60 tasks. This repository is the WildClawBench benchmark converted to the Harbor task format, so that all 60 tasks can be run directly with harbor run against any Harbor-supported agent (Claude Code, OpenHands, Codex CLI, custom agents…

Tasks

internlm/WildClawBench-Harbor on the Hugging Face Hub