HF RL Explorer

Judge 10 agent-written citations verbatim no match by fetching their sources and applying SKBench's…

Judge 10 agent-written citations verbatim no match by fetching their sources and applying SKBench's…: a task in skbench-env (Harbor dataset). What the agent is asked, how it is graded, what it runs in, and its files.

Part of seekbot/skbench-env.