HF RL Explorer

What is the difference in macro-averaged F1-score between the top-performing and bottom-performing models?

What is the difference in macro-averaged F1-score between the top-performing and bottom-performing models?: a task in data-agent-harbor-train (Harbor dataset). Files (in /home/user/input, no subfolders): - (see /home/user/input)

The task

Files (in /home/user/input, no subfolders): - (see /home/user/input)

Part of FineEnvs/data-agent-harbor-train.