HF RL Explorer

Tests the ability to find a probability distribution satisfying precise dual KL divergence constraints…

Tests the ability to find a probability distribution satisfying precise dual KL divergence constraints…: a task in Terminal-Bench 2.1 (Harbor git-repos dataset) (Harbor dataset). The confidence of a token prediction in an LLM can be quantified using different metrics. This implementation focuses…

The task

The confidence of a token prediction in an LLM can be quantified using different metrics. This implementation focuses on two metrics based on KL divergence from the uniform distribution:

Part of qingyu-lyq/terminal-bench-2.1.