Tests the ability to find a probability distribution satisfying precise dual KL divergence constraints…
Tests the ability to find a probability distribution satisfying precise dual KL divergence constraints…: a task in Terminal-Bench 2.1 (Harbor git-repos dataset) (Harbor dataset). The confidence of a token prediction in an LLM can be quantified using different metrics. This implementation focuses…
The task
The confidence of a token prediction in an LLM can be quantified using different metrics. This implementation focuses on two metrics based on KL divergence from the uniform distribution:
Part of qingyu-lyq/terminal-bench-2.1.