Tests the ability to find a probability distribution satisfying precise dual KL divergence constraints…
Tests the ability to find a probability distribution satisfying precise dual KL divergence constraints…: a task in Terminal-Bench 2.1 (Harbor git-repos dataset) (Harbor dataset). The confidence of a token prediction in an LLM can be quantified using different metrics. This implementation focuses…
Part of harborframework/terminal-bench-2.1.