Implement a shape-aware batching scheduler for static-graph LLM inference that optimally packs requests into…
Implement a shape-aware batching scheduler for static-graph LLM inference that optimally packs requests into…: a task in Terminal-Bench 2.1 (Harbor git-repos dataset) (Harbor dataset). In this task you will implement an LLM inference batching scheduler (shape‑aware) for a static graph LLM…
The task
In this task you will implement an **LLM inference batching scheduler (shape‑aware)** for a static graph LLM inference system. **Background ** When running large language models on hardware accelerators (such as TPUs, RDUs, or bespoke inference chips) the compiled execution graph has to operate on fixed‑sized…
Part of qingyu-lyq/terminal-bench-2.1.