For whom the door-bell tolls In a previous post we extolled the benefits of KV caching, a technique to save the KV states from the prefill step of LLM-based inference to reduce time to first token (TTFT) and skip redundant computation. I co-presented this with Tushar Gohad at Cephalocon . Since then I’ve been thinking a lot about how to improve the state of the art.