vLLM and PagedAttention: Virtual Memory for LLM Serving
Efficient Memory Management for Large Language Model Serving with PagedAttention
Episode participants
A solo episode without invited guests.
What we will explore
A scheduled Research Insights Made Simple episode walks through the SOSP '23 paper by Kwon, Li, Zhuang and co-authors, «Efficient Memory Management for Large Language Model Serving with PagedAttention»: why large language model serving is bound by KV cache memory and how existing systems lost 60-80 % of it to reservation and fragmentation.
The conversation follows the paper's logic: the cost of one token in the KV cache, three kinds of waste in contiguous placement, the idea of paged virtual memory from operating systems, the PagedAttention kernel, block tables, copy on write for parallel sampling and beam search, a scheduler with all-or-nothing eviction, and the architecture of the vLLM engine.
The central question is how a sixty-year-old borrowed idea delivered 2-4× throughput without a single model change, what the kernel's indirection cost, and where, in the authors' own words, the recipe does not apply.