SGLang: Executing Structured Language Model Programs
SGLang: Efficient Execution of Structured Language Model Programs
Episode participants
A solo episode without invited guests.
What we will explore
A scheduled Research Insights Made Simple episode walks through the NeurIPS 2024 paper by Zheng, Yin, Xie, Sheng and co-authors from Stanford and Berkeley, «SGLang: Efficient Execution of Structured Language Model Programs»: why modern applications call the model many times with branching and structured input and output while inference engines see every call in isolation.
The conversation follows the paper's logic: a seven-primitive language embedded in Python, the interpreter as a stream of asynchronous operations, RadixAttention as a radix tree with LRU eviction on top of a paged KV cache, longest-shared-prefix-first scheduling, a compressed finite state machine for schema-constrained decoding and speculative execution for closed APIs.
The central question is how a server that can read the structure of a program got up to 6.4× throughput and up to 3.7× lower latency over vLLM without touching the model, what the tree itself costs, and where, in the authors' own words, the recipe falls short.