Suffix Cache Reuse Explained
Deep Dive in Efficient Serving for Context Language Models
September 28, 2026
“How to stop KV cache from crying
when they’re no longer in an append-only relationship with the
context” 😉
Before publishing: confirm author formatting (the TMax post bolds co-first authors; the paper marks none), the publication date, and the missing links in the resources line (arXiv, tweet). Draft notes like this one only show up in local builds; production builds hide them.
Resources: 📄 Paper · 👨💻 GitHub · 🐦 Tweet
We recently introduced Context Language Models (CLMs),Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin,
Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan
Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, and Pang
Wei Koh. “Context Language
Models.” arXiv preprint arXiv:2609.37725, 2026.
which
treat context as a file and can perform arbitrary manipulations on it.
We showed that CLMs outperform state-of-the-art, human-designed
context-management harnesses at lower cost. In this blog, we dive deeper
into the efficiency side of CLMs: what metric do we use to capture the
realistic serving cost while being cache aware, and how could we further
improve the cache hit rate by designing serving systems for agents?
Prefix-Reuse FLOPs
Background: existing cache reuse often assumes append-only context. Serving engines such as SGLang and vLLM cache KV states, for example in a radix tree, and reuse the longest matching prefix of a new request. Tokens after the first mismatch must be re-prefilled. This works naturally for append-only histories, but after an in-the-middle edit, even unchanged suffix tokens are recomputed.
Below is a simplified example: the context A B C is edited into A B′ C, with the colors matching the token strip. The prefix A still matches, so its cache is reused; the match breaks at B′, so B′ and the unchanged C after it are prefilled together.
this is sentence A A this is sentence B B this is sentence C C
this is sentence A A · hit prefix cache compacted B B′ · prefilled this is sentence C C · re-prefilled
When measuring CLM efficiency, we account for prefix-cache reuse, i.e., which tokens can reuse cached KV states. We capture this with a theoretical inference-cost metric we call Prefix-Reuse FLOPs:
\mathrm{FLOPs}_{\text{Prefix Reuse}} = \underbrace{\mathrm{FLOPs}_{\text{prefill}}\big(\text{unmatched suffix}\big)}_{\text{from the first prefix mismatch onward}} + \underbrace{\mathrm{FLOPs}_{\text{decode}}\big(\text{generated tokens}\big)}_{\text{new output tokens}}
By default, we use this metric to measure CLM efficiency under
standard serving. Thus, when we say CLMs are
cheaper than the baselines, we’ve accounted for the lower cache hit rate
caused by context edits.For example, in these performance-efficiency Pareto
plots, we used prefix-reuse FLOPs with standard serving. 
The visualization below shows how an edit affects the
Prefix-Reuse FLOPs of a single turn.
Prefix Reuse FLOPs of one Qwen3.6-27B turn under standard serving. Move the edit or change its size; the bar splits the Prefix Reuse FLOPs of the turn into the prefill of B′, the re-prefill of the unchanged C, and decoding.
The hit rate does drop with in-the-middle edits and incurs
re-prefilling of the unchanged C.This also happens beyond CLM editing: some chat
endpoints remove the thinking tokens of earlier turns, so the tokens
after them are re-prefilled in the next turn.
In our experiments with a Qwen3.6-27B CLM on
BrowseComp-Plus, standard SGLang serves 72.9% of all prompt tokens from
its prefix cache, but only 24.2% on the turns right after a context
edit. The rest of an edited turn is prefilled again, including the large
part of the context that the edit left unchanged.
Suffix Cache Reuse
Prefix cache reuse has been the tradition in serving engines because context has always been append-only. But we ask: can we adapt serving engines for AI’s convenience? To that end, we propose a simple yet effective method, Suffix Cache Reuse, to make CLM serving even more efficient on the system side.
this is sentence A A this is sentence B B this is sentence C C
this is sentence A A · hit prefix cache compacted B B′ · prefilled this is sentence C C · re-prefilled
this is sentence A A · hit prefix cache compacted B B′ · prefilled this is sentence C C · suffix cache reused
Standard serving versus Suffix Cache Reuse after an edit replaces B with B′. Standard serving hits the prefix cache for A but must prefill B′ and re-prefill all of C; Suffix Cache Reuse also reuses the cached states of C. The × marks the first prefix-mismatch position.
Say the context is A B C, and an edit replaces B with B′. Standard serving reuses the cache for A and then stops, because a prefix cache only matches up to the first changed token: B′ and all of C are prefilled again, even though C did not change. Suffix Cache Reuse (SCR) keeps the cached states for C instead of throwing them away. When the next prompt arrives, SCR compares it with the previous prompt of the same session to find the spans that survived the edit, shifts their rotary position encodings to their new positions, and splices them in after B′. Only B′ and newly appended tokens are prefilled. Below is an interactive visualization of SCR serving compared with prefix-reuse-only serving.
Adding Suffix Cache Reuse on top of prefix caching. The same turn as in the first figure; the bars compare prefix caching with and without SCR, and the dashed green outline is the re-prefill of C that SCR removes. SCR is counted with B′ prefilled through every layer and C relocated as one span.
Technical details. For full-attention layers, SCR
reuses the cached keys and values of C as they are, except for position: after
the edit, C sits earlier by the
length difference between B and
B′, so SCR re-rotates the rotary
position encodings of its cached keys by that offset. Because rotary
encodings depend only on position, this rotation is exact and far
cheaper than recomputing the entries. SCR is an approximation: C’s cached states were computed under the
old context, before the edit. To bound how much approximation one edit
can introduce, SCR relocates at most six surviving spans per edit, the
longest first, and re-prefills the rest.More details, such as how the cap on relocated spans is
set and what effect it has, are in Appendix B of the paper; the main
text uses k = 6.
Qwen3.6-27B is a hybrid model: 48 of its 64 layers use linear attention, which keeps a fixed-size recurrent state rather than a per-token cache, so there are no per-token entries to move. For those layers, SCR continues from a snapshot of the recurrent state taken before the edit. The edit itself is seen by the 16 full-attention layers, and through their outputs it still reaches the later linear-attention layers.
Results on BrowseComp-Plus
We perform an end-to-end evaluation of SCR on BrowseComp-Plus with
the CLM
harness, by simply switching the Qwen3.6-27B endpoint from standard
SGLang to our patched SGLang with SCR (Section 5.3 of the paper).Our implementation of Suffix Cache Reuse is available
at facebookresearch/context-language-models/suffix_cache_reuse.
Accuracy is identical, 60.2% both ways, while compute drops from 10.98 to 7.14 PFLOPs per question. On the turns right after an edit, SCR serves an extra 28.2% of the prompt from relocated cache that standard serving would have recomputed. CLMs were already cheaper than the baselines under standard serving, through better context management alone; SCR brings their serving cost down to 65% of that.
Bonus: reasoning-token stripping is a context edit too. Other cache misses arise from reasoning-token stripping in some chat-template serving. For models such as Qwen3.6, earlier reasoning blocks are removed from subsequent prompts, forcing the unchanged suffix to be prefilled again. SCR reuses this suffix cache as well. In fact, most of SCR’s extra reuse comes from stripped reasoning rather than CLM edits. The figure below shows a decomposition of the tokens hit by Suffix Cache Reuse.
Improvement space for cache misses. We also looked at the unchanged tokens that are still prefilled with SCR. SCR removes most re-prefilling of unchanged suffixes after an edit. Much of the remaining overhead comes from unchanged prefixes that standard SGLang fails to reuse efficiently for hybrid models, because linear-attention states are cached only at request boundaries. This leaves substantial room for improvement: finer-grained recurrent-state checkpoints could increase cache hit rates for both standard prefix caching and SCR. More broadly, serving models with editable context opens many new systems research directions.
Related work
With standard prefix caching, changing an early part of a prompt forces the serving system to recompute the KV states of everything that follows, even when the later text is unchanged. Prior work relaxes this requirement in different settings. Prompt Cache precomputes attention states for predefined prompt modules, allowing a module to be reused in prompts that do not share the same preceding text. In retrieval-augmented generation, the same document may appear after different documents or instructions. CacheBlend and EPIC reuse cached document chunks in these new contexts, recomputing selected tokens to account for the changed surroundings.
PIE studies cache reuse when a user modifies previously processed code and requests a new completion. It retains cached states for unchanged text after an edit and corrects their rotary positions, avoiding suffix recomputation. Memento evicts each completed reasoning block from the KV cache but keeps the cached states of its summary, which were computed while the block was still in context, and finds that these states retain useful information from the evicted block. Concurrently, KV-streams, released in the last few days, keeps the KV cache across agentic compaction during reinforcement learning instead of flushing it, and reports a 2× speed-up on SWE tasks with any compaction method. Suffix Cache Reuse applies the same reuse principle to an agent’s live context: when the agent replaces a span, the unchanged suffix retains its cached states rather than being prefilled again. We integrate this mechanism into SGLang for agent-driven context editing and further extend it to hybrid architectures that combine full-attention layers with linear-attention layers.
We believe cache space can enable more context management opportunities than pure token space. We are excited to see more work along this line.
Limitations
Suffix Cache Reuse makes more approximations than prefix cache reuse. A reused prefix is exact, since its states depend only on tokens that have not changed. A relocated suffix is not. Its cached states were computed while the old span was still in context, so they may still reflect text the edit removed and miss text it added. The approximation is larger for hybrid models, whose linear-attention layers resume from a recurrent state taken before the edit and see the new span only through the full-attention layers. We cap the number of relocated spans per edit to limit this, and on BrowseComp-Plus SCR matches standard serving in accuracy. We have not yet characterized where these approximations break down. More work is needed to profile their failure cases and to make reuse adaptive, deciding for each edit what to reuse and what to recompute.
References
- Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, and Pang Wei Koh. Context Language Models. arXiv:2609.37725, 2026.
- Lianmin Zheng et al. SGLang: Efficient Execution of Structured Language Model Programs. arXiv:2312.07104, 2023.
- Woosuk Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM). SOSP 2023.
- Radix tree. Wikipedia.
- In Gim et al. Prompt Cache: Modular Attention Reuse for Low-Latency Inference. arXiv:2311.04934, 2023.
- Jiayi Yao et al. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion. arXiv:2405.16444, 2024.
- Junhao Hu et al. EPIC: Efficient Position-Independent Caching for Serving Large Language Models. arXiv:2410.15332, 2024.
- Zhenyu He et al. Let the Code LLM Edit Itself When You Edit the Code (PIE). arXiv:2407.03157, 2024.
- Vasilis Kontonis et al. MEMENTO: Teaching LLMs to Manage Their Own Context. arXiv:2604.09852, 2026.
- Emiliano Penaloza et al. KV-streams for Efficient Compaction in Agentic Reinforcement Learning. arXiv:2609.35750, 2026.
Citation
Please cite this work as:
@article{shao2026context,
title = {Context Language Models},
author = {Shao, Rulin and Shen, Shannon Zejiang and Yin, Junjie Oscar and
Li, Yuetai and Wang, Minheng and Ivison, Hamish and
Poovendran, Radha and Lambert, Nathan and Xiao, Teng and
Lewis, Mike and Yih, Wen-tau and Zettlemoyer, Luke and
Koh, Pang Wei},
journal = {arXiv preprint arXiv:2609.37725},
year = {2026}
}
Add the tweet link to the resources line once it is posted.