← All publications

2024 · Article

Prompt Cache: Modular attention reuse for low-latency inference

In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal and Lin Zhong

MLSys, 2024

Abstract

Prompt Cache accelerates large-language-model inference by reusing attention states across prompts. Many prompts share text such as system messages, templates, and context documents. By precomputing and storing attention states for these recurring segments, an inference server can reuse them when they appear in later prompts. Prompt Cache provides a schema for defining reusable prompt modules, ensuring positional accuracy and giving users an interface to access cached states. Across several models, the prototype reduces time-to-first-token latency by up to 8 times on GPU inference and 60 times on CPU inference while preserving output accuracy and requiring no model-parameter changes.

Publication details

Venue
MLSys
Publication year
2024

BibTeX

@article{promptcache,
  title = {{Prompt Cache: Modular attention reuse for low-latency inference}},
  author = {Gim, In and Chen, Guojun and Lee, Seung-seob and Sarda, Nikhil and Khandelwal, Anurag and Zhong, Lin},
  journal = {MLSys},
  month = may,
  year = {2024}
}