This article was originally published on my blog. For the latest version and future updates, please visit the original post: https://jaketao.com/language/en/kv-cache-vs-prompt-cache/ . Every time a large language model generates a token, it draws on the content that came before it.