Reusing the Prompt Prefix with a Key-Value Cache for SLM Optimization
A practical walk-through showing how to avoid recomputing static prompt tokens in small language models by caching per-layer key and value vectors for a fixed prefix.

A practical walk-through showing how to avoid recomputing static prompt tokens in small language models by caching per-layer key and value vectors for a fixed prefix.

Set up a reproducible token split: the prefix/suffix boundary must be token-clean so cached keys align with full prompt encoding.
Using a constrained scoring approach (comparing first-token logits of distinct labels) makes per-item processing a single forward pass after key-value prefill.
The article benchmarks with Qwen2.5-0.5B-Instruct (float16) on an M2 MacBook Air (24GB RAM, 16-core Neural Engine) and includes concrete code snippets for tokenization and prompt construction.
# Why cache the prompt prefix Narrow automation prompts tend to be mostly static: an instruction block, a taxonomy, and a few examples. Only a short tail changes per item. Transformers build a key and a value vector for each token at every layer, and those vectors depend only on tokens to the left. If the prefix is fixed across calls, those vectors are identical each time. Computing them once and reusing them reduces per-item work to the few tokens that actually vary.
# Baseline and environment The article uses Qwen2.5-0.5B-Instruct in float16 via Hugging Face Transformers. Benchmarks run on an M2 MacBook Air with 24GB RAM and a 16-core Neural Engine. The required Python packages shown are torch, transformers, and accelerate (pip install torch transformers accelerate). A short, realistic few-shot prompt is re-encoded for every ticket in the baseline.
# Constrained scoring to minimize work To keep the evaluation cheap, the author carries forward a constrained scoring trick introduced earlier: if label first tokens are distinct, comparing logits for each label's first token is enough to choose the label. That turns the decision into a single forward pass after the prompt prefill is handled.
# Token-clean prefix/suffix split The chat prompt uses a ChatML-style layout written out by hand so the script can split the prompt at a known boundary. The split must be token-clean: encoding the two halves separately must give the same token ids as encoding the whole prompt in one go. If that check fails, cached keys will not align with the model's expectations.
# Concrete elements included in the walkthrough
# What caching changes in practice Instead of re-encoding and recomputing keys/values for every token of the entire prompt for each ticket, cache the per-layer key and value tensors for the fixed prefix. For each new ticket you only encode the ticket-specific tail, run the smaller per-item prefill and the constrained scoring forward pass, and pick the label by comparing logits for the stored label-first-token ids.
# Implementation details to watch for
# Practical outcome When your automation prompt is largely static and the per-item change is a short tail, caching prefix key/value tensors cuts the repeated computation substantially. The article illustrates this with code and a working validation that the prefix/suffix split is token-clean, using Qwen2.5-0.5B-Instruct as the reference model.
# Next steps suggested by the series

Most people cache the whole prompt, which catches almost nothing because real prompts carry timestamps and request IDs. The fix is four lines, and the clever upgrade after it is a trap.
As Large Language Models (LLMs) become increasingly integrated into enterprise applications, optimizing response time and reducing operational costs have become critical priorities. One of the most effective techniques for achieving both is Prompt Caching. Instead of processing identical prompt segments repeatedly, pro
This article was originally published on my blog. For the latest version and future updates, please visit the original post: https://jaketao.com/language/en/kv-cache-vs-prompt-cache/ . Every time a large language model generates a token, it draws on the content that came before it. If it had to compute everything from

Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integ

This article covers five prompt optimization strategies such as: prompt optimization, prompt engineering, LLM output quality, few-shot prompting, chain-of-thought, structured outputs.
I built a prompt dependency graph that separates everything a component can reach from the smaller set that actually needs targeted evaluation. The post Changing One Prompt Can Affect 50 Others — I Built a Prompt Dependency Graph to Find What Needs Retesting appeared first on Towards Data Science .
Loading more related stories...
Open the app view to save this story, compare related coverage, and continue from the same source.