Engineering Insights

Semantic Caching: Cut LLM Cost Without Making Answers Stale

July 15, 2026 Umut Kaya

Semantic caching stores a previous model response and returns it when a new prompt is close enough in meaning, not only when the string matches. Exact-match caches miss “How do I reset my password?” versus “I forgot my password.” Those two should not cost you two full generations.

We added it after a week of Polylingo support traffic where half the questions were the same five intents in different wording. Token spend dropped without touching the model. The failure mode is also obvious: cache a wrong answer and you serve it faster.

Similarity is a product knob

We embed the user turn (and only the user turn — not the whole system prompt) and look up a vector index keyed by product, locale, and feature. Thresholds are per feature. Password-reset copy can be aggressive. A medical-adjacent explanation, if we ever shipped one, would be off. A lesson that depends on the learner’s last mistake must include that mistake in the cache key or you will “help” the wrong child.

Invalidation is the real work

When the knowledge base changes, we bump a generation id on that source and drop cache entries that cited it. When we change a prompt template, we bump another. TTL alone is not enough — a 24-hour TTL on a pricing answer the day you change plans is a support ticket factory. We treat cache entries like HTML: they have an etag made of template + index generation + locale.

Where it lives in our .NET stack

A small service in front of the LLM client: Redis for the payload, a compact vector store for the lookup. The ASP.NET handler never talks to the model directly. That makes it testable — we replay a week of anonymised prompts and assert hit rate and staleness. If hit rate is under 15% for a feature, we turn the cache off rather than pretend.

Questions we keep getting

Is this the same as RAG? No. RAG fetches source documents. Semantic cache reuses a finished answer. We often run both: RAG first, then cache the grounded reply.

What about personal data? We never cache prompts that contain emails, names, or learner free text that could identify a child. The classifier that decides “safe to cache” is itself a small model with a deny-by-default list.

How much did you save? On Polylingo support, enough to pay for the Redis box several times over. On authoring tools with unique prompts, almost nothing — and we left it off.