Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Semantic Caching

By Anup Rai7 min readReviewed September 2026

Semantic caching reuses a previously computed result when a new request is judged equivalent enough for that result to remain valid. It usually uses embeddings to find candidates, then applies additional eligibility and validation checks. Similar wording or meaning is evidence for a match, not proof that two requests need the same answer.

Remember: Retrieval asks “is this relevant evidence?” Response caching asks the stricter question “may I return this result for this request?”

Distinguish the three kinds of reuse

Mechanism What is reused Match condition Main risk
Exact response cache A saved output Same correctly constructed request key Missing dependency or stale data
Semantic response cache Output for a related request Accepted equivalence under the product's policy False match in addition to staleness
Prefix/KV cache Intermediate input computation Eligible matching prefix and runtime conditions Incorrect assumptions about eligibility, lifetime or cost

An exact cache is not risk-free. Hashing only “What is my balance?” can return another account's result, even if the strings are identical. Correct key construction includes the relevant identity, context and data version. Even a perfect hash cannot repair a missing dependency.

Pick an appropriate first use case

Start with an assistant answering public questions about a versioned product manual. Candidate requests “How do I reset the device?” and “What are the reset instructions?” may be equivalent for the same device model, firmware, locale and reset type.

Contrast “restart the device” with “factory reset the device.” They share many terms but require different instructions; the latter may erase data. A high similarity score is not sufficient.

Functional requirements

  1. Reuse only eligible validated answers.
  2. Restrict candidates to the correct product, locale and knowledge revision.
  3. Fall back to normal generation when equivalence is uncertain.
  4. Invalidate results when source content or answer policy changes.
  5. Record why each answer was reused or rejected.

Non-functional requirements

  1. Set a maximum acceptable false-reuse rate for the use case.
  2. Protect tenant/private context and prevent cache poisoning.
  3. Keep cache lookup overhead below the expected benefit.
  4. Preserve acceptable miss-path and tail latency.
  5. Bound memory, retention and concurrent miss amplification.

Cache read-only answers first. Reusing a response must never substitute for executing a newly requested business action or checking its current authorization. An answer saying “payment sent” is not a reusable payment implementation.

Build the lookup path

Architecture / visual model
flowchart TD A[Request plus trusted scope and versions] --> E{Eligible for response reuse?} E -->|No| G[Generate normally from current evidence] E -->|Yes| K[Try correctly scoped exact cache] K -->|Valid hit| R[Return saved result] K -->|Miss| V[Embed request and search permitted candidates] V --> C{Fresh and equivalent under policy?} C -->|Yes| R C -->|No or uncertain| G G --> Q{Quality and cache eligibility checks pass?} Q -->|Yes| S[Store result with dependencies and expiry] Q -->|No| N[Apply normal response policy without caching] S --> D[Return generated result]
Read diagram source
flowchart TD
    A[Request plus trusted scope and versions] --> E{Eligible for response reuse?}
    E -->|No| G[Generate normally from current evidence]
    E -->|Yes| K[Try correctly scoped exact cache]
    K -->|Valid hit| R[Return saved result]
    K -->|Miss| V[Embed request and search permitted candidates]
    V --> C{Fresh and equivalent under policy?}
    C -->|Yes| R
    C -->|No or uncertain| G
    G --> Q{Quality and cache eligibility checks pass?}
    Q -->|Yes| S[Store result with dependencies and expiry]
    Q -->|No| N[Apply normal response policy without caching]
    S --> D[Return generated result]

Cache metadata can include:

  1. Tenant or public-data scope and authorization version where applicable.
  2. Product/entity IDs, locale and relevant request parameters.
  3. Conversation-state fingerprint when the answer depends on prior turns.
  4. Prompt/policy, model and tool-schema revisions as appropriate.
  5. Source-document or data snapshot versions.
  6. Creation time, expiry and validation provenance.

Not every field must be a hash-key component, but every answer-changing dependency needs a matching or invalidation rule. If tracking the dependencies is too difficult, exclude that request class from response caching.

Use similarity as a candidate score

For nonzero vectors, cosine similarity is:

similarity(q, x) = (q · x) / (||q|| × ||x||)

Under the common cosine-distance convention, distance = 1 − similarity. Higher similarity means closer; lower distance means closer. Check the actual implementation and normalization. A library using Euclidean distance has different units.

RedisVL SemanticCache supports distance thresholds, TTL and filterable metadata. Its documented Redis cosine distance range is 0–2. This is an implementation contract, not a statement that any particular cutoff makes answers safe.

Do not prescribe 0.95 similarity for one field or 0.98 for another without measured data. Embedding-model changes, languages, query lengths and reranking can change score distributions. A number near one is not a calibrated correctness probability.

Construct hard negative pairs

Request pair Why similar text is insufficient
Reset versus restart Different business action and consequences
“Can I cancel?” versus “Can I not cancel?” Negation and intent
Firmware 4.2 versus 4.3 Version-specific behavior
Account A versus account B Different private facts and permission
Before expiry versus after expiry Time changes the applicable answer
“How much did it cost?” in two conversations Pronoun refers to different entities

Normalize only transformations that preserve the required semantics. Removing numbers, punctuation or stopwords can erase version IDs, negation or signs. A semantic cache needs the same discipline as retrieval preprocessing.

Add validation where it earns its cost

Hard metadata checks are often cheaper and more reliable than asking a model to rediscover entity or version mismatches. A calibrated verifier can assess whether the candidate answer covers the request, but it adds latency and cost and can make mistakes. For some sensitive or dynamic classes, bypassing response reuse is the simpler design.

The verifier must see the actual request context and dependencies. Asking “are these sentences similar?” does not test answer validity. Keep the verifier's instructions separate from the cached text to reduce injection risk.

Admission also matters: only cache results that pass the product's quality checks. A cached hallucination can be repeated thousands of times without further model calls. Invalidating a poisoned entry must reach replicas and any downstream exact cache.

Calculate the break-even point

Let:

  • h = fraction of eligible requests served by accepted cache hits.
  • C_lookup = average embedding/search/validation cost per eligible request.
  • C_generate = average cost avoided by an accepted hit.

Ignoring fixed infrastructure and miss-write cost for the moment:

C_with_cache = C_lookup + (1 − h) × C_generate

Caching saves money when h > C_lookup / C_generate

Illustrative values: lookup costs $0.0001, generation costs $0.002 and accepted hit rate is 40%:

$0.0001 + 0.60 × $0.002 = $0.0013 per request

That is a 35% variable-cost reduction versus $0.002. The variable-cost break-even hit rate is 5%. Now add fixed cache hosting, storage, writes, invalidations, engineering and the cost of erroneous reuse. Low volume is not inherently unprofitable; fixed costs and request repetition determine the outcome. High volume does not guarantee a high hit rate.

If lookup takes 20 ms and normal generation takes 800 ms, a simplified serial model gives 20 + 0.60 × 800 = 500 ms mean latency. A miss takes 820 ms, so average improvement can coexist with a worse miss path. Measure percentiles; do not average percentile numbers as though they were mean samples.

Measure correctness alongside savings

Metric Denominator Interpretation
Accepted hit rate Eligible requests How often reuse actually occurs
False-reuse rate Accepted hits Fraction of reused answers that were invalid
Cache-caused error rate All measured requests Product-level impact of reuse errors
Avoided generation cost Matched generation baseline Gross saving before cache overhead
Stale-answer rate Returned answers Invalidation/freshness failures
Miss latency Cache misses Cost imposed when reuse fails

At 40% accepted hits and 0.5% false reuse among those hits, 10,000 eligible requests produce approximately 20 incorrect reused answers: 10,000 × 0.40 × 0.005. A high hit rate alone can conceal an unacceptable error count.

Calibrate on representative labeled request/answer pairs, keep a held-out set, and shadow-test candidate reuse before returning it to users. Sample accepted hits for ongoing review. Include rare but consequential mismatches, not just easy paraphrases.

Handle freshness, stampedes and modalities

TTL bounds age under its implementation; it does not prove the answer stayed valid during that interval. Invalidate on important source/policy changes. Extending TTL merely because an answer is popular can preserve obsolete information longer.

For identical concurrent misses, single-flight coordination can let one request compute the result while others wait. Set a wait deadline and failover behavior. Do not merge merely similar in-flight requests before establishing their equivalence.

Multimodal reuse needs additional caution. Two similar screenshots can contain different balances, error codes or names. Similar audio can contain a different account number or a negation. Exact content hashes plus model/preprocessing/task versions may be appropriate for repeat transcription of the identical asset; semantic similarity alone is not an audio fingerprint proving identity.

Interview practice

Q1: Why is an exact response cache still risky?

The key may omit account, conversation, permissions or data versions, and a saved answer can become stale. Exact matching only solves matching the chosen key, not the validity of that key's design.

Q2: Why not return the closest vector result every time?

Nearest does not mean equivalent. There may be no valid candidate. Apply scope, freshness and equivalence checks, and allow a miss.

Q3: Should we cache answers that trigger tool writes?

Do not use answer reuse as a substitute for the requested operation. Execute through the authorized workflow and handle retries using business-operation identity. Read-only explanatory content can have a separate cache policy.

Q4: How do you select the threshold?

Label representative positive and hard-negative pairs, measure false reuse against accepted hit rate, and choose a policy fitting the error budget. Recalibrate after model/index changes; a fixed score is not universally meaningful.

Q5: When would you remove the cache?

When low reuse, expensive validation, invalidation complexity or correctness failures outweigh its measured benefit. Compare full cost and latency, including misses, with the uncached baseline.

Q6: What is the most important closing statement?

State which request classes are eligible, which dependencies must match, how invalidation works and the measured false-reuse rate. Savings without that validity boundary are not a complete design.

Final notes

Recall card: Scope → freshness → equivalence → quality → economics. A cache hit is valuable only when the reused result is still valid.

Next: state management.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Agentic Memory with Mem0
NEXT LESSONState Management Patterns →

Explore the diagram