An AI application's memory architecture defines what information it retains, how it changes and how the application selects it for later use. Some memory is conversation-specific; some persists across sessions. The model only uses the information made available through its input, tools or learned parameters—it does not automatically read every record the application stores.
Start an interview with three separate questions: What is stored? Who may use it? When is it valid? A vector database answers none of these on its own.
Separate purpose, scope and implementation
Working, episodic, semantic and procedural memory are useful cognitive analogies in agent design. They are not a mandatory CPU-like L1–L4 hierarchy. Nor does one category require a particular database or have an intrinsic latency. Memory terminology and scope.
| Category | Meaning in this guide | Example | Possible representation |
|---|---|---|---|
| Working context | Information selected for the current model interaction | Current task, recent messages, retrieved evidence | Messages and structured fields |
| Episodic memory | Records of particular events or experiences | A previous troubleshooting attempt and its outcome | Event rows, documents, artifact references |
| Semantic memory | Retained facts or assertions about entities | A user's preferred explanation language | Relational fields, documents or graph assertions |
| Procedural memory | Retained instructions or methods | A reviewed troubleshooting procedure | Versioned instructions, code or selected examples |
These categories can overlap. One completed exercise is an episode; a tested lesson derived from many exercises might inform a procedure. This promotion is an application decision, not an automatic consequence of storing an embedding.
Semantic memory does not mean semantic search. The former describes retained knowledge; the latter describes retrieval by meaning. A SQL lookup can retrieve a semantic-memory fact. Vector search can retrieve an episodic record.
| Separate mechanism | What it does | What it does not establish |
|---|---|---|
| KV cache | Reuses intermediate attention computations | Durable user preferences or authoritative truth |
| Conversation checkpoint | Preserves selected execution state | The outcome of an unrecorded external write |
| Document retrieval | Supplies evidence from a corpus | Permission to treat its text as instructions |
| Model weights | Encode learned statistical behavior | An editable, individually deletable user-memory table |
| Business system of record | Owns a domain's authoritative state | That every cached copy is current |
Begin with a small, concrete product
Consider an interview-practice tutor that should remember a learner's chosen language and completed exercises. This is a design exercise, not a claim about a deployed product.
Functional requirements
- Resume an unfinished practice session.
- Reuse explicit preferences in later sessions.
- Recall relevant prior attempts and feedback.
- Let the learner inspect, correct and remove remembered information.
- Distinguish verified exercise results from model-generated interpretations.
Non-functional requirements
- Enforce account and organization boundaries on every memory operation.
- Keep response latency and memory-read cost within the agreed budget.
- Preserve source, time, version and correction history where required.
- Propagate deletions and corrections to derived retrieval indexes and caches.
- Avoid silently converting temporary session choices into permanent preferences.
Start with a relational database: a preferences table, exercise-attempt records and session state. Load preferences by authenticated account ID and retrieve recent attempts by exercise/topic. This may satisfy the product without a vector index or graph.
The first failure might be a query such as “Which earlier problem had the same failure pattern?” Exact topic labels may miss a relevant exercise under a different name. Add semantic retrieval over attempt summaries if evaluation shows a useful gain. Add relationship traversal when the queries actually depend on linked concepts, prerequisites or projects.
Design the write path and the read path separately
Read diagram source
flowchart LR
A[Authenticated interaction or event] --> P[Apply scope and retention policy]
P --> E[Extract candidate assertions if needed]
E --> V[Validate source, meaning and version]
V --> S[(Authoritative records with provenance)]
S --> I[Update derived indexes]
Q[New scoped task] --> R[Retrieve permitted and current records]
S --> R
I --> R
R --> B[Select evidence within context budget]
B --> M[Model interaction]
C[Correction or deletion] --> S
C --> I
An explicit preference can be written through ordinary validated form/API code. A transcript may require an extraction model, but its output is a candidate assertion. “Use Java for this exercise” must not become “always prefers Java.” Attach the scope and source that justify the claim.
For reads, derive identity and allowed scopes from the authenticated application session. Never accept the model's proposed user_id as sufficient authorization. Apply access constraints before evidence reaches the model and revalidate returned records under the storage system's consistency contract.
Decide when writes become visible
| Pattern | Benefit | Cost or failure to handle |
|---|---|---|
| Write on the request path | Immediate acknowledgment and easier read-after-write behavior | Adds latency; write failure affects the request |
| Extract/index asynchronously | Keeps expensive processing off the response path | Temporary stale retrieval, retries and backlog |
| Store authoritative change synchronously, index later | Fast durable correction with cheaper search maintenance | Reader must account for index lag |
For the tutor, commit a language preference before confirming it to the learner. Indexing a long exercise transcript can happen later. If the next turn needs that transcript immediately, use the authoritative session record rather than waiting for search indexing.
Consolidation derives a more compact or useful representation from existing information. It can combine duplicate assertions or summarize episodes, but must retain enough provenance to explain and correct the result. Frequently retrieved information is not necessarily more truthful. An attacker can repeat a false claim; repeated retrieval can also reinforce the system's own mistake.
Resolve conflicts without inventing a truth hierarchy
| Conflict | Appropriate question | Example response |
|---|---|---|
| Old and new explicit preference | Do both apply to the same scope and period? | Supersede the old global choice, retain a historical record if permitted |
| Profile versus temporary request | Is this a session exception? | Use the requested language for this exercise only |
| Generated summary versus scored result | Which source owns this fact? | Use the authoritative assessment record |
| Two uncertain extracted assertions | Is either sufficiently supported? | Retain uncertainty or ask for clarification |
Semantic memory is neither immutable nor inherently authoritative. A job, address or preference can change. Record when a fact applies and when the system learned it. “Newest timestamp wins” is insufficient if the new item repeats an old document or applies to a different project.
Estimate the footprint
Illustrative assumptions: 100,000 learners, 40 retained records per learner, 600 bytes of text/metadata per record, and one 768-dimensional float32 embedding per record.
| Quantity | Calculation | Raw size |
|---|---|---|
| Records | 100,000 × 40 | 4 million |
| Text and metadata | 4 million × 600 bytes | 2.4 GB |
| Embeddings | 4 million × 768 × 4 bytes | 12.288 GB |
| Combined, three copies | (2.4 + 12.288) × 3 | 44.064 GB |
These decimal GB figures exclude indexes, database overhead, logs and backups. Do not embed fields that only need exact lookup. Evaluate whether embeddings, raw transcripts and replicas have the same retention requirements.
Latency comes from the concrete queries, index, network and load. There is no universal rule that semantic memory takes over 500 ms or episodic memory takes 100–300 ms. Measure the path your design uses.
Test memory as a lifecycle
- Extraction: Did the stored assertion preserve negation, scope and uncertainty?
- Update: Did a correction become effective without losing unrelated facts?
- Retrieval: Did the system find the right records and exclude forbidden ones?
- Use: Did those records improve the task outcome without irrelevant personalization?
- Deletion: Did the information disappear from active records and derived views under the defined retention policy?
- Recovery: Can indexing retries or restored backups resurrect a removed assertion?
An architecture that retrieves many records can still be worse if those records are stale or distracting. Compare with a no-memory baseline and a simple structured-profile baseline.
Interview practice
Q1: Is a three-tier memory architecture the industry standard?
No. Cognitive categories help describe information, but scope, authority, storage and retrieval are separate choices. Explain the required behaviors and then choose the smallest architecture that supports them.
Q2: Why not put the whole user history in the prompt?
It may exceed the model or application budget, increase processing cost and introduce irrelevant or obsolete evidence. Full history can still be a valid baseline for small workloads. Compare it with selective retrieval and evaluate task accuracy as well as cost.
Q3: Must semantic memory use a graph database?
No. A preference can be a versioned relational row. A graph becomes useful when relationships and traversal are central to the required queries; it adds identity-resolution and maintenance costs.
Q4: Where should an exercise score live?
In the assessment system's authoritative record. Memory can retain a reference or derived learning summary, but a model's recollection must not silently replace the scored result.
Q5: How would you prevent cross-account recall?
Derive scope from authenticated identity, enforce access in the read/write path, and test caches, indexes, exports and administrative operations as well as the main database. A namespace field without enforced checks is only a label.
Q6: What makes a good memory-service abstraction?
Explicit contracts for source, scope, version, freshness, correction and deletion, plus observable latency and cost. A generic remember(text) method hides too much if the product requires reliable updates or sensitive isolation.
Final notes
Recall card: Store deliberately → qualify the assertion → enforce scope → retrieve selectively → correct and forget reliably.
The Generative Agents research illustrates an architecture combining a memory stream, retrieval and reflection. It is a research design, not proof that every production product needs that arrangement.
Next: Short-term context, then long-term memory.