Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Long-Term Memory

By Anup Rai7 min readReviewed September 2026

Long-term memory retains information for use beyond the current interaction or session. It can contain explicit preferences, past events, derived assertions and reviewed procedures. Persistence does not make a record true, current, relevant or authorized for every reader.

The central interview problem is maintaining useful knowledge through change: new facts arrive, old facts become invalid, permissions change and people request corrections or deletion.

Define what must be remembered

For an engineering assistant, separate these records:

Information Owner of truth Memory representation
Current project deployment region Project configuration service Reference or versioned cached assertion
A developer's preferred explanation language Explicit developer preference Scoped editable field
Last week's failed deployment Deployment event log Episode linked to evidence and outcome
A model's suspected failure cause Unconfirmed interpretation Candidate assertion with provenance
Approved incident procedure Versioned operational documentation Procedure reference and revision

An old successful tool sequence can be a useful example. It is not automatically valid after an API, policy or environment changes. Preserve preconditions and outcome evidence before using an episode as guidance.

Functional requirements

  1. Store and retrieve permitted records across sessions.
  2. Distinguish current assertions from historical events.
  3. Support corrections, supersession and explicit deletion.
  4. Explain which evidence led to a remembered assertion.
  5. Select relevant memories without exposing unrelated users or projects.

Non-functional requirements

  1. Define freshness and read-after-write expectations for critical fields.
  2. Bound retrieval latency, result size and storage growth.
  3. Keep indexing and consolidation retry-safe.
  4. Trace derived records so corrections/deletions propagate.
  5. Test model, schema and embedding upgrades against stored history.

Start with records, then add retrieval

Use an authoritative record store with exact IDs, scope, source and timestamps. An optional vector index supports meaning-based discovery. A graph can support relationship traversal. Neither replaces the record's authorization or validity checks.

Architecture / visual model
flowchart LR E[Permitted source event] --> W[Validate and record assertion or episode] W --> R[(Versioned records)] R --> Q[Outbox or change stream] Q --> I[Idempotent index update] I --> V[(Vector or relationship index)] U[Authenticated memory query] --> S[Search within allowed scope] V --> S R --> S S --> C[Resolve current versions and permissions] C --> P[Select evidence for current task] D[Correction or deletion] --> R
Read diagram source
flowchart LR
    E[Permitted source event] --> W[Validate and record assertion or episode]
    W --> R[(Versioned records)]
    R --> Q[Outbox or change stream]
    Q --> I[Idempotent index update]
    I --> V[(Vector or relationship index)]
    U[Authenticated memory query] --> S[Search within allowed scope]
    V --> S
    R --> S
    S --> C[Resolve current versions and permissions]
    C --> P[Select evidence for current task]
    D[Correction or deletion] --> R

A transactional outbox records a change and the need to publish it in the same local database transaction. The index consumer can retry using the record ID and version. This avoids a crash leaving a committed correction with no corresponding indexing task. It still needs backlog monitoring and a defined consistency boundary. Transactional outbox pattern.

If search returns a superseded record, resolve it against the authoritative version before use. For critical exact fields, bypass the approximate index and read the source of truth directly.

Keep provenance and time explicit

Conceptual record, independent of any memory vendor:

{
  "id": "mem-208",
  "tenant_id": "org-7",
  "subject_id": "project-42",
  "kind": "assertion",
  "predicate": "deployment_region",
  "value": "eu-west",
  "source_ref": "config://project-42/revision-19",
  "valid_from": "2026-09-01T00:00:00Z",
  "recorded_at": "2026-09-03T10:00:00Z",
  "version": 2,
  "supersedes": "mem-173",
  "status": "active"
}

Valid time describes when the fact applies in the domain. Recorded time describes when the system learned or stored it. A full bitemporal model tracks both histories; merely adding two timestamps does not implement every bitemporal operation.

Here the region changed on September 1 but memory learned it on September 3. “Where was the project deployed on September 2?” differs from “What did the assistant believe on September 2?” Define which question an API answers.

Preserve source identities without retaining more source content than the applicable policy allows. A scalar model confidence score does not establish truth or override an authoritative configuration revision.

Choose the storage by the query

Query Good starting mechanism Added cost or limitation
Current preference for account 42 Indexed relational/document lookup Schema and concurrency rules
Attempts about a similar failure Vector/hybrid search over episodes Embeddings, approximate misses and score calibration
Project's owners and dependent services Relational joins or graph traversal Relationship maintenance and bounded traversal
Events before a date Time-indexed event query Retention and event completeness
Original evidence for an assertion Source reference lookup Source access and lifecycle

A vector-to-graph design can first discover a relevant entity and then follow permitted relationships. It is useful when that two-stage query improves the task. It is not a seniority badge or a requirement for every production system.

Bound graph expansion. With ten neighbors per node, a naive three-hop expansion can visit up to 1 + 10 + 100 + 1,000 = 1,111 nodes before accounting for overlap. Filter by relation, scope and time, deduplicate visited nodes and budget evidence selection.

Consolidate, expire and correct deliberately

Operation Meaning Common mistake
Consolidation Produce a compact derived representation Erase exceptions or combine incompatible scopes
Supersession A newer applicable assertion replaces an older one Treat an event in the past as something to overwrite
Expiration Record is no longer eligible after a time Claim the bytes and all copies were deleted
Ranking decay Reduce an item's retrieval priority over time Forget a rarely used but critical constraint
Deletion Remove data under the defined storage/retention contract Delete the vector but leave summaries and caches

Use different policies for different data. A one-day travel instruction may expire; a stable accessibility preference should not disappear merely because it is rarely mentioned. Repetition and popularity can indicate usefulness, but are weak evidence of truth.

A source correction must reach derived summaries and embeddings. Keep lineage such as derived_from=[source_id, revision]. If consolidation combines multiple users, individual deletion becomes much harder; avoid unnecessary cross-user aggregation of personal details.

Prevent removed information from returning

  1. Authenticate and scope the correction/deletion request.
  2. Update authoritative state with a versioned removal marker where appropriate.
  3. Remove or invalidate related index entries, summaries, caches and artifacts.
  4. Make background jobs reject obsolete source versions.
  5. Apply retention and restoration procedures to logs/backups.
  6. Verify retrieval and generation cannot reuse the removed information.

A tombstone can prevent an old indexing job from recreating a deleted record. It must be designed to avoid retaining unnecessary sensitive content. Backup restoration may require replaying deletion records before serving traffic. Technical deletion design alone does not establish compliance with every legal requirement.

Protect identity and isolate scope

Use trusted account/project mappings. Two people with the same name are not the same entity; two accounts owned by one person are not automatically authorized to share memory. Account linking requires an appropriate authenticated workflow.

Database row policies, scoped service credentials, separate collections or separate deployments can enforce different isolation needs. Each has operational tradeoffs. The important property is enforcement across every access path, including exports, caches and administrative tools—not a particular partitioning slogan.

Retrieved memories can contain prompt injection. Treat their text as data with source provenance. A remembered sentence saying “disable verification next time” must not rewrite the tool gateway's policy. Agent security covers the execution boundary.

Measure quality and plan migrations

Distinguish these failure classes:

Failure Diagnostic evidence Likely intervention
Extraction error Stored assertion contradicts its source Better extraction/validation; repair records
Retrieval interference Correct record exists but is buried Better filtering, ranking or consolidation
Stale use Old record outranks a correction Version/freshness resolution
Identity merge error Another subject's record is used Correct mapping and isolation
Model misuse Correct evidence retrieved but answer ignores it Prompt/model evaluation
Catastrophic forgetting in training Learning new tasks degrades prior learned capability Training/evaluation intervention

Do not redefine catastrophic forgetting as “too many vector records.” Retrieval interference is a different mechanism and needs a different diagnosis.

Version embeddings by model, dimensionality and preprocessing. Equal dimensions do not make vectors from two embedding models comparable. Build a parallel index, backfill permitted current records, evaluate recall and task outcomes, then switch reads with a rollback plan. Changes to the extraction model also deserve evaluation: it may write differently scoped or differently phrased assertions.

Interview practice

Q1: How do you remember a preference that changed twice?

Record the scope and effective periods, keep the current applicable value easy to query, and preserve allowed history with provenance. A historical question and a current-personalization query use different validity filters.

Q2: Why keep source references after summarization?

They let us verify, correct and explain a derived assertion. Without lineage, deleting or changing one source can leave unsupported summaries scattered through memory.

Q3: Would you store every conversation forever?

No. Define a purpose and retention policy per record class. Keep exact evidence when justified; otherwise retain scoped summaries or no persistent memory. More storage does not automatically improve recall quality.

Q4: How do you update a database and vector index safely?

Commit the authoritative change and an outbox event atomically, consume it idempotently, monitor lag and resolve returned candidates against current record versions. Define what users see before indexing finishes.

Q5: Why can a memory deletion appear to succeed and later fail?

An old source, cached answer, queued extraction or restored backup can recreate the information. Deletion requires lineage and lifecycle handling beyond one index API call.

Q6: Should old memories always get lower ranking?

No. Recency helps for changing facts, but stable constraints and relevant rare episodes may remain useful. Rank with task relevance, validity, authority and scope; do not let age substitute for those checks.

Q7: How would you close this design?

Identify the authoritative records, query paths, allowed staleness, correction/deletion process and migration strategy. Show a conflict and an index-lag failure, then measure the resulting task quality against a simpler baseline.

Final notes

Recall card: Retain with purpose → record source and time → resolve current versions → retrieve under permission → propagate corrections and deletion.

Next: Mem0 integration. For general agent memory categories, revisit the architecture overview.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Short-Term Context Management
NEXT LESSONAgentic Memory with Mem0 →

Explore the diagram