Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Agent Memory and State

By Anup Rai12 min readReviewed September 2026

Agent memory is information retained from earlier interactions or observations and made available for later tasks. State is the information needed to represent and continue the current computation. They overlap, but have different contracts: a preference can help a future answer, while an operation record determines whether a payment has already been submitted.

Memory does not necessarily change model weights. Most application memory systems store records outside the model and select relevant records into its input. The KV cache stores intermediate attention computations; it is not a durable database of task outcomes.

Separate the four common memory categories

The names below are useful analogies from cognitive science, not a mandatory four-level hardware hierarchy. One database can support several categories; one category can use several stores.

Category Meaning Example Typical implementation choice
Working memory Information available for the current task Current goal, recent observations and selected evidence Bounded model context plus explicit task state
Episodic memory Records of particular past events A failed deployment and its observed error Event records, summaries and searchable artifacts
Semantic memory Retained facts or propositions A user's preferred output format Scoped records with provenance and validity
Procedural memory Reusable instructions for performing a task A tested deployment procedure Versioned playbooks, skills or workflow definitions

“Semantic memory” does not mean that every stored statement is true. A user preference, authoritative configuration value and model inference have different sources and verification requirements. “Procedural memory” does not grant permission to execute a procedure.

Keep these mechanisms distinct

Mechanism What it provides What it does not provide
Conversation history Previous messages and observations A reliable record of external completion
Context summary Compact account of selected history Lossless preservation or guaranteed correctness
Runtime checkpoint Stored execution position and state Automatic rollback of remote side effects
Long-term memory store Cross-task records and retrieval Automatic source authority or freshness
Prefix/KV caching Reuse of eligible model computation Durable, independently editable business facts
Fine-tuned weights Learned behavior or statistical information Ordinary record-level update/deletion semantics

Use explicit results and task records for recovery. A hidden reasoning trace is neither required for memory nor a reliable substitute for an auditable state machine.

Begin with a concrete memory contract

For a developer assistant, distinguish these inputs:

  1. “Use compact tables in this conversation.” Store as session scope.
  2. “Use Python examples by default in future lessons.” Store as an explicit user preference, with a way to inspect and change it.
  3. “The service currently allows 1,000 requests per minute.” Retrieve current authoritative configuration when the decision depends on it; a cached observation must carry time and version.
  4. “This build failed because a package was unavailable.” Keep an event and its evidence. Do not automatically generalize to “never use that package.”
  5. “Deploy by following this reviewed runbook.” Keep a versioned procedure with prerequisites, permitted tools and verification steps.

Functional requirements

  1. Add memories with source, scope and validity information.
  2. Retrieve records relevant to an authorized task.
  3. Correct, supersede, retract and delete records.
  4. Explain which stored evidence influenced an answer.
  5. Resume tasks using durable execution state independently of semantic retrieval.

Non-functional requirements

  1. Isolate users, tenants, projects and sessions as required.
  2. Bound retrieval latency, prompt size and write cost.
  3. Preserve source lineage through summaries and consolidation.
  4. Prevent stale or deleted records from reappearing through caches and derived indexes.
  5. Evaluate correctness at extraction, update, retrieval and answer stages.

The need for memory depends on the product. A stateless one-shot classifier may need none. A long-running assistant needs deliberate state and retention rules, not simply a large context window.

Build a controlled write and read path

Architecture / visual model
flowchart LR O[Observation with identity and source] --> X[Extract candidate records] X --> V[Validate scope and evidence] V --> C[Resolve corrections and versions] C --> S[(Authoritative memory records)] S --> I[Search index and derived summaries] Q[Authorized task] --> F[Scope and validity filters] F --> I I --> R[Rank relevant candidates] R --> A[Recheck current records and permissions] A --> B[Assemble bounded context with source IDs] B --> M[Model response]
Read diagram source
flowchart LR
    O[Observation with identity and source] --> X[Extract candidate records]
    X --> V[Validate scope and evidence]
    V --> C[Resolve corrections and versions]
    C --> S[(Authoritative memory records)]
    S --> I[Search index and derived summaries]
    Q[Authorized task] --> F[Scope and validity filters]
    F --> I
    I --> R[Rank relevant candidates]
    R --> A[Recheck current records and permissions]
    A --> B[Assemble bounded context with source IDs]
    B --> M[Model response]

A simple baseline can be a relational table for explicit preferences and a separate event log. Add embeddings when matching by meaning improves retrieval; add a graph when relationship traversal is needed. A graph built from incorrect extracted relationships still returns incorrect facts.

Example record

This is an application schema, not a vendor API:

{
  "memory_id": "mem-42",
  "tenant_id": "tenant-7",
  "subject_id": "user-18",
  "scope": "user",
  "kind": "preference",
  "key": "example_language",
  "value": "Python",
  "status": "active",
  "source_id": "message-203",
  "source_type": "explicit_user_statement",
  "observed_at": "2026-09-24T14:00:00Z",
  "valid_from": "2026-09-24T14:00:00Z",
  "valid_to": null,
  "record_version": 1
}

Derive identity from authentication. Do not trust model-supplied tenant IDs. Preserve the supporting source according to the retention policy, or keep an appropriate source reference. A numeric model confidence field can assist triage, but it does not establish truth or access rights.

Resolve conflicts using meaning and time

Situation Suitable action Incorrect shortcut
Explicit replacement preference Supersede the old value within the same scope Leave both active and hope ranking chooses correctly
Temporary exception Apply narrower scope or validity interval Overwrite the user's global preference
Correction of an extraction error Retract/correct with lineage Treat the false value as historically true
New authoritative configuration Use the source revision and effective date Trust the newest model-generated sentence
Unresolved contradiction Preserve competing evidence and seek the relevant authority Average conflicting statements into a fact

Valid time describes when a fact applies in the modeled world. Transaction time describes when the database recorded it. Bitemporal storage records both, allowing questions such as “what was effective on September 1?” and “what did the system believe on September 5?”

For example, a team membership change takes effect September 1 but is entered September 10. A historical query using only the insertion timestamp cannot represent both dates correctly. If corrected later, retain the old recorded version for permitted audit queries while excluding it from the current view. A valid_to field alone does not make a system bitemporal.

Statuses such as active, superseded and retracted are useful application choices. They do not by themselves implement a formal belief-revision system. Nor is “user statement beats tool output” a universal ordering: a user controls their preference, while a service's authorized configuration controls its actual rate limit.

Retrieve useful memory within a budget

  1. Resolve identity and allowed scope.
  2. Apply current status, time and authorization constraints.
  3. Retrieve exact keys for known attributes and use search for less structured history.
  4. Rank by task relevance, source quality and appropriate recency.
  5. Remove duplicates and recheck authoritative versions.
  6. Fit evidence into a token budget, keeping source IDs and important exceptions.

A possible heuristic is a weighted score over normalized relevance, recency and importance. Tune it on representative tasks; recent irrelevant observations should not outrank older decisive evidence. The Generative Agents paper explored retrieval and reflection over past experiences, but its architecture is not proof that one scoring formula is universal. Research paper.

Suppose a 16,000-token input allocation reserves 3,000 for instructions/tools, 4,000 for current task evidence and 1,000 for the immediate conversation. That leaves 8,000 tokens for selected history and memory. A separate output allowance must still fit the model's actual context constraints. Retrieving five records is not inherently better than twenty: record size and relevance determine the budget.

Fast-changing facts can be stored as timestamped observations or cached with a suitable freshness contract. They should not silently become timeless current truth. An old stock price is valid historical data; a live trading decision requires appropriately current data from its source.

Consolidate without inventing facts

Consolidation reduces repeated or verbose episodes into useful records. It is a new transformation that can introduce errors.

Stage Check Example failure
Candidate extraction Is the statement supported? Turning a question into a preference
Deduplication Is this the same fact and scope? Merging two people's preferences
Conflict handling Which source and time apply? Keeping a superseded limit active
Summary Were exceptions and uncertainty retained? Removing “only for this session”
Procedure proposal Does the lesson generalize? Inferring a permanent rule from one outage
Publication Has the new version been evaluated? Replacing a working runbook with an untested reflection

One explicit statement can establish a preference. Repeating an unsupported inference five times does not validate it. There is no universal “three to five observations” threshold. Use domain-specific evidence requirements and test procedure changes before promotion.

Background extraction reduces user-facing latency but introduces a delay before the memory is visible. Synchronous writes can support immediate read-after-write expectations at higher latency. For an explicit preference change, acknowledge only after the intended durable update succeeds; background consolidation can follow later.

Handle poisoning and tenant isolation

Memory poisoning persists untrusted instructions or false experiences and reintroduces them in future context. MINJA studies query-only injection, while MemoryGraft studies poisoned experience retrieval. Their results establish attack mechanisms under tested conditions, not a universal production attack-success rate. MINJA, MemoryGraft.

Useful controls include:

  1. Preserve source and trust metadata through every derived record.
  2. Keep retrieved text separate from executable policy and trusted instructions.
  3. Restrict which component can publish procedures or change permissions.
  4. Enforce tool permissions independently of remembered recommendations.
  5. Test delayed attacks, cross-session contamination and correction behavior.
  6. Quarantine or retract compromised records and invalidate their derivatives.

A sanitizer or classifier can reduce risk, but cannot guarantee that all malicious text is recognized. A benign-looking false fact may be enough to change a later action.

Tenant isolation can use properly enforced row-level policies, scoped namespaces or separate stores. Choose according to the store's guarantees and required blast radius. A metadata label alone is insufficient; physical separation is not automatically safe if one privileged service can read every collection.

Check retrieval, record reads, caches, traces, backups, exports and shared summaries. Cache keys must incorporate the required security scope. Prefix-cache isolation addresses a different layer from application-result caching. Per-tenant encryption helps only if the key-access boundary remains enforced; an application holding every decryption key can still disclose data through an authorization bug.

Plan deletion and storage growth

Deleting a memory must account for the canonical row, search entries, summaries, caches and permitted backup-retention behavior. Track lineage so that a deleted observation cannot be re-extracted from an old summary. Use a durable deletion marker or equivalent control during asynchronous cleanup, and verify retrieval no longer returns the record.

For illustration, 50 million memories × 1 KB of serialized text/metadata is 50 GB in decimal units. One 1,536-dimensional float16 embedding per memory adds 153.6 GB of raw vectors. Three stored copies of both payloads total 610.8 GB, before indexes, graph edges, logs and backups. Averages hide tenant skew; measure the largest tenants separately.

Memory quality can decline when stale and low-quality records accumulate, but there is no fixed thirty-day failure threshold. This is retrieval/data-maintenance degradation, not necessarily catastrophic forgetting in model weights. Prune according to retention, relevance and evidential value; evaluate what useful information pruning removes.

Compare implementation approaches

These are starting points to evaluate as of this September 2026 review, not interchangeable promises or a leaderboard.

Approach What it offers What the application still decides
Relational records plus search Explicit schema, versions and flexible retrieval Extraction, ranking and lifecycle policy
Mem0 Extraction, consolidation and retrieval components Scope, correctness thresholds and deployment needs
Letta Persistent memory blocks and agent context management What belongs in context, who may edit it and how changes are reviewed
Graphiti Temporal entities/relationships and source-linked episodes Ontology, extraction quality and operational boundaries
LangMem Memory tools and background/procedural update patterns Promotion policy, evaluation and permissions
Versioned memory files/skills Inspectable records and procedures Access enforcement, conflicts, indexing and safe execution

See the primary descriptions for Mem0, Letta memory blocks, Graphiti and LangMem. Managed services and open-source libraries have different operating responsibilities. Consumer assistant or coding-product memory behavior does not establish the contract of an application API.

Evaluate the operations, then the outcome

Evaluation Question
Extraction precision/recall Did stored candidates preserve the supported facts and scope?
Update correctness Was a correction, expiration or replacement applied correctly?
Retrieval quality Was the right authorized current record selected?
Answer grounding Does the answer accurately use the selected memory?
Deletion test Can the deleted fact reappear through another path?
Longitudinal test Does quality hold after many sessions, updates and pruning cycles?
Isolation test Can one identity influence or read another's private memory?

HaluMem separates extraction, updating and memory question answering to localize errors that end-to-end scores can hide. Use this distinction without assuming a fixed percentage of all production errors comes from extraction. Calibrate automated judges with reviewed examples. HaluMem.

Compare no memory, full recent history, structured preferences and selective retrieval on the same tasks. Measure latency, storage, write amplification and failures alongside answer quality. Preserve tests for corrected facts, temporary preferences, ambiguous names and deleted sources.

Research extension: learned and adaptive memory

A-MEM explores dynamically organized notes and links; HippoRAG explores graph-based long-term retrieval. These are useful alternatives to a flat nearest-neighbor store, not reasons to remove provenance or access checks. A-MEM, HippoRAG.

TTT-E2E updates model weights from the supplied context. The paper distinguishes linear-time prefill from constant-cost decoding with respect to prior context length under its setup; “constant inference cost” should not be read as zero cost to ingest a longer history. It reports limitations on exact needle retrieval outside the attention window. This is a different mechanism from editable memory records. TTT-E2E.

For an application using such adaptation, define session/tenant isolation, adapted-state lifetime, reset/rollback and evidence retention. Deleting one database row does not demonstrate removal of its influence from adapted weights. Retain an external authoritative store when exact facts, corrections or attributable evidence are required.

Interview practice

Q1: Is the KV cache the agent's short-term memory?

It is a computational cache for attention, not the application's durable task state. The model's current context supplies working information; explicit checkpoints and operation records preserve what is needed to resume safely.

Q2: Where should a user preference live?

Use a scoped record with its supporting statement and a correction path. A conversation-only preference belongs to that session. Do not promote it to a global preference merely because a summarizer omitted the qualifier.

Q3: How do you resolve conflicting memories?

Identify whether the conflict is a temporal change, correction, scope difference or unresolved disagreement. Use the relevant source authority and effective time. Preserve versions where needed; neither recency nor model confidence alone decides truth.

Q4: Must semantic memory use a graph database?

No. A keyed preference can be a relational row. A graph helps when relationship traversal is central. Compare query needs, extraction cost, operational complexity and correctness before adding it.

Q5: When does a failed episode become a procedure?

After a supported lesson is proposed, checked for scope and tested against relevant cases. One outage does not justify a permanent universal rule. Publish a versioned procedure and retain a rollback path.

Q6: How do you prevent a deleted memory from returning?

Track derived records and indexes, invalidate caches and block re-extraction from retained summaries. Verify the deletion through actual retrieval paths. Define how backups and audit records are handled under the product's retention policy.

Q7: Is a separate collection per tenant sufficient isolation?

Only if access to that collection is enforced throughout the system. Check service privileges, cache keys, exports, logs and key access. A shared store with correctly enforced policies can also provide isolation; labels alone cannot.

Q8: What makes long-lived memory worse over time?

Unsupported writes, stale versions, lost scope and noisy retrieval can accumulate. Instrument each stage, consolidate with lineage and test across many sessions. Do not assume a specific day count or call every retrieval failure model forgetting.

Q9: What changes when context is learned into weights?

The information no longer has ordinary record-level lifecycle semantics. Evaluate exact recall, isolation and reset behavior, and retain authoritative evidence externally when the application needs it. The adaptation mechanism does not replace a business-state database.

Final notes

Recall card: Scope → source → validity → retrieval → correction → deletion. A useful memory system remembers the right information, can explain where it came from and can stop using it when it is no longer applicable.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Multi-Agent Orchestration
NEXT LESSONPlanning and Decomposition →

Explore the diagram