Ilustración editorial para Memoria de agentes de IA: separar contexto, preferencias y hechos persistentes, y decidir cuándo caducan
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

Remembering does not mean retaining the history

An agent that serves the same person or team on multiple occasions needs continuity: it may be reasonable to remember a preferred language, the usual format for a deliverable, or the name of a project. However, turning every conversation, retrieved document, and model inference into reusable memory creates a governance problem. The system may retrieve information that is irrelevant, stale, incorrect, sensitive, or applicable only to an earlier circumstance.

The architectural question is not whether the agent has memory, but which specific assertion it retains, who may use it, for what purpose, and for how long. A statement such as “the customer prefers to approve changes by email” may be a declared preference, an observation inferred from a single exchange, or a current operating rule. Those interpretations have different value and should not be stored or applied in the same way.

Persistent memory does not gain authority merely because it has been stored. An instruction appearing in a past conversation, an imported document, or a retrieval result remains content from the source that supplied it. It must not displace higher-authority instructions or enable actions the agent cannot perform in the current context. This separation is particularly important when retrieved content has not been validated or may have been manipulated.

For the purposes of this guide, it is useful to distinguish design facts from policy decisions. It is an operational fact that a system can assign creation, review, or expiration dates to records. By contrast, deciding that a preference expires after ninety days is an organizational policy: it should be justified by risk, data volatility, and the intended experience, not presented as a universal rule.

02

Four objects worth keeping separate

The word “memory” often combines components with very different life cycles. Separating them reduces improper retrievals and makes agent behavior easier to understand. Session context contains the immediate exchange needed to interpret the current request. It should be limited to the session and disappear or become inaccessible when it ends, unless part of it is deliberately promoted to another category.

Task state represents progress on a specific piece of work: a request identifier, items already processed, a draft in progress, a partial result, or a pending step. It may need to survive a brief interruption, but that does not make it a preference or a fact that should be available for future tasks. Google’s documentation for agent execution tools provides an example of workflow state tied to a session and of explicit or time-to-live-based cleanup when the task or conversation ends.

Declared memory is information that an individual or team has supplied for continuity, such as a preferred language, a time zone, or a naming convention. It should retain the original declaration or a reference to it, along with its scope. A setting specified for one project does not necessarily extend to every project, and one team member’s preference must not be attributed to the whole organization.

Retrievable knowledge consists of assertions obtained from identified sources: a current internal policy, the status of a service, or a list of accountable owners. Its design is less like a user profile and more like a provenance record: source, version or time of consultation, responsible party, scope, and update conditions. The W3C PROV-O ontology provides concepts for representing entities, activities, and agents, as well as generation, derivation, and invalidation relationships. It does not require any particular database, but it helps make traceability explicit.

Practical separation of information objects

ObjectPrimary purposeIndicative persistenceRisk if reused without control
Session contextUnderstand the current turn and its immediate referencesUntil the session endsCarrying details from one conversation into another
Task stateResume bounded workUntil the task is completed, cancelled, or expiresMistaking technical progress for a stable preference
Declared memoryAdapt future interactions within a scopeWhile a purpose and applicable consent or basis existApplying a preference outside its context
Retrievable knowledgeGround responses or decisions in sourcesAccording to source validity and reviewActing on stale information or information with uncertain provenance
03

The minimum record for governable memory

A record does not need to retain full conversational text to be useful. In many cases, a normalized assertion and metadata are enough. For example, instead of keeping an entire dialogue, a record could state that a person chose to receive summaries in Spanish for a specific workspace. Reducing content does not eliminate every risk, but it limits exposure and makes inspection easier.

At a minimum, every item should include a stable identifier; content or a reference to the content; the owner or subject to whom it is attributed; permitted purpose; scope of application; source and acquisition method; confidence level; creation date; last verification date; expiration rule; and deletion or restriction status. If it is derived from other data, it should also retain its dependencies. This makes it possible to find which memories require review when a source changes or when a person requests correction.

It is useful to distinguish confidence from authorization to use. A preference declared directly by the person may have high confidence regarding what they expressed, while still having limited scope. A fact automatically extracted from a document may be potentially useful, but require human validation or consultation of the primary source before it can trigger an action. Confidence must not become an opaque score that substitutes for provenance.

For personal information, the European Union’s General Data Protection Regulation establishes principles of purpose limitation, data minimization, accuracy, and storage limitation, together with rights related to rectification, erasure, and restriction of processing under applicable conditions. A product operating under that framework must translate those principles into real processes; adding a field called “TTL” does not, on its own, demonstrate legal compliance.

04

What may persist and what should disappear

Persistence should be decided by functional need and risk, not by ease of storage. An explicit, low-impact preference may persist if it has a clear owner, a specific purpose, and an accessible way to edit it. Task state will usually be deleted once work is finished. A changing external fact, such as a price, policy, or service availability, may be retained as a retrieval cue, but should not be treated as current evidence when a consequential decision is about to be made.

Inferences deserve a separate category. A model’s inference that someone prefers brief communications is not equivalent to a declaration by that person. If a product chooses to store inferences, it should label them as such, narrow their scope, set a short review period, and provide an easy mechanism to confirm, correct, or discard them. In high-impact scenarios, it is more prudent not to turn behavioral inferences into persistent memory without an explicit product decision and risk assessment.

Sensitive data require stricter assessment than presentation settings. Sensitivity depends both on the nature of the data and on context, recipient, purpose, and jurisdiction. This guide does not replace legal or security analysis. As a product criterion, the presence of sensitive data does not justify persistence: the team must demonstrate a specific purpose, access controls, limited retention, and an effective process for handling changes or deletion where appropriate.

It is also important not to hide this classification behind one “memory” label. The interface and internal APIs should reflect the differences: a user may want to edit a preference, cancel a paused task, or challenge the accuracy of a fact from an external source. These are different operations and require different audit trails.

Decision process before storing an item

  1. 01Identify whether the data is context, task state, a declared preference, an inference, or knowledge from a source.
  2. 02Define the owner, purpose, and minimum scope in which the data would be useful.
  3. 03Check whether persistence is necessary or whether retention for the session or task is sufficient.
  4. 04Record provenance, date, confidence, and dependencies; explicitly mark inferences.
  5. 05Assign an expiration, a review condition, and an action at expiry: delete, restrict, or reverify.
  6. 06Offer inspection and correction when the data is attributed to a person or affects that person’s experience.
05

Expiration: expiry, review, and events

An expiration date answers the question of when data should stop being retrieved automatically. Mandatory review answers another question: when it must be checked again before being considered valid. Both can coexist. For example, a formatting preference may remain available until the person changes or deletes it, while an operating policy may stay indexed but require consultation of its source before it is used to approve an action.

Event-based rules complement the clock. A change of project, role, provider, account, or document version may invalidate associated memories. If an assertion depends on a specific source, updating or withdrawing that source should trigger a review of its derivatives. Modeling derivations and invalidations makes it possible to find the affected set instead of waiting for each item to expire separately.

Expiration should not be confused with immediate physical deletion. For operational or regulatory reasons, there may be separate states such as “not retrievable,” “pending deletion,” or “retained under a specific policy.” What matters for agent behavior is that an expired or restricted item does not silently re-enter response context. The implementation should document who can access each state and for what purpose.

Specific time periods cannot be inferred from a general technical source. They must arise from the purpose, data type, risk of becoming outdated, applicable obligations, and operational needs. A team may define retention classes, but it should measure their effects: how many memories expire unused, how many are corrected, and how many retrievals are blocked because validity is lacking.

Indicative matrix for expiration and revalidation

ClassRetention ruleBefore actingEvent requiring review
Session contextDelete or isolate at the end of the sessionUse only in the current sessionClosure, abandonment, or identity change
Task stateExpire when completed or after defined inactivityConfirm that the task is still currentCancellation, error, or request change
Declared preferenceMaintain with scope and an edit optionCheck whether it conflicts with the current requestExplicit change, departure from the project, or deletion request
Changing external factRetain provenance and apply a short review cycleConsult the primary source if it conditions an actionNew version, provider change, or signal of conflict
Model inferenceRetain only if policy permits, and for a short periodDo not use for sensitive actions without confirmationCorrection, lack of evidence, or new contradictory behavior
06

Conflicts: current instructions, preferences, and sources

A common conflict arises when a remembered preference contradicts the present request. The simplest operating rule is that the person’s current request, within the system’s authorization and safety limits, takes precedence over an earlier preference. If someone previously asked for brief answers but now requests a detailed analysis, the agent should meet the current request and, where appropriate, offer to update the preference rather than changing it silently.

Instruction hierarchy is independent of memory. OpenAI’s model specification describes authority levels for instructions and notes that untrusted content does not gain authority by appearing in supplied or retrieved data. Therefore, an instruction stored in declared memory or incorporated from a document cannot override higher-authority rules. The design must preserve the origin of every instruction and prevent retrieval from presenting it as a system command.

When two knowledge sources differ, the agent should not resolve the discrepancy by inventing a synthesis. It should identify the conflict, prefer a primary or more recent source according to a defined policy, or request human intervention if the decision has material consequences. The memory record should mark disputed items so they are not reused as established facts.

Correction has two dimensions. Correcting the primary record prevents the data from appearing in new retrievals. Correcting its derivatives prevents it from surviving in summaries, indexes, caches, evaluations, or retrieval-training data, depending on the architecture. A rectification or deletion process that only updates a visible table while leaving the data in indexes that feed the agent does not meet the operational objective of preventing reuse.

Flow for rectification or deletion

  1. 01Authenticate and log the request, identifying the affected item and requested scope.
  2. 02Locate the original record, its versions, its derivations, and associated retrieval indexes or caches.
  3. 03Immediately change the usage state to block new retrievals while the request is processed.
  4. 04Correct, restrict, or delete according to the applicable decision; retain only the necessary operational evidence under a separate policy.
  5. 05Propagate the change to summaries, vectors, caches, and evaluation sets that could reintroduce the data.
  6. 06Run non-reuse tests and record the result and any known limitation.
07

Revalidate before taking an external action

Memory can help formulate a response, but it is not always enough to carry out an external action. If the agent is going to send information, modify a record, initiate a transaction, change permissions, or make an impactful decision, it must assess the validity of the data that conditions that action. The higher the impact and the more changeable the source, the greater the requirement for consultation or confirmation.

Evidence of validity is not a generic statement that a record has “high confidence.” It may be a recent query to the authorized primary source, explicit confirmation from the competent person, or signed data that remains valid under the rules of the domain. The system should record what check was performed, when, against which source, and what decision it permitted. If it cannot check, it should abstain, reduce the scope of the action, or request confirmation.

The OWASP Gen AI Security project identifies risks associated with memory and context poisoning in agentic applications. In design terms, this reinforces the need to separate retrieved content from authorized instructions, record provenance, and observe which memories influenced an action. A semantic filter alone does not guarantee that a memory is reliable or appropriate for the current goal.

NIST proposes governance, mapping, measurement, and management activities in its AI risk management framework, including monitoring and documentation during operation. Applied to memory, this suggests treating failed retrievals, stale memories, and repeated corrections as measurable risk signals, rather than merely isolated quality incidents.

08

Product controls and operational testing

A memory view is not merely a settings page. It should make it possible to understand which data the agent uses and distinguish, at a minimum, declared preferences, retrievable facts, inferences, and task states when they are visible to the person. For each item, a useful presentation includes understandable content, source or acquisition method, scope, last review, expiration date, and options to edit, restrict, or delete where appropriate.

The usage log is the technical complement to that view. When a problematic response occurs, the team needs to reconstruct which items were retrieved, which were discarded, which influenced the decision, and whether revalidation occurred. It is not necessary to expose all internal data to every operator: the logs themselves require access controls, minimization, and retention periods. Traceability should support investigation without becoming another unlimited memory store.

Testing should cover more than correct retrieval. Include cases where an old memory contradicts the current request; a source has changed; an inference is wrong; a preference belongs to another project; and a deleted item appears in a summary or cache. For external actions, test that the absence of current evidence blocks the action or requests appropriate confirmation.

Measure the rates of retrievals of expired or out-of-scope memories, detected and unresolved conflicts, rectifications, completed deletions, blocks caused by lack of revalidation, and actions that depended on memory. These metrics do not by themselves prove that the system is safe or compliant, but they reveal where memory is becoming a source of faulty decisions.

09

Deployment checklist

Before enabling persistent memory, the team should be able to answer several questions in a verifiable way: which data classes it stores; who owns them; what purpose justifies each class; how the source is identified; when it expires; which events invalidate it; who can view or change it; and what happens in indexes, caches, and summaries when it is corrected or deleted.

The architecture does not need to solve every risk through automation. In some domains, the right response is to limit memory to low-impact preferences, keep sensitive facts outside agent persistence, or require human review for high-impact changes. Operational continuity does not depend on remembering more, but on remembering only what can be governed.

To broaden the framework, connect this practice with Learn content on agent design, with Safety content on operational risks, and with the Glossary to align on terms such as provenance, retention, scope, and revalidation. Maintaining a common vocabulary prevents product, engineering, data, and security teams from using “memory” to mean incompatible mechanisms.

Pre-deployment checklist

  1. 01Every memory class has a documented definition, owner, purpose, scope, and retention rule.
  2. 02Items include provenance, review date, confidence, and usage or deletion status.
  3. 03Stored or retrieved instructions cannot elevate their authority by being in memory.
  4. 04External actions require current evidence from the appropriate source or confirmation.
  5. 05An interface or process exists for inspection, correction, restriction, and deletion.
  6. 06Deletion propagates to indexes, caches, summaries, and other derivatives defined by the architecture.
  7. 07Tests verify that expired, out-of-scope, or deleted data is not retrieved and does not condition actions.
  8. 08The system monitors stale retrievals, conflicts, revalidation failures, and correction requests.

Open questions

  • Specific expiration periods and categories of sensitive data depend on the use case, jurisdiction, contractual obligations, and architecture; this guide does not establish universal time limits.
  • A source can provide provenance without guaranteeing accuracy or current validity. Traceability enables review, but does not replace validation against the primary source.
  • The feasibility of deleting derivatives depends on the specific components, including indexing systems, caches, logs, and evaluation mechanisms. The scope must be documented and tested.
  • Rights and obligations arising from data-protection rules require contextual legal assessment; this description of regulatory principles is not legal advice.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction