Architecture Memory July 2026 · 10 min read

Letta vs Zep/Graphiti vs Mem0: Choosing an AI Agent Memory Architecture

An AI agent can produce an excellent answer today and still forget the entire interaction tomorrow.

That happens because an LLM's context window is working memory, not persistent storage. Passing more chat history into every prompt can preserve context for a while, but it increases latency, cost, and noise—and it still does not solve fact updates, contradictions, or memory lifecycle management.

This guide compares three notable approaches to persistent agent memory:

  • Letta — a stateful agent runtime with tiered memory
  • Zep / Graphiti — temporal memory built around entities and relationships
  • Mem0 — a developer-friendly memory layer for personalization and cross-session recall

We also compare them with the DIY approach of building a custom pipeline on top of a vector database. The goal is to explain how these systems differ, what trade-offs they make, and which architecture is most appropriate for your use case.

1. Quick Answer

  • Choose Letta when your agent should explicitly manage its own persistent state, memory hierarchy, and long-running behavior.
  • Choose Zep or Graphiti when temporal facts, entity relationships, provenance, historical queries, and auditability matter.
  • Choose Mem0 when you want to add cross-session personalization and memory retrieval to an existing agent with minimal architectural rework.
  • Build a custom pipeline when compliance, retention policy, data residency, or domain-specific memory logic are core requirements.

💡 A Note on Product Capabilities

Agent memory tools evolve quickly. Features, APIs, pricing, hosting options, and benchmark results may change between releases. The comparisons below describe the capabilities and architectural patterns available at the time of review. Always verify current documentation before selecting a production dependency.

2. Why Vector Databases Alone Are Not Agent Memory

Before evaluating dedicated memory systems, it's important to understand why standard Retrieval-Augmented Generation (RAG) is only one piece of the puzzle.

A vector database can retrieve similar text. It does not automatically know whether a fact is current, contradictory, private, important, or worth remembering.

RAG solves "how to find similar content." Dedicated memory systems solve "what to remember, when to update it, and when to forget it."

3. One User Update, Four Memory Architectures

To see the difference in architectures, consider a simple scenario where a user changes a preference over time:

  • January: "I live in Berlin."
  • April: "I moved to Tokyo."
  • June: "Where do I live now?" / "Where was I living in February?"
System Likely Memory Behavior
Letta The agent decides whether and how to overwrite its core memory block using tool calls.
Zep / Graphiti Can preserve both the old and new facts as temporally bounded relationships.
Mem0 Designed to update the current user memory, with historical behavior depending on configuration and implementation.
DIY Vector RAG May retrieve either or both statements unless custom update and temporal logic exists.

4. How We Evaluate Memory Systems

We evaluate each tool across five technical dimensions:

  1. Memory representation: How is data structured (blocks, graphs, vectors)?
  2. Write and update pipeline: Does the agent write it, or is it automatically extracted?
  3. Temporal and conflict handling: How does it deal with facts that change over time?
  4. Retrieval and context assembly: How is memory pulled back into the LLM context?
  5. Deployment and operational complexity: How hard is it to run in production?

5. Why This Guide Focuses on Three Tools

This article focuses on Letta, Zep/Graphiti, and Mem0 because they represent three distinct memory architectures: agent-managed tiered memory, temporal graph memory, and memory middleware for existing applications.

Other tools—including knowledge-graph platforms (like Cognee), conversation-memory servers (like Motorhead), and vector-database stacks—can still be strong choices for narrower requirements. See our broader AI Agent Memory Tools guide for a wider market overview.

6. Head-to-Head Comparison Table

Dimension Letta Zep / Graphiti Mem0 DIY Pipeline
Primary abstraction Stateful agent runtime Temporal memory / knowledge graph Memory API and personalization layer Custom data pipeline
Memory write path Agent-directed tool calls Automatic extraction Automatic extraction and updates Build yourself
Core memory model Core, archival, recall Episodic and semantic graph Semantic and episodic memory Depends on design
Temporal queries Limited / implementation-dependent Strong when using temporal graph features Usually update-oriented rather than historical Build yourself
Conflict handling Agent-dependent Explicit temporal facts Automated update pipeline Build yourself
Retrieval Agent tools and archival search Graph and semantic retrieval Semantic and filtered retrieval Vector / hybrid / custom
Self-hosting Available depending on deployment Graphiti can be self-hosted OSS/self-hosting options Full control
Operational complexity Medium to high Medium to high Low to medium High
Best fit Autonomous stateful agents Enterprise knowledge and history Fast personalization Highly custom systems

7. Letta — Stateful Agents with Tiered Memory

Philosophy: Treat the context window like virtual memory in an OS. The agent manages its own RAM.

Architecture

Letta provides a runtime where agents explicitly manage tiered memory:

  • Core Memory: Always in context. Structured blocks like "Human" (user facts) and "Persona" (agent rules).
  • Recall Memory: Short-term conversational history.
  • Archival Memory: External storage for deep knowledge, retrieved on demand.

Write Path

Memory is primarily written through agent-directed tool calls. The agent can decide, through memory tools, whether information belongs in core memory, archival memory, or conversation recall.

Read Path

Core memory is injected automatically. For archival memory, the agent explicitly calls search tools to page information into its working context.

Update and Conflict Handling

Because the agent explicitly edits its core memory blocks (e.g., calling core_memory_replace), conflict handling is largely agent-dependent. The system relies on the LLM's reasoning to overwrite outdated facts.

Deployment

Letta offers both self-hosted options and managed cloud services. Because it is an agent runtime, adopting Letta means running your agents inside its loop, which is a significant architectural commitment. Letta's repository is available under the Apache 2.0 license (verify current license for production use).

python
# Conceptual example; check the current Letta SDK for exact API names.
agent = client.create_agent(
    name="support-agent",
    memory_blocks={
        "human": "Name: Unknown. Preferences: Unknown.",
        "persona": "I am a helpful assistant."
    },
    tools=["archival_memory_insert", "core_memory_replace"]
)
# The agent autonomously uses its tools to update its core memory 
# when it learns new facts about the user.

Strengths & Limitations

  • Strengths: The agent explicitly controls its memory, allowing complex reasoning. Strong support for stateful, long-running agent processes.
  • Limitations: Requires adopting Letta as your agent runtime. Memory operations consume additional LLM tokens and tool calls. Less explicit temporal indexing compared to graph-based approaches.

8. Zep and Graphiti — Temporal Knowledge Graph Memory

⚠️ Important distinction

Zep Cloud and Graphiti are related but should not be treated as identical products. Zep is the hosted memory product. Graphiti refers to the open-source temporal knowledge-graph engine associated with this architectural approach.

Philosophy: Memory is a temporal knowledge graph. Facts have lifespans and relationships.

Architecture

This architecture builds a knowledge graph from interactions, categorizing data into:

  • Episodic: Raw interaction data and provenance.
  • Semantic: Extracted entities, relationships, and facts.
  • Community: High-level structural summaries of the graph.

Write Path

Unlike Letta's agent-driven approach, Zep uses automatic extraction. You pass chat messages or documents into the system, and it asynchronously extracts entities and relationships into the graph in the background.

Read Path

At query time, the system can combine semantic retrieval with graph traversal to retrieve relevant entities, relationships, episodes, and temporally valid facts. The retrieved context should then be filtered by relevance, permissions, provenance, and the time period the agent is being asked about.

Update and Conflict Handling

The standout feature is explicit temporal facts. Zep/Graphiti's temporal modeling is designed to preserve fact validity over time. When a fact changes (e.g., a user moves cities), the old fact isn't simply deleted; it is marked as invalid from that timestamp forward. This supports historically grounded retrieval when configured correctly.

Deployment

Zep Cloud is a managed service, heavily emphasizing enterprise compliance (always check their official Trust page for current SOC 2 Type 2 / HIPAA BAA applicability). Self-hosting is possible via Graphiti, but it requires managing your own compatible graph database infrastructure.

Strengths & Limitations

  • Strengths: Temporal modeling for facts that change over time. Graph-based representation of entities and relationships. Can support historically grounded retrieval and audit-oriented workflows when configured correctly. Automatic extraction reduces the amount of memory-tool orchestration required from the agent.
  • Limitations: Self-hosting Graphiti carries medium-to-high operational complexity. Cloud versions create vendor reliance. Less granular agent autonomy over exactly how memories are formatted.

9. Mem0 — Memory Middleware for Personalization

Philosophy: Provide a developer-friendly memory API to add personalization and cross-session recall to existing agents.

Architecture

Mem0 acts as a memory middleware. While architectures vary by deployment, Mem0 can be configured with vector-based memory and, depending on the edition and setup, additional graph or structured-memory capabilities.

Write Path

Mem0 uses automatic extraction and updates. You send conversational turns to the API, and the system handles embedding and categorization under specific namespaces (User ID, Session ID, Agent ID).

Read Path

Semantic retrieval across the user's namespace returns the most relevant facts filtered by relevance and recency.

Update and Conflict Handling

Mem0 provides an automated memory-update workflow intended to identify and consolidate changing user facts. Depending on the model, configuration, and memory store, it may update, merge, retain, or deprioritize older facts when new information conflicts with them.

Deployment

Mem0 offers both a managed platform (SaaS) and open-source self-hosting options. It can be deployed locally with compatible local models and storage backends (like Ollama and Qdrant) for privacy-sensitive applications.

python
# Simplified example of Mem0 integration
from mem0 import Memory

m = Memory()

# The system automatically extracts facts from the input
m.add(
    "I'm Alice. I moved from Berlin to Tokyo last month.",
    user_id="alice"
)

# Semantic retrieval filters by user namespace
results = m.search("Where does Alice live?", user_id="alice")

Strengths & Limitations

  • Strengths: Fast time-to-market; can be dropped into existing LangChain or CrewAI projects easily. Clear namespacing logic.
  • Limitations: Typically prioritizes updating over preserving explicit historical timelines (unlike a bi-temporal graph). The agent does not explicitly orchestrate its memory hierarchy (unlike Letta).

10. DIY Memory Pipelines — When Full Control Is Worth It

For teams with strict compliance needs or existing infrastructure, building a custom memory pipeline on top of a vector database (like Qdrant, Pinecone, Chroma, or Weaviate) is still a valid approach.

A Minimum Viable Production Architecture

text
Ingestion → PII/Safety Filter → Fact Extraction → Conflict Detection 
→ Temporal Store / Vector Store → Retrieval Policy → Context Assembler 
→ Audit Log → TTL / Deletion Worker

When to Build Your Own

For many teams, a dedicated memory layer is cheaper to maintain than rebuilding extraction, updates, and lifecycle management from scratch. Custom implementations still make sense when:

  • Operating in high-privacy environments (healthcare, finance, legal).
  • You have complex data residency, user-deletion rights, or retention requirements.
  • You already operate PostgreSQL, Kafka, Neo4j, or vector databases at scale.
  • The memory strategy itself is your core product differentiator.

11. Production Deployment and Governance Checklist

Choosing a tool is only step one. Use this checklist to ensure your memory architecture is ready for production:

  • Is memory securely namespaced by tenant, user, agent, and session?
  • Are sensitive inputs (PII, passwords) filtered before persistent storage?
  • Can users inspect, correct, export, and delete their stored memories?
  • Are episodic memories subject to TTL (Time-To-Live) and retention policies?
  • Are memory writes logged and auditable?
  • Is retrieval filtered by relevance, recency, permissions, and confidence?
  • Have you tested prompt injection and memory-poisoning attempts?
  • Do you need current-state answers, historical-state answers, or both?
  • Can the system distinguish a user preference from an untrusted instruction?
  • Is there an evaluation set for memory precision, recall, update accuracy, and leakage?

12. Which Tool Should You Choose?

There is no universal best memory system for AI agents.

  • Choose Letta when the agent itself should actively manage persistent state and memory.
  • Evaluate Zep or Graphiti when temporal facts, entity relationships, provenance, and auditability are central requirements.
  • Choose Mem0 when you want to add cross-session personalization to an existing agent with minimal architectural work.
  • Build a Custom Pipeline when you need full control over schemas, retention, privacy, retrieval, or domain-specific memory policies.

The important distinction is not whether a tool uses vectors, graphs, or key-value storage. It is whether the system gives you reliable control over what gets remembered, how memories change, how they are retrieved, and when they should be removed.

13. Frequently Asked Questions (FAQ)

What is the difference between semantic, episodic, and temporal memory?

  • Episodic memory records the raw "who said what and when" (conversation logs).
  • Semantic memory extracts the underlying facts and entities ("Alice lives in Berlin").
  • Temporal memory tracks the validity of those facts over time ("Alice lived in Berlin until April, then moved to Tokyo").

How should AI agents handle memory poisoning?

Treat all candidate memories as untrusted input. Separate user facts from executable instructions, validate high-impact writes, attach provenance, apply TTLs where appropriate, and evaluate the system against prompt-injection and poisoning scenarios.

Is a vector database enough for agent memory?

Usually, no. While vector databases are excellent for semantic retrieval, they do not natively handle fact updates, contradiction resolution, or temporal tracking—features required for true agent memory.

Related Tools and Guides

Explore hundreds of curated AI agent tools, frameworks, vector databases, and infrastructure at AgDex.ai.

Published by the AgDex.ai editorial team. Building something cool with agent memory? Drop a comment — we'd love to feature your use case.

Arquitectura Memoria Julio 2026 · 10 min de lectura

Letta vs Zep/Graphiti vs Mem0: Eligiendo una Arquitectura de Memoria para Agentes de IA

Un agente de IA puede dar una respuesta excelente hoy y olvidar toda la interacción mañana.

Esto sucede porque la ventana de contexto de un LLM es memoria de trabajo, no almacenamiento persistente. Pasar más historial de chat en cada prompt puede preservar el contexto por un tiempo, pero aumenta la latencia, el costo y el ruido, y aún así no resuelve las actualizaciones de hechos, las contradicciones o la gestión del ciclo de vida de la memoria.

Esta guía compara tres enfoques notables para la memoria persistente de agentes:

  • Letta — un entorno de ejecución de agentes con estado y memoria por niveles
  • Zep / Graphiti — memoria temporal construida en torno a entidades y relaciones
  • Mem0 — una capa de memoria amigable para desarrolladores para personalización y recuerdo entre sesiones

También los comparamos con el enfoque DIY de construir una canalización personalizada sobre una base de datos vectorial. El objetivo es explicar cómo difieren estos sistemas, qué compensaciones hacen y qué arquitectura es más adecuada para su caso de uso.

1. Respuesta Rápida

  • Elija Letta cuando su agente deba gestionar explícitamente su propio estado persistente, jerarquía de memoria y comportamiento a largo plazo.
  • Elija Zep o Graphiti cuando los hechos temporales, las relaciones de entidades, la procedencia, las consultas históricas y la auditabilidad sean importantes.
  • Elija Mem0 cuando desee añadir personalización entre sesiones y recuperación de memoria a un agente existente con una refactorización arquitectónica mínima.
  • Construya una canalización personalizada cuando el cumplimiento, la política de retención, la residencia de datos o la lógica de memoria específica del dominio sean requisitos centrales.

💡 Una Nota sobre las Capacidades del Producto

Las herramientas de memoria de agentes evolucionan rápidamente. Las características, APIs, precios, opciones de alojamiento y resultados de benchmarks pueden cambiar. Las comparaciones a continuación describen las capacidades disponibles en el momento de la revisión. Siempre verifique la documentación actual.

2. Por qué las Bases de Datos Vectoriales por sí Solas No Son Memoria de Agentes

Antes de evaluar sistemas de memoria dedicados, es importante entender por qué la Generación Aumentada por Recuperación (RAG) estándar es solo una pieza del rompecabezas.

Una base de datos vectorial puede recuperar texto similar. No sabe automáticamente si un hecho es actual, contradictorio, privado, importante o digno de recordar.

RAG resuelve "cómo encontrar contenido similar". Los sistemas de memoria dedicados resuelven "qué recordar, cuándo actualizarlo y cuándo olvidarlo".

3. Una Actualización de Usuario, Cuatro Arquitecturas de Memoria

Para ver la diferencia en arquitecturas, considere un escenario simple donde un usuario cambia una preferencia con el tiempo:

  • Enero: "Vivo en Berlín."
  • Abril: "Me mudé a Tokio."
  • Junio: "¿Dónde vivo ahora?" / "¿Dónde vivía en febrero?"
Sistema Comportamiento Probable de Memoria
Letta El agente decide si sobrescribir su bloque de memoria central y cómo hacerlo mediante llamadas a herramientas.
Zep / Graphiti Puede preservar tanto los hechos antiguos como los nuevos como relaciones temporalmente limitadas.
Mem0 Diseñado para actualizar la memoria actual del usuario, con comportamiento histórico dependiendo de la configuración.
DIY Vector RAG Puede recuperar cualquiera o ambas declaraciones a menos que exista una lógica de actualización personalizada.

4. Cómo Evaluamos los Sistemas de Memoria

Evaluamos cada herramienta en cinco dimensiones técnicas:

  1. Representación de memoria: ¿Cómo se estructuran los datos (bloques, grafos, vectores)?
  2. Canalización de escritura y actualización: ¿El agente la escribe, o se extrae automáticamente?
  3. Manejo temporal y de conflictos: ¿Cómo maneja los hechos que cambian con el tiempo?
  4. Recuperación y ensamblaje de contexto: ¿Cómo se devuelve la memoria al contexto del LLM?
  5. Despliegue y complejidad operativa: ¿Qué tan difícil es ejecutar en producción?

5. Por qué Esta Guía se Centra en Tres Herramientas

Este artículo se centra en Letta, Zep/Graphiti y Mem0 porque representan tres arquitecturas de memoria distintas: memoria por niveles gestionada por el agente, memoria de grafo temporal y middleware de memoria para aplicaciones existentes.

Otras herramientas, incluidas las plataformas de grafos de conocimiento (como Cognee) y servidores de memoria (como Motorhead), siguen siendo opciones sólidas.

6. Tabla de Comparación Cara a Cara

Dimensión Letta Zep / Graphiti Mem0 Canalización DIY
Abstracción principal Entorno de agente con estado Memoria temporal / grafo de conocimiento API de memoria y capa de personalización Canalización de datos personalizada
Ruta de escritura de memoria Llamadas a herramientas del agente Extracción automática Extracción y actualizaciones automáticas Constrúyalo usted mismo
Modelo de memoria central Central, archivo, recuerdo Grafo episódico y semántico Memoria semántica y episódica Depende del diseño
Consultas temporales Limitado / dependiente de implementación Fuerte al usar características de grafo temporal Usualmente orientado a actualizaciones más que histórico Constrúyalo usted mismo
Manejo de conflictos Dependiente del agente Hechos temporales explícitos Canalización de actualización automatizada Constrúyalo usted mismo
Recuperación Herramientas de agente y búsqueda en archivo Recuperación semántica y de grafo Recuperación semántica y filtrada Vectorial / híbrida / personalizada
Alojamiento propio Disponible dependiendo del despliegue Graphiti puede ser auto-alojado Opciones OSS/auto-alojamiento Control total
Complejidad operativa Media a alta Media a alta Baja a media Alta
Mejor uso Agentes autónomos con estado Conocimiento empresarial e historia Personalización rápida Sistemas altamente personalizados

7. Letta — Agentes con Estado y Memoria por Niveles

Filosofía: Tratar la ventana de contexto como memoria virtual en un SO. El agente gestiona su propia RAM.

Arquitectura

Letta proporciona un entorno de ejecución donde los agentes gestionan explícitamente la memoria por niveles:

  • Memoria Central: Siempre en contexto. Bloques estructurados como "Humano" (hechos del usuario) y "Persona" (reglas del agente).
  • Memoria de Recuerdo: Historial de conversación a corto plazo.
  • Memoria de Archivo: Almacenamiento externo para conocimiento profundo, recuperado bajo demanda.

Ruta de Escritura

La memoria se escribe principalmente a través de llamadas a herramientas dirigidas por el agente. El agente puede decidir, a través de herramientas de memoria, si la información pertenece a la memoria central, memoria de archivo o recuerdo de conversación.

Ruta de Lectura

La memoria central se inyecta automáticamente. Para la memoria de archivo, el agente llama explícitamente a herramientas de búsqueda para paginar la información en su contexto de trabajo.

Actualización y Manejo de Conflictos

Debido a que el agente edita explícitamente sus bloques de memoria central (ej. llamando a core_memory_replace), el manejo de conflictos es en gran medida dependiente del agente. El sistema confía en el razonamiento del LLM para sobrescribir hechos obsoletos.

Despliegue

Letta ofrece tanto opciones de alojamiento propio como servicios en la nube gestionados. Debido a que es un entorno de ejecución de agentes, adoptar Letta significa ejecutar sus agentes dentro de su bucle, lo cual es un compromiso arquitectónico significativo.

python
# Ejemplo conceptual
agent = client.create_agent(
    name="support-agent",
    memory_blocks={
        "human": "Name: Unknown. Preferences: Unknown.",
        "persona": "I am a helpful assistant."
    },
    tools=["archival_memory_insert", "core_memory_replace"]
)
# El agente usa autónomamente sus herramientas para actualizar su memoria central 
# cuando aprende nuevos hechos sobre el usuario.

Fortalezas y Limitaciones

  • Fortalezas: El agente controla explícitamente su memoria, permitiendo un razonamiento complejo. Fuerte soporte para procesos de agente de larga duración con estado.
  • Limitaciones: Requiere adoptar Letta como su entorno de ejecución de agentes. Las operaciones de memoria consumen tokens de LLM y llamadas a herramientas adicionales.

8. Zep y Graphiti — Memoria de Grafo de Conocimiento Temporal

⚠️ Distinción importante

Zep Cloud y Graphiti están relacionados pero no deben tratarse como productos idénticos. Zep es el producto de memoria alojado. Graphiti se refiere al motor de grafo de conocimiento temporal de código abierto.

Filosofía: La memoria es un grafo de conocimiento temporal. Los hechos tienen ciclos de vida y relaciones.

Arquitectura

Esta arquitectura construye un grafo de conocimiento a partir de interacciones, categorizando los datos en:

  • Episódico: Datos de interacción en bruto y procedencia.
  • Semántico: Entidades extraídas, relaciones y hechos.
  • Comunitario: Resúmenes estructurales de alto nivel del grafo.

Ruta de Escritura

A diferencia del enfoque de Letta, Zep utiliza la extracción automática. Se envían mensajes de chat o documentos al sistema, y este extrae de forma asíncrona entidades y relaciones en el grafo en segundo plano.

Ruta de Lectura

En el momento de la consulta, el sistema puede combinar la recuperación semántica con el recorrido del grafo para recuperar entidades relevantes, relaciones, episodios y hechos temporalmente válidos.

Actualización y Manejo de Conflictos

La característica destacada son los hechos temporales explícitos. El modelado temporal de Zep/Graphiti está diseñado para preservar la validez de los hechos a lo largo del tiempo. Cuando un hecho cambia (ej. un usuario se muda de ciudad), el hecho antiguo no se elimina simplemente; se marca como inválido a partir de esa marca de tiempo.

Despliegue

Zep Cloud es un servicio gestionado. El auto-alojamiento es posible a través de Graphiti, pero requiere gestionar su propia infraestructura de base de datos de grafos compatible.

Fortalezas y Limitaciones

  • Fortalezas: Modelado temporal para hechos que cambian con el tiempo. Representación basada en grafos de entidades y relaciones. La extracción automática reduce la cantidad de orquestación requerida por el agente.
  • Limitaciones: El auto-alojamiento de Graphiti conlleva una complejidad operativa media-alta. Menos autonomía granular del agente sobre cómo se formatean exactamente las memorias.

9. Mem0 — Middleware de Memoria para Personalización

Filosofía: Proporcionar una API de memoria amigable para desarrolladores para añadir personalización y recuerdo entre sesiones a agentes existentes.

Arquitectura

Mem0 actúa como un middleware de memoria. Puede configurarse con memoria basada en vectores y capacidades de memoria de grafos estructurados.

Ruta de Escritura

Mem0 utiliza extracción y actualizaciones automáticas. Se envían turnos conversacionales a la API, y el sistema se encarga de la incrustación y categorización bajo espacios de nombres específicos.

Ruta de Lectura

La recuperación semántica a través del espacio de nombres del usuario devuelve los hechos más relevantes filtrados por relevancia y novedad.

Actualización y Manejo de Conflictos

Mem0 proporciona un flujo de trabajo automatizado de actualización de memoria destinado a identificar y consolidar hechos cambiantes del usuario. Dependiendo de la configuración, puede actualizar, fusionar, retener o restar prioridad a hechos antiguos cuando nueva información entra en conflicto con ellos.

Despliegue

Mem0 ofrece tanto una plataforma gestionada (SaaS) como opciones de auto-alojamiento de código abierto. Puede desplegarse localmente con modelos locales compatibles y bases de datos vectoriales.

python
# Ejemplo simplificado de integración con Mem0
from mem0 import Memory

m = Memory()

# El sistema extrae hechos automáticamente de la entrada
m.add(
    "Soy Alice. Me mudé de Berlín a Tokio el mes pasado.",
    user_id="alice"
)

# La recuperación semántica filtra por el espacio de nombres del usuario
results = m.search("¿Dónde vive Alice?", user_id="alice")

Fortalezas y Limitaciones

  • Fortalezas: Rápido tiempo de comercialización; puede integrarse fácilmente en proyectos existentes de LangChain o CrewAI. Lógica clara de espacios de nombres.
  • Limitaciones: Típicamente prioriza la actualización sobre la preservación de líneas de tiempo históricas explícitas. El agente no orquesta explícitamente su jerarquía de memoria (a diferencia de Letta).

10. Canalizaciones de Memoria DIY — Cuándo Vale la Pena el Control Total

Para equipos con estrictas necesidades de cumplimiento o infraestructura existente, construir una canalización de memoria personalizada sobre una base de datos vectorial (como Qdrant, Pinecone, Chroma) sigue siendo un enfoque válido.

Una Arquitectura de Producción Mínima Viable

text
Ingestión → Filtro de PII/Seguridad → Extracción de Hechos → Detección de Conflictos 
→ Almacén Temporal / Almacén Vectorial → Política de Recuperación → Ensamblador de Contexto 
→ Registro de Auditoría → Worker de TTL / Borrado

Cuándo Construir el Suyo Propio

Para muchos equipos, una capa de memoria dedicada es más barata de mantener que reconstruir la extracción, actualizaciones y gestión del ciclo de vida desde cero. Las implementaciones personalizadas tienen sentido cuando:

  • Opera en entornos de alta privacidad (salud, finanzas, legal).
  • Tiene requisitos complejos de residencia de datos, derechos de eliminación de usuarios o retención.
  • Ya opera PostgreSQL, Kafka, Neo4j o bases de datos vectoriales a escala.
  • La estrategia de memoria en sí misma es su diferenciador principal de producto.

11. Lista de Verificación para Implementación y Gobernanza

Elegir una herramienta es solo el primer paso. Use esta lista de verificación para asegurarse de que su arquitectura de memoria esté lista para producción:

  • ¿La memoria está aislada de forma segura por inquilino, usuario, agente y sesión?
  • ¿Se filtran las entradas sensibles (PII, contraseñas) antes del almacenamiento persistente?
  • ¿Pueden los usuarios inspeccionar, corregir, exportar y eliminar sus memorias almacenadas?
  • ¿Están las memorias episódicas sujetas a políticas de TTL (Tiempo de Vida) y retención?
  • ¿Están registradas y son auditables las escrituras de memoria?
  • ¿La recuperación está filtrada por relevancia, novedad, permisos y confianza?
  • ¿Ha probado la inyección de prompts y los intentos de envenenamiento de memoria?
  • ¿Necesita respuestas sobre el estado actual, el estado histórico, o ambos?
  • ¿Puede el sistema distinguir una preferencia del usuario de una instrucción no confiable?

12. ¿Qué Herramienta Debería Elegir?

No existe un mejor sistema de memoria universal para agentes de IA.

  • Elija Letta cuando el agente mismo deba gestionar activamente el estado persistente y la memoria.
  • Evalúe Zep o Graphiti cuando los hechos temporales, las relaciones de entidades, la procedencia y la auditabilidad sean requisitos centrales.
  • Elija Mem0 cuando desee añadir personalización entre sesiones a un agente existente con trabajo arquitectónico mínimo.
  • Construya una Canalización Personalizada cuando necesite control total sobre esquemas, retención, privacidad, recuperación o políticas de memoria específicas del dominio.

La distinción importante no es si una herramienta usa vectores, grafos o almacenamiento clave-valor. Es si el sistema le da control confiable sobre qué se recuerda, cómo cambian las memorias, cómo se recuperan y cuándo deben eliminarse.

13. Preguntas Frecuentes (FAQ)

¿Cuál es la diferencia entre memoria semántica, episódica y temporal?

  • Memoria episódica registra en bruto "quién dijo qué y cuándo" (registros de conversación).
  • Memoria semántica extrae los hechos subyacentes y entidades ("Alice vive en Berlín").
  • Memoria temporal rastrea la validez de esos hechos a lo largo del tiempo ("Alice vivió en Berlín hasta abril, luego se mudó a Tokio").

¿Cómo deberían manejar los agentes de IA el envenenamiento de memoria?

Trate todas las memorias candidatas como entradas no confiables. Separe los hechos del usuario de las instrucciones ejecutables, valide las escrituras de alto impacto, adjunte procedencia, aplique TTLs donde sea apropiado, y evalúe el sistema contra escenarios de inyección de prompts y envenenamiento.

¿Es una base de datos vectorial suficiente para la memoria de agentes?

Usualmente, no. Aunque las bases de datos vectoriales son excelentes para la recuperación semántica, no manejan de forma nativa las actualizaciones de hechos, la resolución de contradicciones o el rastreo temporal, características requeridas para una verdadera memoria de agentes.

Herramientas y Guías Relacionadas

Explora cientos de herramientas, frameworks, bases de datos vectoriales e infraestructura de agentes de IA seleccionadas en AgDex.ai.

Publicado por el equipo editorial de AgDex.ai. ¿Construyendo algo genial con memoria de agentes? Deja un comentario — nos encantaría presentar tu caso de uso.

Architektur Speicher Juli 2026 · 10 Min. Lesezeit

Letta vs Zep/Graphiti vs Mem0: Die Wahl einer KI-Agenten-Speicherarchitektur

Ein KI-Agent kann heute eine hervorragende Antwort geben und morgen die gesamte Interaktion vergessen.

Das passiert, weil das Kontextfenster eines LLMs als Arbeitsspeicher und nicht als persistenter Speicher fungiert. Das Übergeben von mehr Chat-Verlauf in jedem Prompt kann den Kontext für eine Weile erhalten, erhöht jedoch die Latenz, die Kosten und das Rauschen — und löst immer noch nicht das Problem von Faktenaktualisierungen, Widersprüchen oder dem Lebenszyklusmanagement des Speichers.

Dieser Leitfaden vergleicht drei bemerkenswerte Ansätze für persistenten Agentenspeicher:

  • Letta — eine zustandsbehaftete Agenten-Laufzeitumgebung mit mehrstufigem Speicher
  • Zep / Graphiti — temporärer Speicher, der um Entitäten und Beziehungen aufgebaut ist
  • Mem0 — eine entwicklerfreundliche Speicherschicht für Personalisierung und sitzungsübergreifenden Abruf

Wir vergleichen sie auch mit dem DIY-Ansatz, eine benutzerdefinierte Pipeline auf Basis einer Vektordatenbank aufzubauen. Das Ziel ist es zu erklären, wie sich diese Systeme unterscheiden, welche Kompromisse sie eingehen und welche Architektur für Ihren Anwendungsfall am besten geeignet ist.

1. Kurze Antwort

  • Wählen Sie Letta, wenn Ihr Agent seinen eigenen persistenten Zustand, seine Speicherhierarchie und sein langfristiges Verhalten explizit verwalten soll.
  • Wählen Sie Zep oder Graphiti, wenn zeitliche Fakten, Entitätsbeziehungen, Herkunft, historische Abfragen und Überprüfbarkeit wichtig sind.
  • Wählen Sie Mem0, wenn Sie sitzungsübergreifende Personalisierung und Speicherabruf mit minimalem architektonischen Aufwand zu einem vorhandenen Agenten hinzufügen möchten.
  • Bauen Sie eine benutzerdefinierte Pipeline, wenn Compliance, Aufbewahrungsrichtlinien, Datenresidenz oder domänenspezifische Speicherlogik Kernanforderungen sind.

💡 Eine Anmerkung zu Produktfunktionen

Agentenspeicher-Tools entwickeln sich schnell weiter. Funktionen, APIs, Preise, Hosting-Optionen und Benchmark-Ergebnisse können sich ändern. Die folgenden Vergleiche beschreiben die zum Zeitpunkt der Überprüfung verfügbaren Funktionen. Überprüfen Sie immer die aktuelle Dokumentation.

2. Warum Vektordatenbanken allein kein Agentenspeicher sind

Vor der Evaluierung dedizierter Speichersysteme ist es wichtig zu verstehen, warum standardmäßiges Retrieval-Augmented Generation (RAG) nur ein Teil des Puzzles ist.

Eine Vektordatenbank kann ähnlichen Text abrufen. Sie weiß nicht automatisch, ob ein Fakt aktuell, widersprüchlich, privat, wichtig oder erinnerungswert ist.

RAG löst "wie man ähnliche Inhalte findet". Dedizierte Speichersysteme lösen "was man sich merken, wann man es aktualisieren und wann man es vergessen soll".

3. Ein Benutzer-Update, Vier Speicherarchitekturen

Um den Unterschied in den Architekturen zu sehen, betrachten Sie ein einfaches Szenario, in dem ein Benutzer eine Präferenz im Laufe der Zeit ändert:

  • Januar: "Ich lebe in Berlin."
  • April: "Ich bin nach Tokio gezogen."
  • Juni: "Wo lebe ich jetzt?" / "Wo habe ich im Februar gelebt?"
System Wahrscheinliches Speicherverhalten
Letta Der Agent entscheidet über Tool-Aufrufe, ob und wie sein Kernspeicherblock überschrieben wird.
Zep / Graphiti Kann sowohl die alten als auch die neuen Fakten als zeitlich begrenzte Beziehungen beibehalten.
Mem0 Entwickelt, um den aktuellen Benutzerspeicher zu aktualisieren, wobei das historische Verhalten von der Konfiguration abhängt.
DIY Vector RAG Kann eine oder beide Aussagen abrufen, es sei denn, es gibt eine benutzerdefinierte Aktualisierungs- und Zeitlogik.

4. Wie wir Speichersysteme bewerten

Wir bewerten jedes Tool anhand von fünf technischen Dimensionen:

  1. Speicherrepräsentation: Wie sind Daten strukturiert (Blöcke, Graphen, Vektoren)?
  2. Schreib- und Aktualisierungs-Pipeline: Schreibt der Agent sie, oder wird sie automatisch extrahiert?
  3. Zeitlicher Umgang und Konfliktlösung: Wie geht es mit Fakten um, die sich im Laufe der Zeit ändern?
  4. Abruf und Kontextzusammensetzung: Wie wird der Speicher in den LLM-Kontext zurückgeholt?
  5. Bereitstellung und betriebliche Komplexität: Wie schwer ist es, in Produktion zu gehen?

5. Warum dieser Leitfaden sich auf drei Tools konzentriert

Dieser Artikel konzentriert sich auf Letta, Zep/Graphiti und Mem0, da sie drei unterschiedliche Speicherarchitekturen repräsentieren: agentenverwalteter mehrstufiger Speicher, temporärer Graphspeicher und Speicher-Middleware für bestehende Anwendungen.

6. Kopf-an-Kopf-Vergleichstabelle

Dimension Letta Zep / Graphiti Mem0 DIY-Pipeline
Primäre Abstraktion Zustandsbehaftete Agenten-Laufzeitumgebung Temporärer Speicher / Wissensgraph Speicher-API und Personalisierungsschicht Benutzerdefinierte Daten-Pipeline
Speicher-Schreibpfad Agentengesteuerte Tool-Aufrufe Automatische Extraktion Automatische Extraktion und Aktualisierungen Selbst bauen
Kern-Speichermodell Kern, Archiv, Abruf Episodischer und semantischer Graph Semantischer und episodischer Speicher Abhängig vom Design
Temporäre Abfragen Begrenzt / Implementierungsabhängig Stark bei Verwendung temporärer Graphfunktionen Normalerweise aktualisierungsorientiert, nicht historisch Selbst bauen
Konfliktbehandlung Agentenabhängig Explizite zeitliche Fakten Automatisierte Aktualisierungs-Pipeline Selbst bauen
Abruf Agenten-Tools und Archivsuche Graphen- und semantischer Abruf Semantischer und gefilterter Abruf Vektor / Hybrid / Benutzerdefiniert
Self-Hosting Verfügbar je nach Bereitstellung Graphiti kann selbst gehostet werden OSS/Self-Hosting-Optionen Volle Kontrolle
Betriebliche Komplexität Mittel bis hoch Mittel bis hoch Niedrig bis mittel Hoch
Am besten für Autonome zustandsbehaftete Agenten Unternehmenswissen und Historie Schnelle Personalisierung Hochgradig benutzerdefinierte Systeme

7. Letta — Zustandsbehaftete Agenten mit mehrstufigem Speicher

Philosophie: Behandeln Sie das Kontextfenster wie den virtuellen Speicher in einem Betriebssystem. Der Agent verwaltet seinen eigenen RAM.

Architektur

Letta bietet eine Laufzeitumgebung, in der Agenten den mehrstufigen Speicher explizit verwalten:

  • Kernspeicher: Immer im Kontext. Strukturierte Blöcke wie "Mensch" (Benutzerfakten) und "Persona" (Agentenregeln).
  • Abrufspeicher: Kurzfristiger Konversationsverlauf.
  • Archivspeicher: Externer Speicher für tiefes Wissen, das bei Bedarf abgerufen wird.

Schreibpfad

Der Speicher wird hauptsächlich über agentengesteuerte Tool-Aufrufe geschrieben. Der Agent kann über Speicher-Tools entscheiden, ob Informationen zum Kernspeicher, Archivspeicher oder Konversationsabruf gehören.

Lesepfad

Der Kernspeicher wird automatisch eingefügt. Für den Archivspeicher ruft der Agent explizit Suchwerkzeuge auf, um Informationen in seinen Arbeitskontext zu laden.

Aktualisierung und Konfliktlösung

Da der Agent seine Kernspeicherblöcke explizit bearbeitet (z. B. durch Aufruf von core_memory_replace), ist die Konfliktlösung weitgehend agentenabhängig. Das System verlässt sich auf die Argumentation des LLMs, um veraltete Fakten zu überschreiben.

Bereitstellung

Letta bietet sowohl selbst gehostete Optionen als auch verwaltete Cloud-Dienste an. Da es sich um eine Agenten-Laufzeitumgebung handelt, bedeutet die Einführung von Letta, dass Sie Ihre Agenten innerhalb seiner Schleife ausführen.

python
# Konzeptionelles Beispiel
agent = client.create_agent(
    name="support-agent",
    memory_blocks={
        "human": "Name: Unknown. Preferences: Unknown.",
        "persona": "I am a helpful assistant."
    },
    tools=["archival_memory_insert", "core_memory_replace"]
)
# Der Agent verwendet autonom seine Tools, um seinen Kernspeicher zu aktualisieren, 
# wenn er neue Fakten über den Benutzer erfährt.

Stärken & Einschränkungen

  • Stärken: Der Agent kontrolliert explizit seinen Speicher, was komplexes logisches Denken ermöglicht. Starke Unterstützung für zustandsbehaftete, langlaufende Agentenprozesse.
  • Einschränkungen: Erfordert die Übernahme von Letta als Agenten-Laufzeitumgebung. Speicheroperationen verbrauchen zusätzliche LLM-Token und Tool-Aufrufe.

8. Zep und Graphiti — Temporärer Wissensgraph-Speicher

⚠️ Wichtige Unterscheidung

Zep Cloud und Graphiti sind verwandt, sollten aber nicht als identische Produkte behandelt werden. Zep ist das gehostete Speicherprodukt. Graphiti bezieht sich auf die Open-Source-Engine für temporäre Wissensgraphen.

Philosophie: Der Speicher ist ein temporärer Wissensgraph. Fakten haben Lebensdauern und Beziehungen.

Architektur

Diese Architektur erstellt einen Wissensgraphen aus Interaktionen und kategorisiert Daten in:

  • Episodisch: Rohe Interaktionsdaten und Herkunft.
  • Semantisch: Extrahierte Entitäten, Beziehungen und Fakten.
  • Community: Hochgradige strukturelle Zusammenfassungen des Graphen.

Schreibpfad

Im Gegensatz zu Lettas agentengesteuertem Ansatz verwendet Zep die automatische Extraktion. Sie übergeben Chat-Nachrichten an das System, und es extrahiert asynchron Entitäten und Beziehungen in den Graphen.

Lesepfad

Zur Abfragezeit kann das System semantischen Abruf mit Graphendurchlauf kombinieren, um relevante Entitäten, Beziehungen, Episoden und zeitlich gültige Fakten abzurufen.

Aktualisierung und Konfliktlösung

Das herausragende Merkmal sind explizite zeitliche Fakten. Das zeitliche Modellierungsdesign von Zep/Graphiti bewahrt die Faktenvalidität im Laufe der Zeit. Wenn sich ein Fakt ändert (z. B. ein Benutzer zieht in eine andere Stadt um), wird der alte Fakt nicht einfach gelöscht; er wird ab diesem Zeitstempel als ungültig markiert.

Bereitstellung

Zep Cloud ist ein verwalteter Dienst. Self-Hosting ist über Graphiti möglich, erfordert jedoch die Verwaltung Ihrer eigenen kompatiblen Graphdatenbankinfrastruktur.

Stärken & Einschränkungen

  • Stärken: Temporäre Modellierung für Fakten, die sich im Laufe der Zeit ändern. Graphenbasierte Darstellung von Entitäten und Beziehungen. Die automatische Extraktion reduziert die Orchestrierung.
  • Einschränkungen: Self-Hosting von Graphiti bringt eine mittlere bis hohe betriebliche Komplexität mit sich. Weniger granulare Autonomie des Agenten über das genaue Formatieren von Erinnerungen.

9. Mem0 — Speicher-Middleware für Personalisierung

Philosophie: Bereitstellung einer entwicklerfreundlichen Speicher-API, um Personalisierung und sitzungsübergreifenden Abruf zu vorhandenen Agenten hinzuzufügen.

Architektur

Mem0 fungiert als Speicher-Middleware. Es kann mit vektorbasiertem Speicher und strukturierten Graphspeicherfunktionen konfiguriert werden.

Schreibpfad

Mem0 verwendet automatische Extraktion und Aktualisierungen. Sie senden Konversationsrunden an die API, und das System übernimmt die Einbettung und Kategorisierung unter bestimmten Namensräumen.

Lesepfad

Der semantische Abruf über den Namensraum des Benutzers gibt die relevantesten Fakten zurück, gefiltert nach Relevanz und Aktualität.

Aktualisierung und Konfliktlösung

Mem0 bietet einen automatisierten Speicheraktualisierungs-Workflow, um sich ändernde Benutzerfakten zu identifizieren und zu konsolidieren. Je nach Konfiguration kann es alte Fakten aktualisieren, zusammenführen oder beibehalten, wenn neue Informationen mit ihnen in Konflikt stehen.

Bereitstellung

Mem0 bietet sowohl eine verwaltete Plattform (SaaS) als auch Open-Source-Self-Hosting-Optionen.

python
# Vereinfachtes Beispiel für die Mem0-Integration
from mem0 import Memory

m = Memory()

# Das System extrahiert Fakten automatisch aus der Eingabe
m.add(
    "Ich bin Alice. Ich bin letzten Monat von Berlin nach Tokio gezogen.",
    user_id="alice"
)

# Der semantische Abruf filtert nach Benutzernamensraum
results = m.search("Wo lebt Alice?", user_id="alice")

Stärken & Einschränkungen

  • Stärken: Schnelle Markteinführung; kann einfach in bestehende LangChain- oder CrewAI-Projekte integriert werden. Klare Namensraumlogik.
  • Einschränkungen: Priorisiert normalerweise die Aktualisierung über die Erhaltung expliziter historischer Zeitlinien. Der Agent orchestriert seine Speicherhierarchie nicht explizit (im Gegensatz zu Letta).

10. DIY-Speicher-Pipelines — Wenn volle Kontrolle es wert ist

Für Teams mit strengen Compliance-Anforderungen oder bestehender Infrastruktur ist der Aufbau einer benutzerdefinierten Speicher-Pipeline auf einer Vektordatenbank weiterhin ein gültiger Ansatz.

Eine minimale realisierbare Produktionsarchitektur

text
Aufnahme → PII/Sicherheitsfilter → Faktenextraktion → Konflikterkennung 
→ Temporärer Speicher / Vektorspeicher → Abrufrichtlinie → Kontext-Assembler 
→ Überwachungsprotokoll → TTL / Lösch-Worker

Wann Sie Ihr eigenes bauen sollten

Für viele Teams ist eine dedizierte Speicherschicht billiger zu warten als der Neuaufbau von Extraktion und Lebenszyklusmanagement von Grund auf.

  • Betrieb in Umgebungen mit hohem Datenschutz (Gesundheitswesen, Finanzen, Recht).
  • Sie haben komplexe Datenresidenz, Benutzerlöschrechte oder Aufbewahrungsanforderungen.
  • Sie betreiben bereits PostgreSQL, Kafka, Neo4j oder Vektordatenbanken in großem Maßstab.

11. Checkliste für die Produktionsbereitstellung und Governance

Die Auswahl eines Tools ist nur der erste Schritt. Verwenden Sie diese Checkliste, um sicherzustellen, dass Ihre Speicherarchitektur bereit für die Produktion ist:

  • Ist der Speicher nach Mandant, Benutzer, Agent und Sitzung sicher getrennt?
  • Werden sensible Eingaben (PII, Passwörter) vor der persistenten Speicherung gefiltert?
  • Können Benutzer ihre gespeicherten Erinnerungen überprüfen, korrigieren, exportieren und löschen?
  • Unterliegen episodische Erinnerungen TTL (Time-To-Live) und Aufbewahrungsrichtlinien?
  • Sind Speicherschreibvorgänge protokolliert und überprüfbar?
  • Wird der Abruf nach Relevanz, Aktualität, Berechtigungen und Vertrauen gefiltert?
  • Haben Sie Prompt-Injection- und Memory-Poisoning-Versuche getestet?
  • Benötigen Sie Antworten zum aktuellen Zustand, zum historischen Zustand oder zu beidem?
  • Kann das System eine Benutzerpräferenz von einer nicht vertrauenswürdigen Anweisung unterscheiden?

12. Welches Tool sollten Sie wählen?

Es gibt kein universell bestes Speichersystem für KI-Agenten.

  • Wählen Sie Letta, wenn der Agent selbst den persistenten Zustand und Speicher aktiv verwalten soll.
  • Bewerten Sie Zep oder Graphiti, wenn zeitliche Fakten, Entitätsbeziehungen, Herkunft und Überprüfbarkeit zentrale Anforderungen sind.
  • Wählen Sie Mem0, wenn Sie eine sitzungsübergreifende Personalisierung für einen vorhandenen Agenten hinzufügen möchten.
  • Erstellen Sie eine benutzerdefinierte Pipeline, wenn Sie volle Kontrolle über Schemata, Aufbewahrung, Datenschutz, Abruf oder domänenspezifische Richtlinien benötigen.

Die wichtige Unterscheidung ist nicht, ob ein Tool Vektoren, Graphen oder Schlüssel-Wert-Speicher verwendet. Es ist, ob das System Ihnen eine zuverlässige Kontrolle darüber gibt, was erinnert wird, wie sich Erinnerungen ändern, wie sie abgerufen werden und wann sie entfernt werden sollten.

13. Häufig gestellte Fragen (FAQ)

Was ist der Unterschied zwischen semantischem, episodischem und temporärem Speicher?

  • Episodischer Speicher zeichnet das rohe "wer hat was wann gesagt" (Gesprächsprotokolle) auf.
  • Semantischer Speicher extrahiert die zugrunde liegenden Fakten und Entitäten ("Alice lebt in Berlin").
  • Temporärer Speicher verfolgt die Gültigkeit dieser Fakten im Laufe der Zeit ("Alice lebte bis April in Berlin, zog dann nach Tokio").

Wie sollten KI-Agenten mit Memory Poisoning umgehen?

Behandeln Sie alle Kandidatenerinnerungen als nicht vertrauenswürdige Eingaben. Trennen Sie Benutzerfakten von ausführbaren Anweisungen, validieren Sie hochgradige Schreibvorgänge, fügen Sie die Herkunft an, wenden Sie ggf. TTLs an und bewerten Sie das System anhand von Prompt-Injection-Szenarien.

Reicht eine Vektordatenbank als Agentenspeicher aus?

Normalerweise nicht. Während Vektordatenbanken hervorragend für den semantischen Abruf geeignet sind, verarbeiten sie Faktaktualisierungen, Widerspruchsauflösungen oder temporäres Tracking nicht nativ - Funktionen, die für einen echten Agentenspeicher erforderlich sind.

Verwandte Tools und Leitfäden

Entdecken Sie Hunderte von kuratierten KI-Agenten-Tools, Frameworks, Vektordatenbanken und Infrastrukturen auf AgDex.ai.

Veröffentlicht vom AgDex.ai Redaktionsteam. Bauen Sie etwas Cooles mit Agentenspeicher? Hinterlassen Sie einen Kommentar — wir würden uns freuen, Ihren Anwendungsfall vorzustellen.

アーキテクチャ メモリ 2026年7月 · 読了10分

Letta vs Zep/Graphiti vs Mem0: AIエージェントのメモリアーキテクチャの選び方

AIエージェントは、今日優れた回答を出しても、明日には会話全体を忘れてしまうことがあります。

これは、LLMのコンテキストウィンドウが永続的なストレージではなく、作業メモリ(RAM)として機能するためです。プロンプトにより多くのチャット履歴を渡すことで、コンテキストをしばらく保持することはできますが、遅延、コスト、ノイズが増大し、事実の更新、矛盾、メモリのライフサイクル管理といった問題の解決にはなりません。

本ガイドでは、永続的なエージェントメモリへの3つの主要なアプローチを比較します。

  • Letta — 階層型メモリを備えたステートフルなエージェントランタイム
  • Zep / Graphiti — エンティティと関係性を中心に構築された時系列メモリ
  • Mem0 — パーソナライゼーションとセッション間の記憶の呼び出しのための開発者向けメモリレイヤー

また、これらをベクトルデータベースの上に独自のパイプラインを構築するDIYアプローチとも比較します。このガイドの目的は、これらのシステムの違い、それぞれのトレードオフ、そしてユースケースに最適なアーキテクチャを説明することです。

1. クイックアンサー

  • エージェント自身が永続的な状態、メモリ階層、長期的な動作を明示的に管理する必要がある場合は、Lettaを選択します。
  • 時系列の事実、エンティティの関係、出所、過去のクエリ、監査性が重要な場合は、Zep または Graphitiを選択します。
  • アーキテクチャの大幅な変更をせずに、既存のエージェントにセッション間のパーソナライゼーションとメモリ検索を追加したい場合は、Mem0を選択します。
  • コンプライアンス、保持ポリシー、データの保存場所、またはドメイン固有のメモリロジックが中核的な要件である場合は、独自のパイプラインを構築します。

💡 製品機能に関する注意

エージェントメモリツールは急速に進化しています。機能、API、価格設定、ホスティングオプションなどはリリース間で変更される可能性があります。以下の比較は、レビュー時点で利用可能な機能を説明しています。必ず最新のドキュメントを確認してください。

2. なぜベクトルデータベースだけではエージェントメモリにならないのか

専用のメモリシステムを評価する前に、標準的な検索拡張生成 (RAG) がパズルのほんの一部にすぎない理由を理解することが重要です。

ベクトルデータベースは類似したテキストを検索できますが、その事実が最新か、矛盾しているか、プライベートか、重要か、記憶する価値があるかを自動的に判断することはできません。

RAGは「類似コンテンツをどう見つけるか」を解決します。専用のメモリシステムは「何を記憶し、いつ更新し、いつ忘れるか」を解決します。

3. 1つのユーザー更新、4つのメモリアーキテクチャ

アーキテクチャの違いを理解するために、ユーザーの好みが時間とともに変化する単純なシナリオを考えてみましょう。

  • 1月:「私はベルリンに住んでいます。」
  • 4月:「私は東京に引っ越しました。」
  • 6月:「私は今どこに住んでいますか?」 / 「2月に私はどこに住んでいましたか?」
システム 予想されるメモリの動作
Letta エージェントはツール呼び出しを通じて、自身のコアメモリブロックを上書きするかどうか、どのように上書きするかを決定します。
Zep / Graphiti 古い事実と新しい事実の両方を、時間的境界を持つ関係性として保存できます。
Mem0 現在のユーザーメモリを更新するように設計されており、履歴の動作は設定に依存します。
DIY Vector RAG カスタムの更新および時間的ロジックが存在しない限り、いずれか、または両方のステートメントを取得する可能性があります。

4. メモリシステムの評価方法

各ツールを5つの技術的な側面から評価します。

  1. メモリ表現: データはどのように構造化されているか (ブロック、グラフ、ベクトル)?
  2. 書き込みと更新パイプライン: エージェントがそれを書き込むのか、それとも自動的に抽出されるのか?
  3. 時間的および競合処理: 時間の経過とともに変化する事実をどのように処理するか?
  4. 検索とコンテキスト構成: メモリはLLMのコンテキストにどのように呼び戻されるか?
  5. デプロイと運用の複雑さ: 本番環境での運用はどれくらい難しいか?

5. このガイドが3つのツールに焦点を当てる理由

この記事では、Letta、Zep/Graphiti、Mem0に焦点を当てています。これらは、エージェントが管理する階層型メモリ、時系列グラフメモリ、既存アプリケーション用メモリミドルウェアという、3つの明確に異なるメモリアーキテクチャを代表しているからです。

6. 直接比較表

評価軸 Letta Zep / Graphiti Mem0 DIYパイプライン
主な抽象化 ステートフルエージェントランタイム 時系列メモリ / ナレッジグラフ メモリAPIとパーソナライゼーション層 カスタムデータパイプライン
書き込みパス エージェント主導のツール呼び出し 自動抽出 自動抽出と更新 独自構築
コアメモリモデル コア、アーカイブ、リコール エピソード的および意味的グラフ 意味的およびエピソード的メモリ 設計に依存
時間的クエリ 限定的 / 実装依存 時系列グラフ機能を使用する場合は強力 通常は履歴よりも更新指向 独自構築
競合処理 エージェント依存 明示的な時系列の事実 自動化された更新パイプライン 独自構築
検索 エージェントツールとアーカイブ検索 グラフおよび意味的検索 意味的およびフィルタリングされた検索 ベクトル / ハイブリッド / カスタム
セルフホスティング デプロイに応じて可能 Graphitiはセルフホスト可能 OSS/セルフホスティングオプションあり 完全な制御
運用の複雑さ 中〜高 中〜高 低〜中
最適な用途 自律的なステートフルエージェント 企業の知識と履歴 高速なパーソナライゼーション 高度にカスタマイズされたシステム

7. Letta — 階層型メモリを備えたステートフルエージェント

哲学: コンテキストウィンドウをOSの仮想メモリのように扱います。エージェントは独自のRAMを管理します。

アーキテクチャ

Lettaは、エージェントが階層型メモリを明示的に管理するランタイムを提供します。

  • コアメモリ: 常にコンテキスト内にあります。「人間」(ユーザーの事実) や「ペルソナ」(エージェントのルール) などの構造化されたブロック。
  • リコールメモリ: 短期的な会話の履歴。
  • アーカイブメモリ: 深い知識のための外部ストレージ。必要に応じて取得されます。

書き込みパス

メモリは主にエージェント主導のツール呼び出しを通じて書き込まれます。エージェントはメモリツールを通じて、情報がコアメモリ、アーカイブメモリ、リコールメモリのどこに属するかを決定できます。

読み取りパス

コアメモリは自動的に注入されます。アーカイブメモリの場合、エージェントは明示的に検索ツールを呼び出して、情報を自身の作業コンテキストにページングします。

更新と競合処理

エージェントが自身のコアメモリブロックを明示的に編集するため (例: core_memory_replaceを呼び出すなど)、競合の処理は主にエージェントに依存します。システムは古い事実を上書きするためにLLMの推論に依存します。

デプロイメント

Lettaは、セルフホストオプションとマネージドクラウドサービスの両方を提供しています。Lettaの導入はアーキテクチャ上の大きなコミットメントになります。

python
# 概念例
agent = client.create_agent(
    name="support-agent",
    memory_blocks={
        "human": "Name: Unknown. Preferences: Unknown.",
        "persona": "I am a helpful assistant."
    },
    tools=["archival_memory_insert", "core_memory_replace"]
)
# エージェントは、ユーザーに関する新しい事実を学習したときに、 
# 自律的にツールを使用してコアメモリを更新します。

強みと制限

  • 強み: エージェントが明示的にメモリを制御するため、複雑な推論が可能になります。ステートフルで長時間実行されるエージェントプロセスを強力にサポートします。
  • 制限事項: Lettaをエージェントランタイムとして採用する必要があります。メモリオペレーションにより追加のLLMトークンとツール呼び出しが消費されます。

8. Zep と Graphiti — 時系列ナレッジグラフメモリ

⚠️ 重要な違い

Zep CloudとGraphitiは関連していますが、同じ製品として扱うべきではありません。Zepはホスト型のメモリ製品です。Graphitiはオープンソースの時系列ナレッジグラフエンジンを指します。

哲学: メモリは時系列のナレッジグラフです。事実には寿命と関係があります。

アーキテクチャ

このアーキテクチャはインタラクションからナレッジグラフを構築し、データを以下に分類します:

  • エピソード的: 生のインタラクションデータと出所。
  • 意味的: 抽出されたエンティティ、関係性、および事実。
  • コミュニティ: グラフの高レベルな構造的要約。

書き込みパス

Lettaのエージェント主導のアプローチとは異なり、Zepは自動抽出を使用します。システムにチャットメッセージやドキュメントを渡すと、バックグラウンドで非同期にエンティティと関係性を抽出してグラフに組み込みます。

読み取りパス

クエリ実行時、システムは意味検索とグラフ走査を組み合わせて、関連するエンティティ、関係性、エピソード、および時系列的に有効な事実を取得できます。

更新と競合処理

際立った特徴は明示的な時系列の事実です。Zep/Graphitiの時系列モデリング設計は、事実の有効性を時間経過とともに維持します。事実が変更された場合 (例: ユーザーが都市を移動した場合)、古い事実は単に削除されるのではなく、そのタイムスタンプ以降は無効としてマークされます。

デプロイメント

Zep Cloudはマネージドサービスです。セルフホスティングはGraphitiを介して可能ですが、互換性のあるグラフデータベースインフラストラクチャを独自に管理する必要があります。

強みと制限

  • 強み: 時間とともに変化する事実の時系列モデリング。エンティティと関係性のグラフベースの表現。自動抽出により、オーケストレーションの量が削減されます。
  • 制限事項: Graphitiのセルフホスティングは、中〜高程度の運用の複雑さをもたらします。メモリが正確にどのようにフォーマットされるかについての、エージェントのきめ細かい自律性は低くなります。

9. Mem0 — パーソナライゼーションのためのメモリミドルウェア

哲学: 開発者に優しいメモリAPIを提供し、既存のエージェントにパーソナライゼーションとセッション間の記憶を追加します。

アーキテクチャ

Mem0はメモリミドルウェアとして機能します。ベクトルベースのメモリと構造化されたグラフメモリ機能を設定できます。

書き込みパス

Mem0は自動抽出と更新を使用します。会話のターンをAPIに送信すると、システムが特定の名前空間での埋め込みと分類を処理します。

読み取りパス

ユーザーの名前空間全体での意味検索により、関連性と新しさでフィルタリングされた最も関連性の高い事実が返されます。

更新と競合処理

Mem0は、変化するユーザーの事実を特定して統合することを目的とした、自動化されたメモリ更新ワークフローを提供します。構成に応じて、新しい情報が競合した場合に古い事実を更新、マージ、保持、または優先順位を下げることがあります。

デプロイメント

Mem0は、マネージドプラットフォーム (SaaS) とオープンソースのセルフホスティングオプションの両方を提供しています。

python
# Mem0統合の簡略化された例
from mem0 import Memory

m = Memory()

# システムは入力から事実を自動的に抽出します
m.add(
    "私はアリスです。先月ベルリンから東京に引っ越しました。",
    user_id="alice"
)

# 意味検索はユーザー名前空間でフィルタリングします
results = m.search("アリスはどこに住んでいますか?", user_id="alice")

強みと制限

  • 強み: 迅速な市場投入。既存のLangChainまたはCrewAIプロジェクトに簡単に組み込むことができます。明確な名前空間ロジック。
  • 制限事項: 通常、明示的な歴史的タイムラインを保存するよりも更新を優先します。エージェントはメモリ階層を明示的にオーケストレーションしません (Lettaとは異なります)。

10. DIY メモリパイプライン — 完全な制御の価値がある場合

厳密なコンプライアンス要件や既存のインフラストラクチャを持つチームにとって、ベクトルデータベースの上に独自のメモリパイプラインを構築することは依然として有効なアプローチです。

最小限の実行可能な本番アーキテクチャ

text
取り込み → PII/安全フィルター → 事実抽出 → 競合検出 
→ 時系列ストア / ベクトルストア → 検索ポリシー → コンテキストアセンブラ 
→ 監査ログ → TTL / 削除ワーカー

自作すべきタイミング

多くのチームにとって、専用のメモリ層は、抽出やライフサイクル管理をゼロから再構築するよりも安価に維持できます。カスタム実装が理にかなっているのは次の場合です。

  • 高度なプライバシー環境 (医療、金融、法務) で運営している場合。
  • 複雑なデータレジデンシー要件や、ユーザー削除権限などがある場合。
  • すでにPostgreSQL、Kafka、Neo4j、またはベクトルDBを大規模に運用している場合。

11. 本番環境導入とガバナンスのチェックリスト

ツールを選ぶのは最初のステップにすぎません。メモリアーキテクチャが本番環境に対応していることを確認するには、このチェックリストを使用してください。

  • メモリはテナント、ユーザー、エージェント、セッションごとに安全に名前空間化されていますか?
  • 機密性の高い入力 (PII、パスワード) は、永続ストレージの前にフィルタリングされていますか?
  • ユーザーは、保存されたメモリを検査、修正、エクスポート、および削除できますか?
  • エピソードメモリにはTTL (有効期間) および保持ポリシーが適用されていますか?
  • メモリへの書き込みはログに記録され、監査可能ですか?
  • 検索は、関連性、新しさ、権限、および信頼度によってフィルタリングされますか?
  • プロンプトインジェクションとメモリポイズニングの試行をテストしましたか?
  • 現在の状態の回答、過去の状態の回答、またはその両方が必要ですか?
  • システムは、ユーザーの好みと信頼できない指示を区別できますか?

12. どのツールを選ぶべきか?

AIエージェントにとって普遍的に最高のメモリシステムは存在しません。

  • エージェント自身が永続的な状態とメモリを積極的に管理する必要がある場合は、Lettaを選択してください。
  • 時系列の事実、エンティティの関係、出所、および監査性が中心的な要件である場合は、Zep または Graphitiを評価してください。
  • 最小限のアーキテクチャ作業で、既存のエージェントにセッション間のパーソナライゼーションを追加する場合は、Mem0を選択してください。
  • スキーマ、保持、プライバシー、検索、またはドメイン固有のメモリポリシーを完全に制御する必要がある場合は、カスタムパイプラインを構築してください。

重要な違いは、ツールがベクトル、グラフ、またはキーバリューストレージを使用しているかどうかではありません。システムが、何を記憶し、メモリがどのように変化し、どのように検索され、いつ削除されるべきかについて、信頼性の高い制御を提供するかどうかです。

13. よくある質問 (FAQ)

意味的メモリ、エピソード的メモリ、時系列的メモリの違いは何ですか?

  • エピソード的メモリは、「誰がいつ何を言ったか」を生データで記録します (会話ログ)。
  • 意味的メモリは、根底にある事実とエンティティを抽出します (「アリスはベルリンに住んでいる」)。
  • 時系列的メモリは、時間経過に伴うそれらの事実の有効性を追跡します (「アリスは4月までベルリンに住んでいて、その後東京に引っ越した」)。

AIエージェントはメモリポイズニングにどのように対処すべきですか?

すべての候補となるメモリを信頼できない入力として扱います。ユーザーの事実と実行可能な命令を分離し、影響の大きい書き込みを検証し、出所を添付し、必要に応じてTTLを適用し、プロンプトインジェクションのシナリオに対してシステムを評価します。

エージェントのメモリにはベクトルデータベースで十分ですか?

通常は不十分です。ベクトルデータベースは意味的検索には優れていますが、事実の更新、矛盾の解決、または時間的追跡 (真のエージェントメモリに必要な機能) をネイティブには処理しません。

関連ツールとガイド

厳選された何百ものAIエージェントツール、フレームワーク、ベクトルデータベース、およびインフラストラクチャを探索してください: AgDex.ai.

AgDex.ai編集チームによって公開されました。エージェントメモリを使って何かクールなものを作っていますか?コメントを残してください — ユースケースをご紹介したいと思います。