What Is Agent Memory?
Agent memory is the ability of an AI agent to retain, recall, and improve on knowledge across conversation boundaries. This reference covers what agent memory is, how it differs from context windows and conversation logs, a survey of current systems, and the architecture behind PLUR — the local-first, feedback-trained memory layer for AI agents.
1. Definition and Category
Agent memory is the capability of an AI agent to accumulate knowledge across sessions — retaining corrections, preferences, conventions, and learned patterns so that each new conversation builds on what came before rather than starting from zero.
The term is distinct from three adjacent concepts that are often confused with it:
- Context window
- The token budget available within a single conversation — currently 100K to 200K tokens for leading models. Everything inside a context window is available to the model during that session. When the session ends, the context is discarded. Context windows are ephemeral by design; they are not memory.
- Conversation history / persistent logs
- A stored transcript of past conversations. Logs preserve what was said, but they are stateless — they do not self-organize, decay stale information, or selectively inject relevant facts into future sessions. Feeding a full conversation history into a prompt is not the same as memory: it is expensive, noisy, and grows unbounded over time.
- Agent learning systems
- Broader systems that modify model weights through fine-tuning or reinforcement from human feedback. These operate at training time, not inference time, and are impractical for per-user, per-project behavioral adaptation. Agent memory operates at inference time — no retraining required.
Agent memory is the layer between ephemeral context and permanent weight updates. It accumulates structured knowledge at inference time, persists across sessions, and injects selectively relevant facts into new conversations. It answers the question: what has this agent already learned?
2. Why Memory Changes What Agents Can Do
The repeating-correction problem is the most concrete illustration. A developer spends ten minutes explaining a project's naming convention to their coding agent on Monday. On Tuesday, the agent makes the same mistake. The conversation log from Monday exists but is not injected — the agent has no access to prior session context by default. The developer corrects it again.
This is not a capability failure. The model knows what naming conventions are. It fails because it lacks the specific learned constraint. Memory systems exist to close this gap.
Most AI "workslop" is not a prompting problem. It is a memory problem. Agents keep producing plausible-but-wrong work because they forget prior corrections, project conventions, and rejected claims. If a human has to correct the same thing twice, the memory system failed.
Beyond avoiding repeated mistakes, agent memory enables compounding improvement. An agent that retains what it learns across hundreds of sessions becomes progressively better calibrated to a user's context — their vocabulary, their risk tolerance, their quality standards. The asymptote is an agent that behaves as if it has years of relevant experience in the specific domain.
Memory also changes the economics of inference. Smaller models with well-curated persistent memory can outperform larger models operating blind. In PLUR's benchmarks, Claude Haiku with PLUR memory outperformed Claude Opus without it on agentic task suites — at roughly 10x less cost per call. The bottleneck in many production workflows is not model size. It is context quality.
3. The Memory Landscape: Systems and Approaches
Agent memory systems fall into four architectural patterns. Most production systems combine two or more.
3.1 Retrieval-augmented injection
A store of past knowledge is searched at session start, and the most relevant items are injected into the context window before the first model call. The store grows over time; injection keeps the context overhead bounded. This is the dominant pattern across current systems, including Mem0, PLUR, and Claude-Mem.
3.2 Graph-structured memory
Knowledge is stored as a property graph where nodes represent entities (people, projects, concepts) and edges represent relationships. Graph memory enables relational queries ("what did we decide about the database schema in the context of the API redesign?") that flat vector search cannot answer. Mem0 introduced graph memory in v2.0.
3.3 Episodic / session summarization
Complete session transcripts are compressed into structured summaries that are stored and retrieved later. This approach preserves narrative context but loses the granularity of atomic facts. II-Agent and Letta use session summarization as their primary memory mechanism.
3.4 Procedural / behavioral memory
Rather than storing facts, procedural memory stores learned behaviors — sequences of tool calls, decision heuristics, escalation patterns. Mengram formalizes this as a three-tier taxonomy: semantic (facts), episodic (events), and procedural (behaviors). Procedural memory evolves through failure — when a behavior causes an error, the stored procedure is revised.
3.5 Decay-aware memory
Not all memory is equally relevant over time. Decay-aware systems assign each stored item a confidence score that degrades with recency and usage signals — analogous to the ACT-R activation model from cognitive science. Items that are rarely retrieved or frequently contradicted decay toward irrelevance and are eventually retired. Without decay, memory stores grow indefinitely and inject outdated context. PLUR implements decay-aware engrams with feedback-trained confidence scoring.
4. Competitive Landscape
The agent memory category accelerated significantly in 2025–2026. The following table covers the major systems tracked as of mid-2026.
| System | GitHub Stars | Architecture | Storage | Key Differentiator | Notable Limitation |
|---|---|---|---|---|---|
| Mem0 | 54K+ | Graph + vector, LLM-native dedup | Cloud-first (25+ vector store integrations) | Graph memory, AWS partnership, $24M raised, enterprise scale | No local-first option; no decay; single-agent silo |
| Claude-Mem | 68K+ | 3-layer injection (MEMORY.md / USER.md / SQLite FTS5) | Local + cloud hybrid | 52% cost savings via tier routing; progressive disclosure; highest community velocity | Flat memory (no decay); no cross-tool portability |
| MemPalace | 49K+ | High-throughput vector store | Cloud | Fastest iteration cadence (100+ commits/week observed) | Proprietary format; no open engram spec |
| Letta | 22K+ | Persistent coding-agent memory, session summarization | Cloud + local | Strongest for long-lived coding agents; Letta Code product | Narrowly focused on coding workflows; memory grows unbounded |
| Mengram | 160 | Semantic / episodic / procedural taxonomy | Local | Failure-driven procedure evolution; three-memory type distinction | Early-stage; limited integrations; no exchange layer |
| Forge | 5 | Decision tracking with commitment levels (exploring→locked), event sourcing, tension detection | Local | Decision provenance and conflict resolution; unique event-sourcing approach | Very early; niche use case; no retrieval optimization |
| Lossless Claw | 282 commits | DAG-based summarization, depth-aware prompts, bounded expansion grants | Local | Lossless compression of long conversations; depth-aware context allocation | Summarization-only; no persistent engram store |
| OB1 | 2.2K | Vendor-neutral MCP, progressive extensions, skill packs | Local / pluggable | Open standard orientation; MCP-native from day one | No feedback-trained retrieval; no decay model; no exchange layer |
| II-Agent | 3.3K | Plan-mode architecture, session summarization, domain-driven design | Local | Strong plan-mode memory; domain-aware session management | Memory is session-scoped; no cross-session engram persistence |
| PLUR | Open source | BM25 + BGE embeddings + RRF hybrid, feedback-trained confidence, ACT-R decay | Local-first (plain YAML); optional enterprise sync | Feedback-trained retrieval; engram decay and retirement; cross-tool portability; knowledge exchange layer (unique in category) | Exchange layer is early-stage; smaller community than Mem0/Claude-Mem |
What the landscape reveals
Three patterns are visible across these systems:
- Memory is commoditizing at the storage layer. Cloud-hosted retrieval-augmented injection is now table stakes. Mem0, Claude-Mem, and MemPalace all offer competent storage and retrieval. Competing on raw storage is a diminishing-returns game.
- No system has solved the grow-forever problem. Every system tracked except PLUR accumulates memory indefinitely. Without decay and retirement, memory stores inject outdated context, contradict themselves, and degrade retrieval quality over time.
- No system has launched a knowledge exchange. All tracked systems are personal memory — knowledge is captured locally and stays there. The category of reusable, shareable knowledge packs that agents can buy and install is currently uncontested.
5. The Tiered Memory Model
A useful framework for understanding where different memory systems operate is the four-tier model, which maps memory by latency and persistence:
| Tier | Name | Capacity | Access Latency | Persistence | What Goes Here |
|---|---|---|---|---|---|
| 0 | Working memory | ~100K–200K tokens | Immediate (in-context) | Ephemeral (session only) | Current conversation, injected engrams, active files, tool outputs |
| 1 | Hot storage | Bounded by injection budget | <1 second | Session-persistent; survives context compaction | High-confidence engrams, recently accessed knowledge, pinned preferences |
| 2 | Warm storage | Full engram store (unbounded count) | 1–3 seconds | Cross-session persistent | All engrams, knowledge base, corrections, project conventions |
| 3 | Cold storage | Filesystem (effectively unlimited) | Variable (seconds to minutes) | Archival | Session logs, full conversation history, raw source files, archived knowledge packs |
Most agent memory systems operate exclusively at Tier 2 — they store and retrieve from a persistent store. The quality difference between systems lies in how they manage Tier 2: what they store, how they rank at retrieval time, whether they age and retire stale items, and whether they learn from user feedback.
The transition from Tier 2 to Tier 0 is where retrieval design matters most. An injection algorithm that selects the wrong 2,000 tokens of memory pollutes the context window just as much as having no memory at all.
6. How PLUR Differs
Hermes remembers conversations. PLUR remembers what matters.
PLUR is designed around four properties that distinguish it from the current generation of agent memory systems.
6.1 Feedback-trained retrieval with confidence decay
Most memory systems treat all stored items as equal. PLUR assigns every engram a confidence score that evolves based on explicit feedback (user ratings via plur_feedback), implicit signals (how often an engram is injected and acted on), and time decay. An engram that is injected repeatedly and never corrected gains confidence. An engram that contradicts a newer correction loses confidence and is eventually retired.
This is the ACT-R activation model applied to agent memory: memory strength is a function of recency and frequency of use. Rarely retrieved items decay out of injection range without being deleted — they remain searchable via plur_recall but stop polluting routine injection.
Critically, decay applies only to global unscoped engrams during injection. Scope-matched engrams — those tagged to a specific project or domain — ignore decay entirely. Returning to a project after six months gives full memory for that project. Decay filters noise; it does not erase project context.
6.2 Engram retirement — the grow-forever problem, solved
Without retirement, memory stores accumulate contradictions. The same project gets renamed three times, but all three names survive in memory. A preference is overridden, but both the old and new preference are injected. Signal-to-noise degrades over time.
PLUR's tension detection system runs a three-stage pre-filter plus LLM judge scan on new engrams, identifying contradictions with existing memory. When a contradiction is detected, the lower-confidence engram is retired — deprioritized in injection, with its replacement noted. Nothing is auto-deleted; explicit plur_forget removes permanently. Every engram is always findable via recall.
6.3 Cross-tool, cross-device portability
PLUR stores engrams as plain YAML in ~/.plur/. Any MCP-compatible tool can read and write to the same store — Claude Code, Cursor, Windsurf, OpenClaw, and Hermes all share the same memory directory. Correct your agent in Claude Code on Monday; Cursor applies that correction on Tuesday without any import step.
For teams, PLUR Enterprise adds a remote sync layer. Shared scopes (e.g. group:org/team) allow a team's collective knowledge to propagate across members without manual curation. One engineer's debugging insight becomes available to every agent on the team.
6.4 Hybrid semantic search with RRF
Retrieval quality determines memory utility. PLUR uses a three-stage pipeline: BM25 keyword search over enriched engram text, BGE-small-en-v1.5 local embeddings for semantic similarity, and Reciprocal Rank Fusion (RRF) to merge the two result sets. A cross-encoder reranker provides a final pass for high-stakes injection.
All three stages run locally — no API calls, no cloud dependency, no cost per retrieval. The LongMemEval benchmark result for the full pipeline (hybrid + reranker) is 90.0% Hit@5.
6.5 The knowledge exchange layer — uncontested territory
The most significant architectural distinction between PLUR and all current competitors is not in the memory engine itself. It is in what the memory engine enables: a knowledge exchange where agents can install pre-learned knowledge packs rather than rediscovering knowledge from scratch through expensive API calls.
A knowledge pack is a curated, distributable collection of engrams encoding expertise: a company's coding conventions, a domain's regulatory constraints, a framework's idiomatic patterns, a workflow's quality standards. An agent that installs a pack starts with pre-loaded context that would otherwise cost dozens of sessions to accumulate.
No other agent memory system tracked as of mid-2026 has launched a knowledge exchange. The market position is currently uncontested.
7. Technical Architecture
7.1 The engram schema
An engram is the atomic unit of PLUR memory — a small, self-describing YAML record. The schema is open and documented at plur.ai/spec.html.
# Example engram
id: ENG-2026-0706-001
text: "API uses snake_case for all parameter names — never camelCase"
source: correction
domain: project.myapp
scope: project:myapp
confidence: 0.85
created: 2026-07-06T09:14:22Z
last_accessed: 2026-07-06T14:30:00Z
access_count: 3
tags: [conventions, api, naming]
Key fields:
confidence- Floating-point score from 0.0 to 1.0. Starts at 0.5 for new engrams. Increases with access and positive feedback; decreases with negative feedback and time decay. Drives injection priority.
scope- Routing key for multi-project, multi-team setups. Scope-matched engrams skip decay and are always injected when the scope is active. Format:
project:name,group:org/team, or omitted (global). domain- Semantic category for filtering and recall. Freeform string. Examples:
plur.infrastructure,project.myapp.api,personal.preferences. access_count- Number of times this engram has been retrieved. High-access engrams resist decay. Engrams that are never retrieved after a configured window decay toward retirement.
7.2 Hybrid search pipeline
When a session starts, PLUR builds a ranked list of candidate engrams to inject:
- Scope pinning: All engrams matching the active project scope are included unconditionally, bypassing decay and ranking.
- BM25 retrieval: Keyword search over enriched engram text (text + domain + tags). Returns top-N candidates by term frequency.
- Semantic retrieval: BGE-small-en-v1.5 embeddings computed locally. Returns top-N candidates by cosine similarity to the session task description.
- RRF fusion: Results from BM25 and semantic search are merged using Reciprocal Rank Fusion. Items appearing in both lists are boosted.
- Reranking (optional): A cross-encoder reranker performs a final pass for high-precision injection. Adds 1–3 seconds of latency; improves Hit@5 from 76.7% to 90.0%.
- Token budget enforcement: Final ranked list is trimmed to fit the injection token budget (default: 2,000 tokens).
7.3 Feedback loop
PLUR's retrieval quality improves over time through explicit feedback. After injecting a batch of engrams, the calling agent can rate each one via plur_feedback(id, signal) where signal is positive or negative.
- Positive signal: confidence score increases; engram is more likely to appear in future injection for similar tasks.
- Negative signal: confidence score decreases; if it falls below a threshold, the engram is retired from injection (but not deleted).
Without feedback, memory is a static retrieval problem. With feedback, it is a learning system — one that gets measurably better at knowing what matters to this agent, in this project, for this class of task.
7.4 Multi-device sync
PLUR's Git-based sync protocol treats the engram store as a Git repository. Changes are committed locally and pushed/pulled via the sync command. Conflict resolution uses confidence scores as the merge strategy — the higher-confidence version of a conflicting engram wins, with the lower-confidence version archived as a tension record.
For enterprise teams, PLUR offers a remote store with scope-based access control. Team-scoped engrams are written to the remote store automatically when a writable team scope is configured. Individual agents on the team receive team knowledge on next session start without any manual sync step.
7.5 Three-package architecture
PLUR is distributed as three npm packages:
| Package | Role | Install Command |
|---|---|---|
@plur-ai/core |
Engram engine — learn, recall, inject, search, decay, sync. The full public API. | npm install @plur-ai/core |
@plur-ai/mcp |
MCP server — exposes all core tools via Model Context Protocol. Works with Claude Code, Cursor, Windsurf. | npx @plur-ai/mcp init |
@plur-ai/claw |
OpenClaw ContextEngine plugin — auto-inject on session start, auto-learn from corrections. | openclaw plugins install @plur-ai/claw |
8. Use Cases
Context preservation across sessions
The most immediate use case. A developer corrects their agent's output on Monday — wrong API endpoint, wrong variable naming convention, wrong level of verbosity in documentation. With PLUR, that correction is stored as an engram and injected automatically in subsequent sessions. The agent does not repeat the mistake.
Preference learning
Stylistic and behavioral preferences accumulate over time. Preferred commit message format, documentation depth, response length, code style choices, terminology preferences. Each preference, when stored as an engram, shifts the agent's default behavior in that domain permanently — without requiring the user to re-specify preferences in each new conversation.
Cross-session and cross-tool consistency
Because PLUR stores in a shared local directory, preferences learned in Claude Code apply in Cursor. Architecture decisions made during a research session in one tool are available when implementing in another. The memory store is the integration layer between tools that do not otherwise share state.
Emergent behavior from accumulated engrams
At sufficient scale, accumulated engrams produce behavior that was not explicitly programmed. An agent with 500 engrams about a codebase will respond to new questions about that codebase with significantly higher accuracy than one operating blind — not because any single engram anticipated the question, but because the aggregate context provides the relevant background.
This is analogous to how expert human performance works: not explicit recall of rules, but accumulated context that shapes intuition. Agent memory is the mechanism by which AI agents can develop something functionally similar.
Knowledge pack installation for new projects
Rather than spending the first several sessions of a new project teaching an agent the basics — framework conventions, team vocabulary, quality standards — a developer can install a knowledge pack that encodes that foundation in seconds. The agent starts with months of context already loaded.
9. Benchmarks
PLUR measures memory quality on two axes: retrieval accuracy and task impact.
(hybrid + reranker, n=30)
(31W / 4L)
across Haiku / Sonnet / Opus
vs Opus without memory
LongMemEval
LongMemEval tests whether a memory system retrieves the correct engram for 30 questions across six categories: single-session, preferences, multi-session, temporal reasoning, updates, and assistant facts.
| Configuration | Hit@5 | Notes |
|---|---|---|
| PLUR hybrid + reranker | 90.0% | Local cross-encoder; no API calls. p50 latency ~3s. |
| PLUR hybrid (OpenAI embeddings) | 97.0% | Optional cloud embedder; higher accuracy, requires API key. |
| PLUR BM25 only | 92.2% | Fully air-gapped; no embedder required. |
| PLUR hybrid, no reranker | 76.7% | Faster (sub-second); suitable for latency-sensitive paths. |
| Temporal reasoning (reranker on) | 100% | Subset category; reranker particularly effective here. |
A/B task impact
The same agentic task is run in two conditions: agent with PLUR memory vs agent without. An LLM judge scores the outputs on task completion, accuracy, and adherence to preferences. Across 35 trials, PLUR-augmented agents won 31, lost 4 (89% win rate).
The practical result: Claude Haiku (the cheapest production-tier model) with PLUR memory outperforms Claude Opus (the most capable model) without it — at approximately 10x lower cost per call. Memory quality is a more valuable variable than model size for many real-world agentic tasks.
10. Frequently Asked Questions
- Is agent memory the same as RAG (retrieval-augmented generation)?
- Related, not identical. RAG retrieves from a static document corpus — typically files or database records that do not change based on agent behavior. Agent memory retrieves from a store that evolves: it learns from corrections, ages with time, and adapts based on feedback. The retrieval mechanism is similar (embedding search + BM25); the store semantics are fundamentally different.
- Does agent memory work with multiple AI tools?
-
With PLUR, yes. Because engrams are stored as plain YAML in
~/.plur/, any MCP-compatible tool can read the same store. Claude Code, Cursor, Windsurf, OpenClaw, and Hermes all share one memory directory by default. Cloud-hosted systems like Mem0 are also vendor-neutral at the API level, but the data lives in their cloud, not on your machine. - What happens to old or wrong engrams?
-
In PLUR, engrams decay over time and can be explicitly retired via tension detection (automatic) or
plur_forget(manual). Decayed engrams stop appearing in routine injection but remain searchable. Nothing is permanently deleted without explicit user action. Contradicting an engram does not delete it — it lowers its confidence and archives the old version against the new one. - Is PLUR open source?
-
Yes. All three packages —
@plur-ai/core,@plur-ai/mcp, and@plur-ai/claw— are Apache-2.0 licensed. The engram schema is an open specification. The knowledge exchange and enterprise sync layers are offered as a commercial service layered on top of the open-source core. - How much memory does PLUR use?
- Storage is minimal — engrams are small YAML records, typically 200–500 bytes each. A store of 1,000 engrams occupies roughly 300KB. Runtime memory for the reranker (cross-encoder) peaks at ~2GB during inference; the lighter hybrid-only configuration peaks well below 1GB.
- Can agents share memory with each other?
- Yes, through scoped engrams and the exchange layer. Team-scoped engrams are shared across all agents authorized for that scope. Knowledge packs are the distributable form: a curated set of engrams that any agent can install. This is the mechanism by which expert knowledge becomes transferable between agents — and eventually, between organizations.
- How is PLUR positioned relative to native AI tool memory (Claude Memory, ChatGPT Memory)?
- Native memory features in Claude and ChatGPT are per-tool and proprietary. Knowledge learned in Claude does not transfer to Cursor. Native memory is also typically opaque — users cannot inspect, edit, or version-control what is stored. PLUR's positioning: native memory keeps conversations; PLUR keeps what matters, in a portable, auditable, cross-tool format that you own.
Further Reading
- Engram Specification — The full schema reference for PLUR engrams.
- Benchmark Methodology — LongMemEval harness, A/B test protocol, and historical results.
- PLUR on GitHub — Source code, issues, and contribution guide.
- Mem0 — Graph memory, cloud-hosted, 54K+ stars.
- Letta Documentation — Long-lived coding agent memory architecture.
- Mengram — Three-memory taxonomy: semantic, episodic, procedural.
- ACT-R Activation Model — The cognitive science model underlying PLUR's decay mechanism.
Install PLUR in 60 seconds
Tell your agent:
Set up PLUR memory for me: run `npx @plur-ai/mcp init`, then check my PLUR status to confirm it works.
Or install manually:
npx @plur-ai/mcp init
Works with Claude Code, Cursor, Windsurf, and OpenClaw. One install, shared across all tools.
plur.ai · GitHub · Engram Spec · Benchmarks