An agent that worked perfectly in a ten-minute demo starts making strange mistakes after an hour: it asks for a preference the user gave at the start, repeats a step it already finished, or loses track of a decision made forty tool calls ago. Nothing is broken. The agent has run into the limit every language model has: it only knows what is in its context window right now, and that window is finite.
The fix is not a bigger model or a longer prompt. It is giving the agent a memory it controls, and keeping its working context clean. This guide explains the three tools Claude's API offers for that in 2026, when to use each, and builds a real per-user memory store in Python that we tested line by line.
Why agents forget: the context window is not memory
Everything a model "knows" during a conversation sits in one place: the context window. Every user message, every tool call and every tool result is added to it. In a long agent run, tool results pile up fast: a web search returns pages, a file read returns a whole file, a database query returns hundreds of rows. Two problems follow.
- The window fills up. Once it is full, older content has to go. Whatever was in it is gone unless something saved it.
- Quality drops before it fills. Anthropic's own documentation describes context as "a finite resource with diminishing returns" and notes that irrelevant content degrades the model's focus. An agent dragging fifty stale search results around is worse at the current step, not just more expensive.
So "memory" really means two separate jobs: keep the working context small and relevant, and store what must survive somewhere outside it.
The three layers, and what each one is for
Claude's API now has a dedicated tool for each job. They solve different problems, and the mistake we see most often is using one where another belongs.
- Context editing: throw away stale tool results. With the strategy
clear_tool_uses_20250919(beta headercontext-management-2025-06-27), the API clears the oldest tool results once the conversation passes a threshold and replaces each one with placeholder text. By default it triggers at 100,000 input tokens and keeps the 3 most recent tool uses. Use it when your agent calls many tools whose output is only useful for a moment. - Compaction: summarize old conversation turns. Compaction replaces older turns with a summary written by the API, so a long conversation stays inside the window without you writing summarization code. You can trigger it yourself (on demand, currently behind the beta header
compact-2026-09-04) or let the API do it when input tokens reach a threshold you set. Use it for long conversations where the gist matters but the exact wording of turn 12 does not. - The memory tool: facts that must survive. The memory tool (
memory_20250818) lets Claude read and write files in a/memoriesdirectory that lives in your storage, not Anthropic's. Those files survive context clearing, compaction and even a brand-new conversation days later. Use it for user preferences, project decisions and task progress.
Anthropic's documentation recommends combining them for long-running agents: compaction "keeps the active context small", and memory "preserves the information that must survive summarization". Context editing goes one step further when paired with memory: as the conversation approaches the clearing threshold, Claude receives an automatic warning so it can save what matters to memory before those tool results disappear.
How the memory tool actually works
The most important thing to understand: Claude never touches your disk or database directly. The memory tool is a client-side tool. Claude sends a command, your code executes it against whatever storage you choose, and you send the result back. There are six commands:
view: list a directory or read a file (optionally only a line range)create: write a new filestr_replace: replace a piece of text in a fileinsert: insert text at a line numberdelete: delete a file or folderrename: rename or move a file
When the tool is enabled, the API automatically adds a short protocol to the system prompt. It tells Claude to "ALWAYS VIEW YOUR MEMORY DIRECTORY BEFORE DOING ANYTHING ELSE" and to "ASSUME INTERRUPTION", that is, to record progress as it goes because the context could be reset at any moment. That is why a well-built memory agent checks its notes first and writes them continuously, not just at the end.
The tool needs no beta header and works with Claude 4 and later models. The Python SDK ships two helpers: BetaLocalFilesystemMemoryTool, which stores memory as files in a local folder, and BetaAbstractMemoryTool, a base class you subclass to store memory anywhere else.
Our position: plain files beat a vector database for this job
Many "agent memory" tutorials go straight to embeddings and a vector database. For knowledge retrieval over thousands of documents, that is the right tool. For the memory an agent actually needs (what this user prefers, what was decided, where the task stopped), it is usually overkill. These facts are few, they change, and they need to be edited precisely, not approximately matched. A small set of readable files that Claude maintains itself is easier to inspect, easier to correct, and easier to delete when a user asks. Start with files; add retrieval only if the memory grows past what Claude can scan with view.
Build it: per-user memory in SQLite
The local-folder helper is fine on your laptop. In a real app with many users you need two things it does not give you: each user's memory must be isolated, and it should live in a database you already back up. So we subclass BetaAbstractMemoryTool and store files in SQLite, keyed by user. You need Python 3.10+ and the Anthropic SDK (we used version 1.11.0):
pip install anthropic
Create sqlite_memory.py:
import posixpath
import re
import sqlite3
from datetime import datetime, timezone
from anthropic.lib.tools import ToolError
from anthropic.tools import BetaAbstractMemoryTool
MAX_FILE_CHARS = 20_000 # one memory file can't grow without limit
SECRET_PATTERNS = [
re.compile(r"sk-[A-Za-z0-9_-]{20,}"), # API-key style secrets
re.compile(r"AKIA[0-9A-Z]{16}"), # AWS access key IDs
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY"), # private keys
]
class SQLiteMemory(BetaAbstractMemoryTool):
"""Claude's memory tool, stored in SQLite, one isolated space per user."""
def __init__(self, db_path: str, user_id: str):
super().__init__()
self.user_id = user_id
self.db = sqlite3.connect(db_path)
self.db.execute(
"CREATE TABLE IF NOT EXISTS memories ("
" user_id TEXT, path TEXT, content TEXT, updated_at TEXT,"
" PRIMARY KEY (user_id, path))"
)
# ---- safety -------------------------------------------------------
def _path(self, raw: str) -> str:
if "%" in raw or "\\" in raw or "\x00" in raw:
raise ToolError(f"Invalid characters in path: {raw}")
clean = posixpath.normpath(raw)
if clean != "/memories" and not clean.startswith("/memories/"):
raise ToolError(f"Path must stay inside /memories: {raw}")
return clean
def _check_text(self, text: str) -> None:
if len(text) > MAX_FILE_CHARS:
raise ToolError(f"Memory file too large (max {MAX_FILE_CHARS} characters)")
if any(p.search(text) for p in SECRET_PATTERNS):
raise ToolError("Refusing to store something that looks like a secret")
# ---- storage helpers ---------------------------------------------
def _get(self, path: str) -> str | None:
row = self.db.execute(
"SELECT content FROM memories WHERE user_id=? AND path=?", (self.user_id, path)
).fetchone()
return row[0] if row else None
def _put(self, path: str, text: str) -> None:
self._check_text(text)
self.db.execute(
"INSERT INTO memories VALUES (?,?,?,?) ON CONFLICT(user_id, path) "
"DO UPDATE SET content=excluded.content, updated_at=excluded.updated_at",
(self.user_id, path, text, datetime.now(timezone.utc).isoformat()),
)
self.db.commit()
def _children(self, folder: str) -> list[str]:
rows = self.db.execute(
"SELECT path FROM memories WHERE user_id=? AND path LIKE ? ORDER BY path",
(self.user_id, folder.rstrip("/") + "/%"),
).fetchall()
return [r[0] for r in rows]
# ---- the six commands Claude can send ------------------------------
def view(self, command):
path = self._path(command.path)
text = self._get(path)
if text is None:
files = self._children(path)
if path != "/memories" and not files:
raise ToolError(f"The path {command.path} does not exist")
return "\n".join(files) if files else "(empty)"
lines = text.split("\n")
start, end = 1, len(lines)
if command.view_range:
start, end = command.view_range
end = len(lines) if end == -1 else end
return "\n".join(f"{i}\t{lines[i - 1]}" for i in range(start, end + 1))
def create(self, command):
path = self._path(command.path)
self._put(path, command.file_text)
return f"File created: {path}"
def str_replace(self, command):
path = self._path(command.path)
text = self._get(path)
if text is None:
raise ToolError(f"The path {command.path} does not exist")
count = text.count(command.old_str)
if count != 1:
raise ToolError(f"old_str must appear exactly once, found {count} times")
self._put(path, text.replace(command.old_str, command.new_str or ""))
return f"File {path} edited"
def insert(self, command):
path = self._path(command.path)
text = self._get(path)
if text is None:
raise ToolError(f"The path {command.path} does not exist")
lines = text.split("\n")
if not 0 <= command.insert_line <= len(lines):
raise ToolError(f"insert_line must be between 0 and {len(lines)}")
lines.insert(command.insert_line, command.insert_text.rstrip("\n"))
self._put(path, "\n".join(lines))
return f"Text inserted in {path}"
def delete(self, command):
path = self._path(command.path)
if path == "/memories":
raise ToolError("Refusing to delete the whole memory directory")
cur = self.db.execute(
"DELETE FROM memories WHERE user_id=? AND (path=? OR path LIKE ?)",
(self.user_id, path, path + "/%"),
)
self.db.commit()
if cur.rowcount == 0:
raise ToolError(f"The path {command.path} does not exist")
return f"Deleted {path}"
def rename(self, command):
old, new = self._path(command.old_path), self._path(command.new_path)
text = self._get(old)
if text is None:
raise ToolError(f"The path {command.old_path} does not exist")
if self._get(new) is not None:
raise ToolError(f"The path {command.new_path} already exists")
self._put(new, text)
self.db.execute("DELETE FROM memories WHERE user_id=? AND path=?", (self.user_id, old))
self.db.commit()
return f"Renamed {old} to {new}"
The storage part is ordinary SQL. The parts worth reading twice are the guards, and every one of them comes from a risk Anthropic's documentation tells you to handle:
- Path traversal. A path like
/memories/../../etc/passwdmust never escape the memory area.posixpath.normpathcollapses the..segments, and we then require the result to start with/memories/. We also reject%outright, because the docs specifically warn about URL-encoded traversal such as%2e%2e%2f. - User isolation. Every query filters on
user_id. Sara's agent can't read Omar's notes, even if a prompt injection convinces it to try, because the boundary is in your code, not in the prompt. - Size caps. A memory file can't grow past 20,000 characters, so a runaway loop can't fill your database or blow up the next request's context.
- No secrets. The docs note that Claude usually refuses to write sensitive information, and recommend validating anyway. We block anything that looks like an API key or private key before it is stored.
When a method raises ToolError, the SDK's tool runner catches it and sends Claude a tool result marked is_error, with your message as the content. Claude reads the reason and adjusts, instead of your app crashing.
Test it without spending a single API call
Because the memory tool is just a handler for six commands, you can test it completely on its own, sending exactly the commands Claude would send. Create test_memory.py:
from anthropic.lib.tools import ToolError
from sqlite_memory import SQLiteMemory
def send(memory, **command):
"""Run one command exactly as Claude would send it, and print the result."""
try:
print("OK ", memory.call(command))
except ToolError as e:
print("ERROR", e)
sara = SQLiteMemory("memory.db", user_id="sara")
omar = SQLiteMemory("memory.db", user_id="omar")
send(sara, command="view", path="/memories")
send(sara, command="create", path="/memories/preferences.md",
file_text="- Prefers replies in Arabic\n- Timezone: Asia/Dubai")
send(sara, command="insert", path="/memories/preferences.md", insert_line=2,
insert_text="- Wants invoices as PDF")
send(sara, command="str_replace", path="/memories/preferences.md",
old_str="as PDF", new_str="as PDF, never as Word")
send(sara, command="view", path="/memories/preferences.md")
print("--- isolation: Omar cannot see Sara's memory")
send(omar, command="view", path="/memories")
print("--- attacks the handler must refuse")
send(sara, command="view", path="/memories/../../etc/passwd")
send(sara, command="view", path="/memories/%2e%2e/secrets")
send(sara, command="create", path="/memories/keys.md",
file_text="openai key: sk-proj-abcdefghijklmnopqrstuvwx")
send(sara, command="delete", path="/memories")
This is the output we got:
OK (empty)
OK File created: /memories/preferences.md
OK Text inserted in /memories/preferences.md
OK File /memories/preferences.md edited
OK 1 - Prefers replies in Arabic
2 - Timezone: Asia/Dubai
3 - Wants invoices as PDF, never as Word
--- isolation: Omar cannot see Sara's memory
OK (empty)
--- attacks the handler must refuse
ERROR Path must stay inside /memories: /memories/../../etc/passwd
ERROR Invalid characters in path: /memories/%2e%2e/secrets
ERROR Refusing to store something that looks like a secret
ERROR Refusing to delete the whole memory directory
Every line matches what we wanted: notes are created and edited precisely, a second user sees an empty memory, and all four attacks are refused with a clear reason. Run this test after every change to the handler; it takes under a second.
Connect it to Claude
With the handler tested, plugging it in takes a few lines. This follows the pattern in Anthropic's memory tool documentation, with our class in place of the local-folder helper. It needs an ANTHROPIC_API_KEY environment variable; keep that key out of your code.
import anthropic
from sqlite_memory import SQLiteMemory
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
memory = SQLiteMemory("memory.db", user_id="sara")
runner = client.beta.messages.tool_runner(
model="claude-opus-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "From now on, send my invoices as PDF."}],
tools=[memory],
)
print(runner.until_done().content)
The runner handles the loop: Claude views /memories, decides to write a note, your class stores it, and Claude confirms. Start a brand-new conversation for the same user_id and ask "How should I send your invoices?", and Claude will find the answer in its notes rather than asking again.
For long-running agents, add context editing next to the memory tool so stale tool results are cleared, with a warning to Claude first. This is the combination shown in the context editing documentation:
response = client.beta.messages.create(
model="claude-opus-5-5",
max_tokens=4096,
messages=[{"role": "user", "content": "Hello"}],
tools=[{"type": "memory_20250818", "name": "memory"}],
betas=["context-management-2025-06-27"],
context_management={"edits": [{"type": "clear_tool_uses_20250919"}]},
)
Five rules for memory that helps instead of hurts
- Tell the agent what is worth remembering. The built-in protocol makes Claude check and update memory, but it doesn't know your product. One line in your system prompt, such as "remember user preferences and decisions, not small talk", keeps notes short and useful.
- Store decisions, not transcripts. "Client approved the blue logo on 3 Oct" is memory. A copy of the whole conversation is not; that is what compaction is for.
- Let users see and delete their memory. Because memory lives in your database, you can show it on a settings page and delete it on request. That is both good practice and, in many places, a legal requirement for personal data.
- Expire what goes stale. The docs suggest periodically deleting files that haven't been accessed. A memory of last year's pricing can be worse than no memory.
- Treat memory as untrusted input. Notes written during one session are read as context in the next. If an attacker gets text into memory through a poisoned web page, it can steer future sessions. Keep memory writes narrow, and never let a note grant permissions.
Where to start
If your agent runs for minutes, you may not need any of this yet. If it runs for hours, or users come back across days, add the three layers in this order: context editing first (one parameter, immediate savings), then the memory tool with a tested handler like the one above, then compaction once single conversations grow long. Test the handler on its own before connecting it to a model; every bug you catch there is one you won't have to debug inside a 200-turn conversation.
If you haven't built a tool server yet, our guide to building an MCP server in 2026 covers the other half of a capable agent: giving it reliable tools.
Sources: Anthropic's documentation for the memory tool, context editing and compaction (platform.claude.com); the Anthropic Python SDK source (version 1.11.0). Code tested on 2026-10-05 with anthropic 1.11.0 and Python 3.13.