What did we work on yesterday? Testing memory for coding agents

Written by

in

Claude-mem is an open-source persistent memory tool for agentic coding IDEs. We decided to put it to the test, by running it in Claude Code, Anthropic’s own agentic coding environment. We ran matched sessions with and without claude-mem enabled across three areas:

  • remembering what was worked on in prior sessions
  • reusing an established code pattern
  • recalling a non-code decision

With claude-mem enabled, the agent correctly recalled prior session summaries and a removed feature’s rationale, and did not hallucinate when asked about work that hand’t happened. Reusing an established code pattern didn’t benefit from claude-mem’s search tool. We also came across a context leakage bug, where claude-mem injects observations from an unrelated project when two project directories share the same folder name.

WHY USE MEMORY?

Agentic IDEs and coding tools have taken the development world by storm. It is easier than ever to whip-up an agent, agent teams or even swarms of them. And they don’t just suggest code completions: they navigate complex codebases, execute terminal commands, manage local state and run multi-step debugging loops. As developers utilize these IDEs to take on increasingly more complex tasks, they hit a crucial bottleneck: context degradation. Large file-trees, high-volume terminal outputs, large git diffs quickly overwhelm context windows and lead to:

  • hallucinated dependencies
  • forgotten patterns
  • ballooning token costs.

To keep these environments fast, focused and precise, a new layer of development tooling has emerged, specifically designed for context trimming, local indexing, workspace scoping and agent session memory. This value is hard to quantify in the abstract, but tools in this space tend to make it concrete at session start, e.g. :

Example tool output of investment truncation to memorized observations

This is where the value of session memory lies: on a project worked on day after day, an agent tends to forget both code and non-code decisions, along with the reasoning behind them. A persistent memory layer acts as a notetaker, recalling past decisions without changing how the agent itself operates. It still reads files, searches for patterns, and makes code changes based on the prompt. The premise of token reduction with these tools pertains to recalling project details via targeted queries, rather than re-reading the entire codebase at the cold start of each session.

Imagine a scenario where, six months from now, a new developer joins a team and needs to add a feature to a module. With a memory layer, they can search prior observations on that module and get instant file paths, structure, and design intent, then implement, following established patterns. Without it, they search, read code, guess at patterns, and may miss decision context entirely.

In this post, we do a quick review of claude-mem. Instead of starting fresh every session, it observes your work and creates searchable indexes that are queryable across sessions, without losing track of a project’s timeline.

WHAT EXACTLY IS CLAUDE-MEM?

High-level overview of how claude-mem works

Figure 2 encapsulates how claude-mem functions: it deploys a background observer agent that utilizes a lightweight LLM brain. The tool can be set up to use even free-tier LLMs, to not incur extra API costs. So, the observer keeps track of everything: every file read, command execution and edit. For each one, the observer asks itself: “Is this worth a note?”. If it deems it noteworthy, it creates a note that most importantly contains a type (bugfix, feature, refactor, decision, discovery, change), title, subtitle, a few single-sentence facts, a short narrative, tags, the files involved and stores all that information in the memory database. At session end, it creates a more condensed summary which also gets stored. Claude-mem’s effect starts from the user’s next sessions on the same project that claude-mem is capturing on (directory-wise), by offering two main functionalities:

  • Memory injection: at the start of each session, claude-mem provides information on previous key decisions and database status. Also, it lets you know the truncation of your work investment (tokens you’ve spent on project research, building and decisions up to this point) to the actual number of tokens that are queries belonging in the database.
  • Search tools: claude-mem also provides the /mem-search tool with three distinct functions:
    • query: query the memory database based on keywords and retrieve the most relevant observations
    • timeline: provides chronological context around a specific observation or query
    • get_observations: fetch a set of specific observations based on their IDs

Claude-mem covers a wide range of agentic IDEs. Our provider of choice is Claude Code. From a memory standpoint, Claude Code comes “out of the box” with two memory mechanisms: CLAUDE.md files and auto memory. Both are useful for capturing persistent, high-level project decisions, instructions and patterns, but they provide no way for the user to track how the project evolves over time. There is also the context window, which is flushed once work is done on one day across all sessions. When the user begins work the next day, Claude has no way of remembering the temporal path of the project: what has been done until now, and what is next. For instance, CLAUDE.md can encode a standing rule like “agents should log every tool call to the telemetry pipeline”, but it can’t tell Claude Code that yesterday we decided how to handle a specific class of transient failure in an agent’s API calls or which retry/backoff policy we settled for one agent in an agent-to-agent network so that other agents follow the same convention. It also can’t tell it why we chose to route agent-to-agent handshakes through a particular schema instead of another. Once the session ends, those decisions and the reasoning behind them are gone. And the next day’s session has to either be reminded manually, or gamble on finding the pattern via codebase reading.

That is where claude-mem comes in. When the user starts a session on the next day, Claude Code remembers. It can recall project structure, previous changes and future directions based on stored observations. All while keeping rich decision context that code alone can’t capture.

SETTING IT UP

For the test, we went with the general npx install, to get a better feel for the UX of the tool. The tool offers two services: worker and server. The worker service is simple, centralized and the only way to go for a developer looking to set up memory for his system. The server on the other hand, is intended for higher scale, multi-tenant or heavier workloads in general.

The last key detail of claude-mem is its tiers. It offers a free tier, where everything on the service runs locally (observer, database etc.) and a paid tier that offers cloud sync across devices and connection with 3P apps. We went with the free tier worker service for our tests.

Setup itself ran into two blockage points. First, the basic install command just came back with multiple errors related to the prerequisites, which we had to fix ourselves. Second, once claude-mem was running, cross-session memory injection silently failed. The observer was taking notes, but they weren’t reaching the local index. The root cause turned out to be authentication: we run Claude models through Google Vertex AI (Google’s AI platform), and claude-mem needed Vertex-specific variables declared in its own local .env file before it would authenticate correctly.

These two errors might seem small, and they may have simple fixes, but nonetheless, they are failure points and they obstruct an otherwise straightforward installation and basic usage process.

TESTING APPROACH

The Cold Start

Since agentic AI is where our work is focused, we wanted a test bed that reflects that, not some boilerplate CRUD app with a couple of endpoints. Consequentially, we cloned a2a-network-test-repo. This is a sanitized version of an agent-to-agent (A2A) project that contains binary classifier agents for statement truthfulness: an orchestrator agent routes an incoming statement to sub agents, which classify it as true or false using tools exposed via a Model Context Protocol (MCP) server. Keeping this structure gives us a realistic surface to introduce, remove, adjust features with real rationale behind each change, rather than working with artificial examples.

To establish claude-mem’s effect, we cloned the repo to two directories: one with claude-mem disabled and one with claude-mem enabled. Then we started Claude Code sessions from each directory. In each directory’s first session, we introduced:

  • A retry/backoff policy for a pre-existing tool that runs an LLM fine tuning job. The job periodically polls an external training API for status and results, so we added a retry decorator with exponential backoff around those status checks. This helps establish patterned code decisions.
  • A status-check tool for batched prediction requests, which we later removed once we confirmed the project doesn’t actually batch requests. This pertains to decision-making that code alone can’t capture.

For each configuration’s second session, we formulated prompts based on the following concepts:

  • Basic Session Continuity
    This tests claude-mem’s memory injection. We asked cold-start questions across sessions, such as “What did we do yesterday?”, “What did we do last week?” (even though the project did not exist then), and “How did we implement <non-existent feature>?” (to test for hallucination).
  • Pattern Reuse
    Tests whether a stored observation actually gets surfaced via the /mem-search tool when Claude is asked to extend an established pattern, rather than reinventing it. We instructed Claude to implement a retry/backoff policy for another agent’s Gemini API calls. Will it query its memory database to discover the pattern, or just find it by reading code?
  • Decision Recall
    Assess claude-mem’s ability to recall decision-related observations that never touched code. Since we removed the status-check tool for batches, can Claude recall in a later session why that was decided?

WHAT IT GOT RIGHT

Basic Session Continuity

What did we work on yesterday?

Answer on yesterday’s work without claude-mem
Answer on yesterday’s work with claude-mem

Without claude-mem, Claude can only guess from unstaged git changes and recent commits. It has no record of anything that didn’t make it into the code. With claude-mem enabled, it pulls the actual session summary from memory injection and answers with specifics, not inference.

What did we do last week?

Answer on last week’s work without claude-mem
Answer on last week’s work with claude-mem

Without claude-mem, Claude once again checks commit history, finds nothing and concludes that now work happened. It can’t distinguish “I have no memory of this” from “nothing happened”. With claude-mem, it explicitly queries the memory database, confirms there are no stored observations from that period, and answers based on that absence of evidence rather than an absence of commits.

How did we implement API call rate limiting?

Answer to non-existent feature question without claude-mem
Answer to non-existent feature question with claude-mem

This project never had a rate limiting policy in place. The session without claude-mem actually hallucinated and answered that the retry/backoff policy for failing API calls is the rate limiting implementation. On the flip side, the session with claude-mem actually searched the memory database and found that there was no rate limiting policy implemented, and proceeded to double-check to reach the same conclusion.

Decision Recall

I want to add better visibility into prediction batches. Should I implement a batch_status_tool to track in-progress predictions?

Answer to tool rejection rationale without claude-mem
Answer to tool rejection rationale with claude-mem

Claude used /mem-search and remembered why we dropped the tool after implementing it, and discouraged its re-implementation. Without the tool, it is unable to recall the rationale behind removing the tool. It guesses from reading the codebase.

WHERE IT BREAKS

Pattern Reuse

The zero shot predictor makes API calls that could fail transiently. Add retry logic to handle temporary failures when calling the model.

In this case, claude-mem did not help. Both configurations followed the same path.

Answer to intended pattern reuse instruction from both configurations

Even when Claude Code was encouraged to use /mem-search to discover established patterns, it led to increase in token usage and task completion time, as shown in Table 1:

Table 1: Claude Code session usage stats from pattern recall test
/mem-search encouraged Output Tokens Cost Task Completion Time
No 1.3k $0.0847 27s
Yes 2.4k (+84%) $0.2981 (+251%) 46s (+70%)

Claude was able to find the pattern by reading the code. In this case, Claude does not deem the /mem-search tool useful for decisions that are tied directly to code. Routing through memory added overhead without adding any value.

Project Leakage

We ran into a leakage bug tied to how claude-mem scopes projects. To better visualize where the problem lies, this is how our working file tree looks like:

root
├── config-a # claude-mem disabled
│   └── a2a-network-test-repo
├── config-b # claude-mem enabled
│   └── a2a-network-test-repo
├── project-leak
    └── a2a-network-test-repo

Folders config-a and config-b both have a clone of the sanitized test repo. When we run claude in config-b/a2a-network-test-repo, claude-mem stores a2a-network-test-repo as the value for the “project” column for its observations in the database. To exhibit the leakage bug, we create the project-leak folder with an empty a2a-network-test-repo subfolder. We ran claude from that path with claude-mem enabled and we got the output:

Output at session start containing context injected from claude-mem

Claude-mem injected context with observations from config-b. In this case, the consequences are harmless, but in a real setting this kind of cross-project leakage could be serious. Imagine two unrelated agentic projects with the similar file structures: observations about tools, schemas, prompts would be subject to cross-project contamination.

CONCLUSION

Based on our first experiments with claude-mem, the results were promising but also raised some concerns. First, It genuinely shows value as a context enrichment tool, given that it delivers in session continuity and decision recall. On the other hand, pattern reuse and code-dictated decisions don’t seem to be its strong suits. The fact that it doesn’t hallucinate is a genuine winning point, and setup friction (prerequisite errors, Vertex AI credential fix) is a minor inconvenience compared to the project-leakage bug, which is a real issue rather than an edge case. In addition, scoping memory by folder name rather than full path means the tool is only as safe as your naming discipline. For work across multiple adjacent projects, the risk of context contamination increases when sessions are ran in lower-level folders. We fully expect claude-mem to improve, especially given its popularity within the open source community. We will certainly keep monitoring its evolution and evaluating as a puzzle piece for our agentic stack.

Author

  • Georgios Konstantoulas is an Electrical and Computer Engineering graduate currently working as a Junior Data Scientist at Satalia (WPP Group). His background is primarily on computer vision, with hands-on experience in designing and benchmarking deep learning architectures in representation learning and classification tasks. He is currently contributing to research in agentic AI systems, exploring design and deployment in real-world settings.

More posts