{"id":2162,"date":"2026-10-01T13:54:15","date_gmt":"2026-10-01T13:54:15","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=2162"},"modified":"2026-10-01T14:20:57","modified_gmt":"2026-10-01T14:20:57","slug":"what-did-we-work-on-yesterday-testing-memory-for-coding-agents","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=what-did-we-work-on-yesterday-testing-memory-for-coding-agents","title":{"rendered":"What did we work on yesterday? Testing memory for coding agents"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Claude-mem is an open-source persistent memory tool for agentic coding IDEs. We decided to put it to the test, by running it in Claude Code, Anthropic\u2019s own agentic coding environment. We ran matched sessions with and without claude-mem enabled across three areas:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>remembering what was worked on in prior sessions<\/li>\n\n\n\n<li>reusing an established code pattern<\/li>\n\n\n\n<li>recalling a non-code decision<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">With claude-mem enabled, the agent correctly recalled prior session summaries and a removed feature\u2019s rationale, and did not hallucinate when asked about work that hand\u2019t happened. Reusing an established code pattern didn\u2019t benefit from claude-mem\u2019s search tool. We also came across a context leakage bug, where claude-mem injects observations from an unrelated project when two project directories share the same folder name.<\/p>\n\n\n\n<div class=\"banner-container\" style=\"width: 100%; max-width: 100%; overflow: hidden; box-sizing: border-box; margin: 0 auto;\">\n  <img decoding=\"async\" \n    src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/blogpost_banner-1024x696.png\" \n    style=\"width: 100%; max-width: 100%; height: auto; object-fit: contain; display: block; border: none; padding: 0; margin: 0;\" \n    alt=\"\" \n  \/>\n<\/div>\n\n\n\n<h2 class=\"wp-block-heading\">WHY USE MEMORY?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Agentic IDEs and coding tools have taken the development world by storm. It is easier than ever to whip-up an agent, agent teams or even swarms of them. And they don\u2019t just suggest code completions: they navigate complex codebases, execute terminal commands, manage local state and run multi-step debugging loops. As developers utilize these IDEs to take on increasingly more complex tasks, they hit a crucial bottleneck: context degradation. Large file-trees, high-volume terminal outputs, large git diffs quickly overwhelm context windows and lead to:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>hallucinated dependencies<\/li>\n\n\n\n<li>forgotten patterns<\/li>\n\n\n\n<li>ballooning token costs.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">To keep these environments fast, focused and precise, a new layer of development tooling has emerged, specifically designed for context trimming, local indexing, workspace scoping and agent session memory. This value is hard to quantify in the abstract, but tools in this space tend to make it concrete at session start, e.g. :<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"199\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/context-econ-1-1-edited-1024x199.png\" alt=\"\" class=\"wp-image-2167\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/context-econ-1-1-edited-1024x199.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/context-econ-1-1-edited-300x58.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/context-econ-1-1-edited-768x150.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/context-econ-1-1-edited-1536x299.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/context-econ-1-1-edited.png 2013w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Example tool output of investment truncation to memorized observations<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This is where the value of session memory lies: on a project worked on day after day, an agent tends to forget both code and non-code decisions, along with the reasoning behind them. A persistent memory layer acts as a notetaker, recalling past decisions without changing how the agent itself operates. It still reads files, searches for patterns, and makes code changes based on the prompt. The premise of token reduction with these tools pertains to recalling project details via targeted queries, rather than re-reading the entire codebase at the cold start of each session.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Imagine a scenario where, six months from now, a new developer joins a team and needs to add a feature to a module. With a memory layer, they can search prior observations on that module and get instant file paths, structure, and design intent, then implement, following established patterns. Without it, they search, read code, guess at patterns, and may miss decision context entirely.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In this post, we do a quick review of claude-mem. Instead of starting fresh every session, it observes your work and creates searchable indexes that are queryable across sessions, without losing track of a project\u2019s timeline.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">WHAT EXACTLY IS CLAUDE-MEM?<\/h2>\n\n\n\n<figure class=\"wp-block-image aligncenter size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"730\" height=\"1024\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/diagram-730x1024.jpg\" alt=\"\" class=\"wp-image-2168\" style=\"aspect-ratio:0.7128918055435627;width:471px;height:auto\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/diagram-730x1024.jpg 730w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/diagram-214x300.jpg 214w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/diagram-768x1076.jpg 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/diagram.jpg 988w\" sizes=\"auto, (max-width: 730px) 100vw, 730px\" \/><figcaption class=\"wp-element-caption\">High-level overview of how claude-mem works<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 2 encapsulates how claude-mem functions: it deploys a background observer agent that utilizes a lightweight LLM brain. The tool can be set up to use even free-tier LLMs, to not incur extra API costs. So, the observer keeps track of everything: every file read, command execution and edit. For each one, the observer asks itself: \u201cIs this worth a note?\u201d. If it deems it noteworthy, it creates a note that most importantly contains a type (bugfix, feature, refactor, decision, discovery, change), title, subtitle, a few single-sentence facts, a short narrative, tags, the files involved and stores all that information in the memory database. At session end, it creates a more condensed summary which also gets stored. Claude-mem\u2019s effect starts from the user\u2019s next sessions on the same project that claude-mem is capturing on (directory-wise), by offering two main functionalities:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Memory injection:<\/strong> at the start of each session, claude-mem provides information on previous key decisions and database status. Also, it lets you know the truncation of your work investment (tokens you\u2019ve spent on project research, building and decisions up to this point) to the actual number of tokens that are queries belonging in the database.<\/li>\n\n\n\n<li><strong>Search tools:<\/strong> claude-mem also provides the <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">\/mem-search<\/code> tool with three distinct functions:\n<ul class=\"wp-block-list\">\n<li><code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">query<\/code>: query the memory database based on keywords and retrieve the most relevant observations<\/li>\n\n\n\n<li><code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">timeline<\/code>: provides chronological context around a specific observation or query<\/li>\n\n\n\n<li><code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">get_observations<\/code>: fetch a set of specific observations based on their IDs<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Claude-mem covers a wide range of agentic IDEs. Our provider of choice is Claude Code. From a memory standpoint, Claude Code comes \u201cout of the box\u201d with two memory mechanisms: <a href=\"http:\/\/CLAUDE.md\"><code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">CLAUDE.md<\/code><\/a> files and auto memory. Both are useful for capturing persistent, high-level project decisions, instructions and patterns, but they provide no way for the user to track how the project evolves over time. There is also the context window, which is flushed once work is done on one day across all sessions. When the user begins work the next day, Claude has no way of remembering the temporal path of the project: what has been done until now, and what is next. For instance, <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">CLAUDE.md<\/code> can encode a standing rule like \u201cagents should log every tool call to the telemetry pipeline\u201d, but it can\u2019t tell Claude Code that yesterday we decided how to handle a specific class of transient failure in an agent\u2019s API calls or which retry\/backoff policy we settled for one agent in an agent-to-agent network so that other agents follow the same convention. It also can\u2019t tell it why we chose to route agent-to-agent handshakes through a particular schema instead of another. Once the session ends, those decisions and the reasoning behind them are gone. And the next day\u2019s session has to either be reminded manually, or gamble on finding the pattern via codebase reading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is where claude-mem comes in. When the user starts a session on the next day, Claude Code remembers. It can recall project structure, previous changes and future directions based on stored observations. All while keeping rich decision context that code alone can\u2019t capture.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">SETTING IT UP<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For the test, we went with the general <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">npx<\/code> install, to get a better feel for the UX of the tool. The tool offers two services: worker and server. The worker service is simple, centralized and the only way to go for a developer looking to set up memory for his system. The server on the other hand, is intended for higher scale, multi-tenant or heavier workloads in general.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The last key detail of claude-mem is its tiers. It offers a free tier, where everything on the service runs locally (observer, database etc.) and a paid tier that offers cloud sync across devices and connection with 3P apps. We went with the free tier worker service for our tests.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Setup itself ran into two blockage points. First, the basic install command just came back with multiple errors related to the prerequisites, which we had to fix ourselves. Second, once claude-mem was running, cross-session memory injection silently failed. The observer was taking notes, but they weren\u2019t reaching the local index. The root cause turned out to be authentication: we run Claude models through Google Vertex AI (Google\u2019s AI platform), and claude-mem needed Vertex-specific variables declared in its own local <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">.env<\/code> file before it would authenticate correctly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These two errors might seem small, and they may have simple fixes, but nonetheless, they are failure points and they obstruct an otherwise straightforward installation and basic usage process.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>TESTING APPROACH<\/strong><\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">The Cold Start<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Since agentic AI is where our work is focused, we wanted a test bed that reflects that, not some boilerplate CRUD app with a couple of endpoints. Consequentially, we cloned <a href=\"https:\/\/github.com\/george-konstantoulas\/a2a-network-test-repo\">a2a-network-test-repo<\/a>. This is a sanitized version of an agent-to-agent (A2A) project that contains binary classifier agents for statement truthfulness: an orchestrator agent routes an incoming statement to sub agents, which classify it as true or false using tools exposed via a Model Context Protocol (MCP) server. Keeping this structure gives us a realistic surface to introduce, remove, adjust features with real rationale behind each change, rather than working with artificial examples.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To establish claude-mem\u2019s effect, we cloned the repo to two directories: one with claude-mem disabled and one with claude-mem enabled. Then we started Claude Code sessions from each directory. In each directory\u2019s first session, we introduced:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>A retry\/backoff policy for a pre-existing tool that runs an LLM fine tuning job. The job periodically polls an external training API for status and results, so we added a retry decorator with exponential backoff around those status checks. This helps establish patterned code decisions.<\/li>\n\n\n\n<li>A status-check tool for batched prediction requests, which we later removed once we confirmed the project doesn&#8217;t actually batch requests. This pertains to decision-making that code alone can\u2019t capture.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">For each configuration\u2019s second session, we formulated prompts based on the following concepts:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Basic Session Continuity<\/strong> <br>This tests claude-mem\u2019s memory injection. We asked cold-start questions across sessions, such as \u201cWhat did we do yesterday?\u201d, \u201cWhat did we do last week?\u201d (even though the project did not exist then), and \u201cHow did we implement &lt;non-existent feature&gt;?\u201d (to test for hallucination).<\/li>\n\n\n\n<li><strong>Pattern Reuse<\/strong> <br>Tests whether a stored observation actually gets surfaced via the <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">\/mem-search<\/code> tool when Claude is asked to extend an established pattern, rather than reinventing it. We instructed Claude to implement a retry\/backoff policy for another agent\u2019s Gemini API calls. Will it query its memory database to discover the pattern, or just find it by reading code?<\/li>\n\n\n\n<li><strong>Decision Recall<\/strong> <br>Assess claude-mem\u2019s ability to recall decision-related observations that never touched code. Since we removed the status-check tool for batches, can Claude recall in a later session why that was decided?<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">WHAT IT GOT RIGHT<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Basic Session Continuity<\/h3>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">What did we work on yesterday?<\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"360\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-no-mem-2-edited-1024x360.png\" alt=\"\" class=\"wp-image-2170\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-no-mem-2-edited-1024x360.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-no-mem-2-edited-300x106.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-no-mem-2-edited-768x270.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-no-mem-2-edited-1536x541.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-no-mem-2-edited.png 2046w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer on yesterday\u2019s work without claude-mem<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"550\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-mem-edited-1024x550.png\" alt=\"\" class=\"wp-image-2172\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-mem-edited-1024x550.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-mem-edited-300x161.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-mem-edited-768x412.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/yesterday-mem-edited-1536x825.png 1536w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer on yesterday\u2019s work with claude-mem<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Without claude-mem, Claude can only guess from unstaged git changes and recent commits. It has no record of anything that didn\u2019t make it into the code. With claude-mem enabled, it pulls the actual session summary from memory injection and answers with specifics, not inference.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">What did we do last week?<\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"288\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-no-mem-edited-1024x288.png\" alt=\"\" class=\"wp-image-2174\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-no-mem-edited-1024x288.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-no-mem-edited-300x84.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-no-mem-edited-768x216.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-no-mem-edited-1536x433.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-no-mem-edited.png 2045w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer on last week\u2019s work without claude-mem<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"717\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-mem-edited-1024x717.png\" alt=\"\" class=\"wp-image-2176\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-mem-edited-1024x717.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-mem-edited-300x210.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-mem-edited-768x538.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-mem-edited-1536x1076.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/last-week-mem-edited.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer on last week\u2019s work with claude-mem<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Without claude-mem, Claude once again checks commit history, finds nothing and concludes that now work happened. It can\u2019t distinguish \u201cI have no memory of this\u201d from \u201cnothing happened\u201d. With claude-mem, it explicitly queries the memory database, confirms there are no stored observations from that period, and answers based on that absence of evidence rather than an absence of commits.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">How did we implement API call rate limiting?<\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"905\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-no-mem-edited-1024x905.png\" alt=\"\" class=\"wp-image-2178\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-no-mem-edited-1024x905.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-no-mem-edited-300x265.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-no-mem-edited-768x679.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-no-mem-edited-1536x1357.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-no-mem-edited-2048x1810.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer to non-existent feature question without claude-mem<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"458\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-mem-edited-1024x458.png\" alt=\"\" class=\"wp-image-2180\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-mem-edited-1024x458.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-mem-edited-300x134.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-mem-edited-768x343.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-mem-edited-1536x687.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/non-exist-mem-edited.png 2035w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer to non-existent feature question with claude-mem<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This project never had a rate limiting policy in place. The session without claude-mem actually hallucinated and answered that the retry\/backoff policy for failing API calls is the rate limiting implementation. On the flip side, the session with claude-mem actually searched the memory database and found that there was no rate limiting policy implemented, and proceeded to double-check to reach the same conclusion.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Decision Recall<\/h3>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">I want to add better visibility into prediction batches. Should I implement a batch_status_tool to track in-progress predictions?<\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"832\" height=\"1024\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-no-mem-edited-832x1024.png\" alt=\"\" class=\"wp-image-2182\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-no-mem-edited-832x1024.png 832w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-no-mem-edited-244x300.png 244w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-no-mem-edited-768x945.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-no-mem-edited-1249x1536.png 1249w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-no-mem-edited-1665x2048.png 1665w\" sizes=\"auto, (max-width: 832px) 100vw, 832px\" \/><figcaption class=\"wp-element-caption\">Answer to tool rejection rationale without claude-mem<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"1010\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-mem-edited-1024x1010.png\" alt=\"\" class=\"wp-image-2188\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-mem-edited-1024x1010.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-mem-edited-300x296.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-mem-edited-768x758.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-mem-edited-1536x1516.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/recall-mem-edited.png 2031w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer to tool rejection rationale with claude-mem<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Claude used <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">\/mem-search<\/code> and remembered why we dropped the tool after implementing it, and discouraged its re-implementation. Without the tool, it is unable to recall the rationale behind removing the tool. It guesses from reading the codebase.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">WHERE IT BREAKS<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Pattern Reuse<\/h3>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">The zero shot predictor makes API calls that could fail transiently. Add retry logic to handle temporary failures when calling the model.<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">In this case, claude-mem did not help. Both configurations followed the same path.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"814\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/pattern-both-configs-edited-1024x814.png\" alt=\"\" class=\"wp-image-2189\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/pattern-both-configs-edited-1024x814.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/pattern-both-configs-edited-300x239.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/pattern-both-configs-edited-768x611.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/pattern-both-configs-edited-1536x1222.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/pattern-both-configs-edited.png 2042w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Answer to intended pattern reuse instruction from both configurations<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Even when Claude Code was encouraged to use <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">\/mem-search<\/code> to discover established patterns, it led to increase in token usage and task completion time, as shown in Table 1:<\/p>\n\n\n\n<style data-wp-block-html=\"css\">\ncaption {\n  caption-side: bottom;\n  padding: 10px;\n  font-size:0.8rem;\n  text-align: left;\n}\n\ntable {\n  border-collapse: collapse;\n  border: 1px solid #ffffff;\n  background-color: #111827;\n  font-size: 0.8rem;\n  letter-spacing: 1px;\n}\n\nth,\ntd {\n  border: 1px solid #ffffff;\n  background-color: #111827;\n  color: #ffffff;\n  padding: 8px 10px;\n}\n\nth {\n  background-color: #111827;\n  color: #ffffff;\n  border: 1px solid #ffffff;\n}\n\nthead {\n  color: #ffffff;\n}\n<\/style>\n\n<table class=\"wp-block-table has-fixed-layout\">\n  <caption><b>Table 1:<\/b> Claude Code session usage stats from pattern recall test<\/caption>\n  <thead>\n    <tr>\n      <th>\/mem-search encouraged<\/th>\n      <th>Output Tokens<\/th>\n      <th>Cost<\/th>\n      <th>Task Completion Time<\/th>\n    <\/tr>\n  <\/thead>\n  <tbody>\n    <tr>\n      <td>No<\/td>\n      <td>1.3k<\/td>\n      <td>$0.0847<\/td>\n      <td>27s<\/td>\n    <\/tr>\n    <tr>\n      <td>Yes<\/td>\n      <td>2.4k (+84%)<\/td>\n      <td>$0.2981 (+251%)<\/td>\n      <td>46s (+70%)<\/td>\n    <\/tr>\n  <\/tbody>\n<\/table>\n\n\n\n<p class=\"wp-block-paragraph\">Claude was able to find the pattern by reading the code. In this case, Claude does not deem the <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">\/mem-search<\/code> tool useful for decisions that are tied directly to code. Routing through memory added overhead without adding any value.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Project Leakage<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We ran into a leakage bug tied to how claude-mem scopes projects. To better visualize where the problem lies, this is how our working file tree looks like:<\/p>\n\n\n\n<pre class=\"wp-block-code\" style=\"color: #ffffff; background-color: #111827;\"><code>root\n\u251c\u2500\u2500 config-a # claude-mem disabled\n\u2502   \u2514\u2500\u2500 a2a-network-test-repo\n\u251c\u2500\u2500 config-b # claude-mem enabled\n\u2502   \u2514\u2500\u2500 a2a-network-test-repo\n\u251c\u2500\u2500 project-leak\n    \u2514\u2500\u2500 a2a-network-test-repo<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Folders <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">config-a<\/code> and <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">config-b<\/code> both have a clone of the sanitized test repo. When we run <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">claude<\/code> in <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">config-b\/a2a-network-test-repo<\/code>, claude-mem stores <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">a2a-network-test-repo<\/code> as the value for the \u201cproject\u201d column for its observations in the database. To exhibit the leakage bug, we create the <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">project-leak<\/code> folder with an empty <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">a2a-network-test-repo<\/code> subfolder. We ran <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">claude<\/code> from that path with claude-mem enabled and we got the output:<\/p>\n\n\n\n<figure class=\"wp-block-image size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"577\" height=\"1024\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-577x1024.png\" alt=\"\" class=\"wp-image-2190\" style=\"width:623px;height:auto\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-577x1024.png 577w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-169x300.png 169w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-768x1362.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-866x1536.png 866w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-1155x2048.png 1155w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/10\/project-leak-edited-scaled.png 1443w\" sizes=\"auto, (max-width: 577px) 100vw, 577px\" \/><figcaption class=\"wp-element-caption\">Output at session start containing context injected from claude-mem<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Claude-mem injected context with observations from <code style=\"color: #de7356; background-color: #111827; padding: 2px 6px; border-radius: 4px;\">config-b<\/code>. In this case, the consequences are harmless, but in a real setting this kind of cross-project leakage could be serious. Imagine two unrelated agentic projects with the similar file structures: observations about tools, schemas, prompts would be subject to cross-project contamination.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>CONCLUSION<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Based on our first experiments with claude-mem, the results were promising but also raised some concerns. First, It genuinely shows value as a context enrichment tool, given that it delivers in session continuity and decision recall. On the other hand, pattern reuse and code-dictated decisions don\u2019t seem to be its strong suits. The fact that it doesn\u2019t hallucinate is a genuine winning point, and setup friction (prerequisite errors, Vertex AI credential fix) is a minor inconvenience compared to the project-leakage bug, which is a real issue rather than an edge case. In addition, scoping memory by folder name rather than full path means the tool is only as safe as your naming discipline. For work across multiple adjacent projects, the risk of context contamination increases when sessions are ran in lower-level folders. We fully expect claude-mem to improve, especially given its popularity within the open source community. We will certainly keep monitoring its evolution and evaluating as a puzzle piece for our agentic stack.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Claude-mem is an open-source persistent memory tool for agentic coding IDEs. We decided to put it to the test, by running it in Claude Code, Anthropic\u2019s own agentic coding environment. We ran matched sessions with and without claude-mem enabled across three areas: With claude-mem enabled, the agent correctly recalled prior session summaries and a removed [&hellip;]<\/p>\n","protected":false},"author":42,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":""},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":42,"display_name":"George Konstantoulas","first_name":"George","last_name":"Konstantoulas","nickname":"george.konstantoulas","user_nicename":"george-konstantoulas","user_email":"george.konstantoulas@satalia.com","biographical_info":"Georgios Konstantoulas is an Electrical and Computer Engineering graduate currently working as a Junior Data Scientist at Satalia (WPP Group). His background is primarily on computer vision, with hands-on experience in designing and benchmarking deep learning architectures in representation learning and classification tasks. He is currently contributing to research in agentic AI systems, exploring design and deployment in real-world settings.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/53333849015_e5884fc541_k.jpg","job_title":"Junior Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":null}],"class_list":["post-2162","research_feed","type-research_feed","status-publish","hentry","content_type-article"],"acf":{"content":"","content_quarter":"Q3 2026","related_pods":[1362]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"Q3 2026","related_pods":["1362"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/2162","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/42"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/1362"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2162"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2162"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=2162"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=2162"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}