{"id":2045,"date":"2026-09-15T08:19:44","date_gmt":"2026-09-15T08:19:44","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=2045"},"modified":"2026-09-16T08:50:30","modified_gmt":"2026-09-16T08:50:30","slug":"simulation-based-verification-for-llm-powered-agents","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=simulation-based-verification-for-llm-powered-agents","title":{"rendered":"Simulation-based Verification for LLM-powered Agents"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>In a previous post we introduced<\/strong> <a href=\"https:\/\/research.wpp.com\/blog\/seven-principles-for-agentic-governance-at-wpp-freedom-to-innovate-confidence-to-scale\">Seven principles for agentic governance<\/a>. A follow-up <a href=\"https:\/\/research.wpp.com\/blog\/from-principles-to-practice-a-governed-multi-agent-community-for-media-campaigns\">post<\/a> then showed them in action in a real setting: a multi-agent community for media campaigns. In this post, we zoom in on a specific member of that community, the Geo agent, and on the sixth principle: continuous verification. We describe how we verify the agent\u2019s behavior by simulating user sessions that include hundreds of questions, designed specifically to stress-test it and expose different types of failure modes. We also discuss how the results of these tests directly inform our efforts to continuously iterate and improve our agents.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A zip code is just a number until you give it meaning<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/research.wpp.com\/blog\/from-principles-to-practice-a-governed-multi-agent-community-for-media-campaigns\">Our community<\/a> includes multiple independent expert agents that collaborate to complete various tasks related to media campaigns. One of these experts is the <strong>Geo agent,<\/strong> designed to answer questions about locations at the ZIP-code level. Examples include:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Which locations in the USA are the most semantically similar to ZIP code 10005?<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Which look nothing like it?<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>How has the semantic profile of the ZIP code evolved in the last 6 months?<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Geo agent is powered by Google\u2019s Earth AI offering, which gives it access to a vast and diverse knowledge base built from the many geo-related data sources in Google\u2019s arsenal. One of these datasets is the <a href=\"https:\/\/developers.google.com\/maps\/documentation\/population-dynamics-insights\/overview\">Population Dynamics Insights<\/a> (PDI), which combines signals such as search interest, mobility, busyness, and environmental conditions like weather and air quality into a single numerical vector for each location (each ZIP code). We refer to such vectors as <em>embeddings<\/em>: semantic representations that encode what it is \u201cto be\u201d that location. Each PDI embedding is a vector of 330 floating point numbers, representing a point in a semantic space of 330 dimensions. Comparing two locations can then be easily achieved by computing the distance between their corresponding embeddings in that space. Google periodically updates these embeddings to reflect the evolution of the underlying signals that produce them.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"512\" height=\"288\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig01-1.png\" alt=\"\" class=\"wp-image-2048\" style=\"width:645px;height:auto\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig01-1.png 512w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig01-1-300x169.png 300w\" sizes=\"auto, (max-width: 512px) 100vw, 512px\" \/><figcaption class=\"wp-element-caption\">Figure 1. Population Dynamics Insights distils signals like search trends, mobility, weather, air quality, and Maps activity into ML-ready embeddings for every location. Source: <a href=\"https:\/\/mapsplatform.google.com\/resources\/blog\/from-static-maps-to-geospatial-ai-announcing-population-dynamics-insights\/\">Google<\/a><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The Geo agent has full access to these evolving PDI embeddings. Its LLM brain allows it to understand questions like the ones we listed above, and its PDI access allows it to answer them. The Geo agent is a critical part of our media-focused agentic community, as it is responsible for providing insights pre-flight, when the community is evaluating various locations to decide which ones should be included in the campaign, as well as post-flight, when the community is trying to explain the observed performance across the targeted locations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This means that continuous verification is critical: if the Geo agent starts producing inaccurate answers, the community\u2019s entire process becomes compromised.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Simulation-Based Verification<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">One of our primary tools for agentic verification is the <a href=\"https:\/\/conscium.com\/verifyax\">VerifyA<\/a>\u03a7 platform, which automatically generates scenarios designed to test the agent in specific ways that fit its description and abilities. We onboard our agents to the platform by presenting their A2A card. The platform then allows us to generate scenarios, simulate them, and receive informative reports that expose key failure modes and suggest ways to address them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Each simulation includes one or more \u201cNPC\u201d (Non-Playable Character) agents that are specifically and automatically generated to create the conditions required to expose various failure modes. Given the Q&amp;A nature of the Geo agent, we opt for simulations in which NPCs ask it various geo-related questions and we compare its responses with the correct reference responses (the ground truth).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The platform automatically generates both questions and their reference \u201cground truth\u201d responses in a federated manner, without actually moving the PDI data that powers the agent.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Having the correct response allows us to evaluate both the accuracy and completeness of the Geo agent\u2019s responses. It also lets us probe its robustness to hallucinations by asking seemingly answerable questions that actually cannot be answered based on the available data. Figure 2 shows an example of such a hallucination trap.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1009\" height=\"543\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig02.png\" alt=\"\" class=\"wp-image-2049\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig02.png 1009w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig02-300x161.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig02-767x413.png 767w\" sizes=\"auto, (max-width: 1009px) 100vw, 1009px\" \/><figcaption class=\"wp-element-caption\">Figure 2. A hallucination trap inside a simulated session. Right after the opening exchange, the NPC asks for similar ZIP codes without naming a target, and the agent asks for the missing ZIP code instead of inventing an answer. <\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The reports returned by VerifyAX include the full transcript of the interaction between our agent and the NPCs. For Q&amp;A-type evaluations, the transcript also includes the correct answer, making it easy to understand exactly what the agent did wrong.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 3 shows a set of historical runs and their corresponding scores.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1249\" height=\"589\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig03.png\" alt=\"\" class=\"wp-image-2050\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig03.png 1249w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig03-300x141.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig03-768x362.png 768w\" sizes=\"auto, (max-width: 1249px) 100vw, 1249px\" \/><figcaption class=\"wp-element-caption\">Figure 3. Historical evaluation runs for the Geo agent in VerifyAX. Each row is one scenario bundle of twenty questions, with the bundle scored out of five.<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Putting the Geo Agent to the Test<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We put the Geo agent through dozens of simulations with hundreds of questions. These runs surfaced multiple critical failure modes. By fixing these issues and re-testing, we significantly improved our agent and prepared it to be confidently used in production. Figure 4 illustrates the continuous process of fixing and retesting the agent.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"585\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig04-1024x585.png\" alt=\"\" class=\"wp-image-2051\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig04-1024x585.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig04-300x171.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig04-767x438.png 767w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig04-1536x877.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig04-2048x1169.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 4. The agent is connected once. After that the loop keeps turning: a fix only counts once it is deployed and the questions are asked again against the running agent. The results in this post are one turn around it.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Next, we discuss some of the critical and recurring failure modes that were surfaced by our simulations. The examples that follow are taken from the first rounds of simulation, before the corresponding fixes were deployed.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Failure Mode 1: Understanding Time<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The agent was asked to answer questions based on past data, anchored to a specific point in time. This is well within the Geo agent\u2019s scope, as it has access to multiple snapshots of the PDI embeddings across time. However, the agent consistently failed to take the given timeframe into account. This showed up in two ways: in some runs it silently answered using the latest (current) snapshot, as shown in Figure 5; in others it declined the question altogether, as shown in Figure 6. The root cause was the same in both cases, as the retrieval tool always queried the most recent snapshot regardless of the date the user asked for.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"535\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig05-1024x535.png\" alt=\"\" class=\"wp-image-2052\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig05-1024x535.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig05-300x157.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig05.png 1113w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 5. Failure Mode 1, silent. Asked about the 2026-05 snapshot, the agent returns 78722 instead of 78751, with the wrong embedding.<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"636\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig06-1024x636.png\" alt=\"\" class=\"wp-image-2056\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig06-1024x636.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig06-300x186.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig06-768x477.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig06.png 1122w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 6. Failure Mode 1, explicit. Asked for the 2023-08 snapshot, the agent declines, saying its tool only uses the latest data.<\/figcaption><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">Failure Mode 2: Accepting Invalid Requests<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The agent was asked to answer questions based on invalid US ZIP codes, such as 00001. Instead of rejecting those outright, the agent accessed its PDI knowledge base and returned a link to a file with the response, without ever protesting the faulty ZIP codes. Thankfully, the file was empty, which means that the agent did not go as far as to hallucinate data and responses for such ZIP codes. It did, however, present that empty file as the requested embedding, which is why the grader marks the answer as fabricated. Still, this is a failure mode that exposes the disconnect between the agent\u2019s natural language interface (which happily accepts the request and returns a response) and the tool that it uses to access the PDI embeddings and actually compute the response. When the tool fails to identify data for a given ZIP code, this should be clearly raised to the user.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An example of this failure mode is shown in Figure 7.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"623\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig07-1024x623.png\" alt=\"\" class=\"wp-image-2057\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig07-1024x623.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig07-300x183.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig07-767x467.png 767w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig07.png 1122w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 7. Failure Mode 2. Asked for invalid postal code 00001, the agent returns a link to a JSON artifact with empty embedding arrays.<\/figcaption><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">Failure Mode 3: Inability to Scale<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The agent was asked multiple consecutive questions in each simulated session. In many of these, it suddenly stopped producing answers and instead returned errors indicating that it had exhausted its context window. This is a sign of poor memory management: the agent never cleaned up or re-summarized its memory as it accumulated across consecutive NPC interactions. Eventually the context window filled up, leaving the agent unable to take on new requests. An example of this is shown in Figure 8.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"795\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig08-1024x795.png\" alt=\"\" class=\"wp-image-2058\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig08-1024x795.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig08-300x233.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig08.png 1121w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 8. Failure Mode 3. Mid-session the agent returns a raw 400 INVALID_ARGUMENT error: history exceeded the one-million-token limit.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">A second sub-issue related to scalability emerged when we started multiple parallel simulations via VerifyAX. As the number of parallel threads increased, the agent reached its limit and started producing capacity errors. An example is shown in Figure 9. We addressed this issue by adding a queue that prevents the agent from getting overwhelmed. Obviously this can increase waiting times. Auto-scaling is another valid option that we apply in our agentic community. However, the queue remains a best-practice feature that allows the agent to operate within its hardware constraints without failing ungracefully (throwing errors).<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"332\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig09-1024x332.png\" alt=\"\" class=\"wp-image-2059\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig09-1024x332.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig09-300x97.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig09-765x248.png 765w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig09.png 1117w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 9. Failure Mode 3. Running in parallel, the agent returns a raw 429 RESOURCE_EXHAUSTED error; 10 of 20 questions went unasked.<\/figcaption><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">Failure Mode 4: Post-Processing Failures<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Due to the nature of the underlying PDI data, the correct answer to certain geospatial questions can be quite long. For instance, asking for the semantic embedding of one or more ZIP codes translates to sharing multiple numeric lists with hundreds of numbers. To avoid cluttering its output, the Geo agent is thus designed to return a JSON file that includes all the relevant info.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The simulations revealed that, in many cases, even though the agent successfully computed the correct response and even returned a signed URL to the file with the supporting data, it was unwilling or unable to actually look at its own data and compute the final response, even if that required only a minor post-processing step.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For example, when asked for the dimensionality of a ZIP code\u2019s embedding (easy, since it\u2019s always 330 for PDI embeddings), the agent correctly retrieved the vector but said that it was unable to count its dimensions. See Figure 10 for an example.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"521\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig10-1024x521.png\" alt=\"\" class=\"wp-image-2060\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig10-1024x521.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig10-767x390.png 767w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig10-300x153.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig10.png 1424w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 10. Failure Mode 4 in VerifyAX. The NPC asks for the dimensionality and three largest-magnitude components of ZIP 90210\u2019s embedding. The agent refuses, saying it cannot read those values from the signed URL its own get_pdfm_embeddings tool produced, even though the dimensionality is always 330 and the components are a direct read-off from the retrieved vector.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">When asked to first find the single ZIP code most similar to a target, then return that neighbor\u2019s embedding, the agent refused the whole request. This task is two steps: find the nearest neighbor, then fetch its embedding. The agent said it could run the search and return a link to the results, but could not then open that link to read off the neighbor and look up its embedding. The agent could do each step on its own, but refused to connect them, inventing a limit that its own earlier runs prove is false. An example of this is shown in Figure 11.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1316\" height=\"428\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig11-1.png\" alt=\"\" class=\"wp-image-2062\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig11-1.png 1316w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig11-1-300x98.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig11-1-766x249.png 766w\" sizes=\"auto, (max-width: 1316px) 100vw, 1316px\" \/><figcaption class=\"wp-element-caption\">Figure 11. Failure Mode 4 in VerifyAX. The NPC asks for the single ZIP code most similar to 90210, then for that neighbor\u2019s embedding. The agent refuses, saying it cannot open its own results link to read off the neighbor and look up its embedding, even though the answer (90049) is available and the same chain succeeded in earlier runs.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Finally, the agent was asked how the drift ranking of a set of ZIP codes changes between a six-month window and a twelve-month window. Drift is how far a ZIP code\u2019s embedding moves between two dates, measured as the distance between its two snapshots. To answer, the agent needs to fetch each ZIP code\u2019s embedding at the relevant dates, work out its drift for each window, and compare the two rankings. Its retrieval tool can fetch every one of those snapshots, so this is just a few ordinary tool calls. Instead it declined, saying it could not compare the two windows \u201cin a single step with the available tools\u201d. This was a false limit, since the same kind of step-by-step work succeeded in other runs. See Figure 12 for an example.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"463\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig12-1024x463.png\" alt=\"\" class=\"wp-image-2063\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig12-1024x463.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig12-300x136.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig12-768x347.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig12.png 1451w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 12. Failure Mode 4 in VerifyAX. The NPC asks whether ZIP 60614 drifted more over the last six months or the last twelve. The agent refuses, saying it cannot compare drift across two time periods in a single step with the available tools, even though its retrieval tool can fetch each snapshot and the correct answer (the six-month window moved more) is a short calculation.<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">The Agent\u2019s Current State<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 13 visualizes the progress of the Geo agent after our first round of simulations and corresponding fixes. We have 18 different simulation bundles, each focused on exposing a different failure mode. The figure has one row for each bundle. The white dot marks the agent\u2019s grade (out of 5, as per the VerifyAX scoring system) before our fixes. The blue dot marks the updated grade after applying fixes and re-running the bundle.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"735\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig13_beforeafter-1024x735.png\" alt=\"\" class=\"wp-image-2064\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig13_beforeafter-1024x735.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig13_beforeafter-300x215.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig13_beforeafter-767x550.png 767w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig13_beforeafter-1536x1102.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/fig13_beforeafter.png 2024w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 13. Bundle scores at baseline and after the first round of verification. Each row is one scenario bundle of twenty graded questions.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The Figure clearly visualizes the agent\u2019s improvement across the board, and illustrates the value of continuous verification. The original version of the agent (white dots) was simply not ready for production. The updated version now scores almost perfectly (5\/5) across the board, with only minor issues remaining in some of the bundles.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What\u2019s Next<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Every agent in our community is currently being onboarded and tested via VerifyAX, to enable continuous verification via simulation. The platform\u2019s full verification functionality can also be accessed programmatically, which enables us to make it a part of our CI\/CD pipeline and block faulty versions from being deployed into production.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">VerifyAX is not the only tool in our agentic verification toolkit. Simulation is a useful approach that allows us to safely stress-test our agents in a sandbox, rather than having them learn \u201con the job\u201d. However, we are also experimenting with additional complementary tools that will allow us to install guardrails and detect regressions while the agents are serving real users and requests in production.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">References<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><a href=\"https:\/\/research.wpp.com\/blog\/from-principles-to-practice-a-governed-multi-agent-community-for-media-campaigns\">From Principles to Practice: A Governed Multi-Agent Community for Media Campaigns<\/a>, the companion post<\/li>\n\n\n\n<li><a href=\"https:\/\/research.wpp.com\/blog\/seven-principles-for-agentic-governance-at-wpp-freedom-to-innovate-confidence-to-scale\">Seven principles for agentic governance at WPP: freedom to innovate, confidence to scale<\/a><\/li>\n\n\n\n<li><a href=\"https:\/\/blog.google\/innovation-and-ai\/products\/google-earth-ai\/\">Google Earth AI<\/a><\/li>\n\n\n\n<li><a href=\"https:\/\/github.com\/google-research\/population-dynamics\">Population Dynamics Foundation Model (PDFM)<\/a><\/li>\n\n\n\n<li><a href=\"https:\/\/developers.google.com\/maps\/documentation\/population-dynamics-insights\/overview\">Population Dynamics Insights<\/a><\/li>\n\n\n\n<li><a href=\"https:\/\/research.wpp.com\/blog\/a-research-agenda-for-expert-agent-communities\">A research agenda for expert agent communities<\/a><\/li>\n\n\n\n<li><a href=\"https:\/\/conscium.com\/verifyax\">VerifyA<\/a>X \u00b7 <a href=\"https:\/\/langfuse.com\/\">Langfuse<\/a> \u00b7 <a href=\"https:\/\/a2a-protocol.org\/\">A2A protocol<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>In a previous post we introduced Seven principles for agentic governance. A follow-up post then showed them in action in a real setting: a multi-agent community for media campaigns. In this post, we zoom in on a specific member of that community, the Geo agent, and on the sixth principle: continuous verification. We describe how [&hellip;]<\/p>\n","protected":false},"author":44,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":"{\"authors\":[59,37],\"author_categories\":{\"37\":\"1\",\"59\":\"1\"},\"fallback_author_user\":\"4\",\"ppma_author_box_select\":\"\",\"selected_authors\":[{\"id\":59,\"display_name\":\"Maria Kalogianni\",\"is_guest\":0,\"category_id\":\"1\"},{\"id\":37,\"display_name\":\"Thanos Lyras\",\"is_guest\":0,\"category_id\":\"1\"}]}"},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":44,"display_name":"Maria Kalogianni","first_name":"Maria","last_name":"Kalogianni","nickname":"maria.kalogianni","user_nicename":"maria-kalogianni","user_email":"maria.kalogianni@satalia.com","biographical_info":"Maria Kalogianni is a Data Scientist at Satalia and part of the WPP Research team. She holds MSc degrees in Pure Mathematics and Big Data and Analytics, and her interests span graph representation learning, natural language processing, and deep learning. Her current work focuses on developing and evaluating AI systems that work together to support marketing decisions.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/image1.png","job_title":"Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":null},{"id":21,"display_name":"Thanos Lyras","first_name":"Thanos","last_name":"Lyras","nickname":"thanos.lyras","user_nicename":"thanos-lyras","user_email":"thanos.lyras@satalia.com","biographical_info":"Thanos Lyras is a data scientist at Satalia specializing in building end-to-end AI pipelines and deploying real-world AI applications. A graduate in Computer Engineering with an MSc in Data Science, his work has led to research contributions in the fields of Big Data, AI, and database performance. Currently, he is focused on pioneering research in agentic projects, exploring the next wave of artificial intelligence.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/04\/profile_photo.png","job_title":"Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":null}],"class_list":["post-2045","research_feed","type-research_feed","status-publish","hentry","content_type-article"],"acf":{"content_quarter":"","related_pods":[2010]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"","related_pods":["2010"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/2045","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/44"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/2010"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2045"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2045"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=2045"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=2045"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}