Author: Maria Kalogianni

  • Simulation-based Verification for LLM-powered Agents

    In a previous post we introduced Seven principles for agentic governance. A follow-up post then showed them in action in a real setting: a multi-agent community for media campaigns. In this post, we zoom in on a specific member of that community, the Geo agent, and on the sixth principle: continuous verification. We describe how we verify the agent’s behavior by simulating user sessions that include hundreds of questions, designed specifically to stress-test it and expose different types of failure modes. We also discuss how the results of these tests directly inform our efforts to continuously iterate and improve our agents.

    A zip code is just a number until you give it meaning

    Our community includes multiple independent expert agents that collaborate to complete various tasks related to media campaigns. One of these experts is the Geo agent, designed to answer questions about locations at the ZIP-code level. Examples include:

    Which locations in the USA are the most semantically similar to ZIP code 10005?

    Which look nothing like it?

    How has the semantic profile of the ZIP code evolved in the last 6 months?

    The Geo agent is powered by Google’s Earth AI offering, which gives it access to a vast and diverse knowledge base built from the many geo-related data sources in Google’s arsenal. One of these datasets is the Population Dynamics Insights (PDI), which combines signals such as search interest, mobility, busyness, and environmental conditions like weather and air quality into a single numerical vector for each location (each ZIP code). We refer to such vectors as embeddings: semantic representations that encode what it is “to be” that location. Each PDI embedding is a vector of 330 floating point numbers, representing a point in a semantic space of 330 dimensions. Comparing two locations can then be easily achieved by computing the distance between their corresponding embeddings in that space. Google periodically updates these embeddings to reflect the evolution of the underlying signals that produce them.

    Figure 1. Population Dynamics Insights distils signals like search trends, mobility, weather, air quality, and Maps activity into ML-ready embeddings for every location. Source: Google

    The Geo agent has full access to these evolving PDI embeddings. Its LLM brain allows it to understand questions like the ones we listed above, and its PDI access allows it to answer them. The Geo agent is a critical part of our media-focused agentic community, as it is responsible for providing insights pre-flight, when the community is evaluating various locations to decide which ones should be included in the campaign, as well as post-flight, when the community is trying to explain the observed performance across the targeted locations.

    This means that continuous verification is critical: if the Geo agent starts producing inaccurate answers, the community’s entire process becomes compromised.

    Simulation-Based Verification

    One of our primary tools for agentic verification is the VerifyAΧ platform, which automatically generates scenarios designed to test the agent in specific ways that fit its description and abilities. We onboard our agents to the platform by presenting their A2A card. The platform then allows us to generate scenarios, simulate them, and receive informative reports that expose key failure modes and suggest ways to address them.

    Each simulation includes one or more “NPC” (Non-Playable Character) agents that are specifically and automatically generated to create the conditions required to expose various failure modes. Given the Q&A nature of the Geo agent, we opt for simulations in which NPCs ask it various geo-related questions and we compare its responses with the correct reference responses (the ground truth).

    The platform automatically generates both questions and their reference “ground truth” responses in a federated manner, without actually moving the PDI data that powers the agent.

    Having the correct response allows us to evaluate both the accuracy and completeness of the Geo agent’s responses. It also lets us probe its robustness to hallucinations by asking seemingly answerable questions that actually cannot be answered based on the available data. Figure 2 shows an example of such a hallucination trap.

    Figure 2. A hallucination trap inside a simulated session. Right after the opening exchange, the NPC asks for similar ZIP codes without naming a target, and the agent asks for the missing ZIP code instead of inventing an answer.

    The reports returned by VerifyAX include the full transcript of the interaction between our agent and the NPCs. For Q&A-type evaluations, the transcript also includes the correct answer, making it easy to understand exactly what the agent did wrong.

    Figure 3 shows a set of historical runs and their corresponding scores.

    Figure 3. Historical evaluation runs for the Geo agent in VerifyAX. Each row is one scenario bundle of twenty questions, with the bundle scored out of five.

    Putting the Geo Agent to the Test

    We put the Geo agent through dozens of simulations with hundreds of questions. These runs surfaced multiple critical failure modes. By fixing these issues and re-testing, we significantly improved our agent and prepared it to be confidently used in production. Figure 4 illustrates the continuous process of fixing and retesting the agent.

    Figure 4. The agent is connected once. After that the loop keeps turning: a fix only counts once it is deployed and the questions are asked again against the running agent. The results in this post are one turn around it.

    Next, we discuss some of the critical and recurring failure modes that were surfaced by our simulations. The examples that follow are taken from the first rounds of simulation, before the corresponding fixes were deployed.

    Failure Mode 1: Understanding Time

    The agent was asked to answer questions based on past data, anchored to a specific point in time. This is well within the Geo agent’s scope, as it has access to multiple snapshots of the PDI embeddings across time. However, the agent consistently failed to take the given timeframe into account. This showed up in two ways: in some runs it silently answered using the latest (current) snapshot, as shown in Figure 5; in others it declined the question altogether, as shown in Figure 6. The root cause was the same in both cases, as the retrieval tool always queried the most recent snapshot regardless of the date the user asked for.

    Figure 5. Failure Mode 1, silent. Asked about the 2026-05 snapshot, the agent returns 78722 instead of 78751, with the wrong embedding.
    Figure 6. Failure Mode 1, explicit. Asked for the 2023-08 snapshot, the agent declines, saying its tool only uses the latest data.

    Failure Mode 2: Accepting Invalid Requests

    The agent was asked to answer questions based on invalid US ZIP codes, such as 00001. Instead of rejecting those outright, the agent accessed its PDI knowledge base and returned a link to a file with the response, without ever protesting the faulty ZIP codes. Thankfully, the file was empty, which means that the agent did not go as far as to hallucinate data and responses for such ZIP codes. It did, however, present that empty file as the requested embedding, which is why the grader marks the answer as fabricated. Still, this is a failure mode that exposes the disconnect between the agent’s natural language interface (which happily accepts the request and returns a response) and the tool that it uses to access the PDI embeddings and actually compute the response. When the tool fails to identify data for a given ZIP code, this should be clearly raised to the user.

    An example of this failure mode is shown in Figure 7.

    Figure 7. Failure Mode 2. Asked for invalid postal code 00001, the agent returns a link to a JSON artifact with empty embedding arrays.

    Failure Mode 3: Inability to Scale

    The agent was asked multiple consecutive questions in each simulated session. In many of these, it suddenly stopped producing answers and instead returned errors indicating that it had exhausted its context window. This is a sign of poor memory management: the agent never cleaned up or re-summarized its memory as it accumulated across consecutive NPC interactions. Eventually the context window filled up, leaving the agent unable to take on new requests. An example of this is shown in Figure 8.

    Figure 8. Failure Mode 3. Mid-session the agent returns a raw 400 INVALID_ARGUMENT error: history exceeded the one-million-token limit.

    A second sub-issue related to scalability emerged when we started multiple parallel simulations via VerifyAX. As the number of parallel threads increased, the agent reached its limit and started producing capacity errors. An example is shown in Figure 9. We addressed this issue by adding a queue that prevents the agent from getting overwhelmed. Obviously this can increase waiting times. Auto-scaling is another valid option that we apply in our agentic community. However, the queue remains a best-practice feature that allows the agent to operate within its hardware constraints without failing ungracefully (throwing errors).

    Figure 9. Failure Mode 3. Running in parallel, the agent returns a raw 429 RESOURCE_EXHAUSTED error; 10 of 20 questions went unasked.

    Failure Mode 4: Post-Processing Failures

    Due to the nature of the underlying PDI data, the correct answer to certain geospatial questions can be quite long. For instance, asking for the semantic embedding of one or more ZIP codes translates to sharing multiple numeric lists with hundreds of numbers. To avoid cluttering its output, the Geo agent is thus designed to return a JSON file that includes all the relevant info.

    The simulations revealed that, in many cases, even though the agent successfully computed the correct response and even returned a signed URL to the file with the supporting data, it was unwilling or unable to actually look at its own data and compute the final response, even if that required only a minor post-processing step.

    For example, when asked for the dimensionality of a ZIP code’s embedding (easy, since it’s always 330 for PDI embeddings), the agent correctly retrieved the vector but said that it was unable to count its dimensions. See Figure 10 for an example.

    Figure 10. Failure Mode 4 in VerifyAX. The NPC asks for the dimensionality and three largest-magnitude components of ZIP 90210’s embedding. The agent refuses, saying it cannot read those values from the signed URL its own get_pdfm_embeddings tool produced, even though the dimensionality is always 330 and the components are a direct read-off from the retrieved vector.

    When asked to first find the single ZIP code most similar to a target, then return that neighbor’s embedding, the agent refused the whole request. This task is two steps: find the nearest neighbor, then fetch its embedding. The agent said it could run the search and return a link to the results, but could not then open that link to read off the neighbor and look up its embedding. The agent could do each step on its own, but refused to connect them, inventing a limit that its own earlier runs prove is false. An example of this is shown in Figure 11.

    Figure 11. Failure Mode 4 in VerifyAX. The NPC asks for the single ZIP code most similar to 90210, then for that neighbor’s embedding. The agent refuses, saying it cannot open its own results link to read off the neighbor and look up its embedding, even though the answer (90049) is available and the same chain succeeded in earlier runs.

    Finally, the agent was asked how the drift ranking of a set of ZIP codes changes between a six-month window and a twelve-month window. Drift is how far a ZIP code’s embedding moves between two dates, measured as the distance between its two snapshots. To answer, the agent needs to fetch each ZIP code’s embedding at the relevant dates, work out its drift for each window, and compare the two rankings. Its retrieval tool can fetch every one of those snapshots, so this is just a few ordinary tool calls. Instead it declined, saying it could not compare the two windows “in a single step with the available tools”. This was a false limit, since the same kind of step-by-step work succeeded in other runs. See Figure 12 for an example.

    Figure 12. Failure Mode 4 in VerifyAX. The NPC asks whether ZIP 60614 drifted more over the last six months or the last twelve. The agent refuses, saying it cannot compare drift across two time periods in a single step with the available tools, even though its retrieval tool can fetch each snapshot and the correct answer (the six-month window moved more) is a short calculation.

    The Agent’s Current State

    Figure 13 visualizes the progress of the Geo agent after our first round of simulations and corresponding fixes. We have 18 different simulation bundles, each focused on exposing a different failure mode. The figure has one row for each bundle. The white dot marks the agent’s grade (out of 5, as per the VerifyAX scoring system) before our fixes. The blue dot marks the updated grade after applying fixes and re-running the bundle.

    Figure 13. Bundle scores at baseline and after the first round of verification. Each row is one scenario bundle of twenty graded questions.

    The Figure clearly visualizes the agent’s improvement across the board, and illustrates the value of continuous verification. The original version of the agent (white dots) was simply not ready for production. The updated version now scores almost perfectly (5/5) across the board, with only minor issues remaining in some of the bundles.

    What’s Next

    Every agent in our community is currently being onboarded and tested via VerifyAX, to enable continuous verification via simulation. The platform’s full verification functionality can also be accessed programmatically, which enables us to make it a part of our CI/CD pipeline and block faulty versions from being deployed into production.

    VerifyAX is not the only tool in our agentic verification toolkit. Simulation is a useful approach that allows us to safely stress-test our agents in a sandbox, rather than having them learn “on the job”. However, we are also experimenting with additional complementary tools that will allow us to install guardrails and detect regressions while the agents are serving real users and requests in production.

    References