This pod focuses on frontier topics in the rapidly expanding domain of LLM-powered agents. Examples of topics include multi-agent communities, collective intelligence, self-organization, trust and reputation, information propagation, agentic governance, agent economies, agent verification and safety, and others.
Author: Ted Lappas
-
From principles to practice: A governed multi-agent community for media campaigns
Agents are being built across WPP at remarkable pace and breadth. In every discipline, on every platform, by both technical and domain experts. This has already revolutionised the way we work. It has also surfaced a hard governance challenge that lurks underneath the excitement, and grows sharper as our agentic footprint expands. In a previous post, we set out the seven principles for agentic governance that every agent at WPP should satisfy. In this post, we show how these principles can work in practice, by applying them to diverse multi-agent community focused on media campaigns.
-
A Research Agenda for Expert Agent Communities
At WPP Research, we are experimenting with peer-to-peer communities of LLM-powered AI agents, each with its own persona and traits, its own abilities and data tools, and its own limitations and blind spots. No central coordinator sits above them, directing who does what. Agents discover one another, decide whom to trust, form teams, share and withhold information, and get work done by talking to their peers. This growing population of diverse expert agents is an ideal testbed for a wide range of exciting research questions.
-
Seven Principles for Agentic Governance: Freedom to Innovate, Confidence to Scale
Agents are being built across WPP at remarkable pace and breadth. In every discipline, on every platform, by both technical and domain experts. This is exactly what should be happening, and it has already revolutionised the way we work.
This rapid, distributed building surfaces a hard governance problem that lurks underneath the excitement, and grows sharper as our agentic footprint expands.
The first instinct in most governance efforts is to centralize everything: restrict which tools people can use, which platforms they can run on, which vendors they can work with. For something as versatile, novel and fast-moving as agentic AI, in a company as large and diverse as WPP, that approach simply would not work. It would stifle innovation and push teams to work around the rules, quietly finding their own comfort zone the moment the official options don’t fit their context.
In WPP Research, we have been deliberate about what shared agentic governance actually means. Rather than dictating how agents are built, we defined a small set of seven core principles that we believe every agent must satisfy to be considered compliant, regardless of who builds it and where it runs.
These principles are designed to answer the following seven questions:
- Is the agent registered and cleared to run?
- How is the agent built, and how has it evolved?
- What is the agent allowed to do, and what’s off-limits?
- How do we manage the agent’s traffic and resource use?
- Can we trace every action and event in the agent’s lifetime?
- Is the agent doing its job correctly, safely, and efficiently?
- What is the agent costing, and to whom?
Answering these questions consistently, across thousands of agents, spread over many teams, runtimes, clouds and clients, is genuinely hard. It is also a prerequisite for building agents we can rely on.
As long as we can accurately answer all seven of these questions for every agent that operates across our ecosystem, we can allow our community full freedom with regard to:
- The building layer: how individual agents are assembled and wired together into a working workflow. Different people prefer different building tools, and that variety is a point of differentiation rather than something to standardise away.
- The runtime: the engine on which the agent actually runs. In WPP Research, we utilize the Gemini Enterprise Agent Platform as the primary runtime for our agents. However, we believe that people across an org as large as WPP should be free to select and manage their own runtime, according to their technical and business requirements.
- The application layer: the apps, products and channels where agents are actually surfaced to and used by people: a chat window, a browser extension, an assistant embedded in a web app. Agents show up wherever the work already happens.
By letting the community choose their own runtimes, tools, and applications, the ecosystem can stay diverse and fast-moving while remaining governed centrally, delivering governance without sacrificing innovation or freedom of choice. This central governance rests on seven core principles, which we describe in detail next.
Principle 1 — Registration and identity
Is the agent registered and cleared to run?
Every agent should be enrolled in a common registry before it does a single piece of work, and none operate off the books. Having one registry answers a question that sounds trivial right up until the day you can’t answer it: which agents are actually running across our ecosystem right now?
Being in the registry means an agent exists. Identity means we know which one it is. Each agent is issued its own credential: a verifiable machine identity, not a shared login borrowed from a team or a generic service account. Every agentic action taken traces back to one specific, known agent.
Identity is a prerequisite for the other principles. Permissions are granted to it, every entry in the audit trail is stamped with it, and every unit of spend is attributed to it.
Principle 2 — Traceable, versioned agent DNA
How is the agent built, and how has it evolved?
Every component that affects an agent’s state and behaviour should be fully version-controlled with complete change history. Examples of such components include:
- the LLMs the agent uses (provider, version, inference parameters);
- the tools and integrations it can access, with their versions and schemas;
- the prompts and skills it uses;
- its knowledge and memory configuration — which data sources and vector stores it’s wired to, the retrieval settings, and the memory schema;
- the agent card defining its roles, boundaries and verification mechanisms;
- a history of the agent’s evaluations.
Respecting this principle allows us to track the configuration of each agent across time: who changed it, how, and when. Every change is attributed and timestamped, and each version is an immutable, addressable snapshot of the full configuration.
Why it matters: without versioned configuration, regression analysis is impossible. When an agent starts failing in flight, the first question is always what changed? Answering this question requires a precise, attributed diff between the last-good and current configuration.
Principle 3 — Access policy
What is the agent allowed to do, and what’s off-limits?
Every agent should follow the policy that was designed for it: a set of rules encoded as policy-as-code, authored and kept apart from the agent itself in a policy engine, versioned like source code, and bound to that agent’s identity. The policy engine’s job is to decide, not to act: for every request it returns a verdict: allow, deny, or escalate to a human.
An agent’s policy can, for example:
- restrict which data it may read, so it never reaches outside its allowed scope or across a client boundary;
- refuse any tool it hasn’t been explicitly granted;
- decide who may invoke it — which people, and which other agents;
- require sign-off from a named human before a high-stakes action, such as sending an output to a client or committing budget.
Keeping the policy separate from the agent is critical. Even if an agent malfunctions or is manipulated into attempting something it shouldn’t, it still can’t step outside its policy. The same rules are enforced identically, regardless of who is using the agent or which runtime it operates on. Policy versioning also allows us to always point to the exact rules an agent was operating under at any point in time.
Principle 4 — Traffic control
How do we manage the agent’s traffic and resource use?
Policy (Principle 3) decides whether a request is allowed. This fourth principle governs the flow of allowed requests: even a fully-permitted action still has to be routed, rate-limited and paid for. An agent is involved in different types of traffic:
- Calls coming in (a person through an app, or another agent, putting the agent to work): each caller gets its own rate limits. This prevents a busy user or an agent stuck in a loop from overwhelming the agent.
- Calls to a model (the agent reaching the LLMs that power its brain): these run through one common interface, so the choice of model becomes a control point. This choice may be automatic, human-approved, or pinned. The same interface caps token spend and prescribes alternatives when a provider goes down.
- Calls to tools and other agents (the agent leveraging external resources to act in the world): each resource has its own rate limits, so a greedy agent cannot drain or overwhelm it. The policy of Principle 3 determines which resources the agent can access; here we govern how much it may lean on them.
Across all three, the levers are the same: how a call is routed, how often it may fire, and how much it may spend. Together they keep a large, busy agentic ecosystem fast, affordable and stable.
Principle 5 — Event-level telemetry
Can we trace every action and event in the agent’s lifetime?
Every agent’s behaviour should be captured at the event level: every message sent or received, every tool call and its result, every retrieval and every other state-changing event is its own record. Together they form a trace: a step-by-step reconstruction of exactly what the agent did.
Each event carries both metadata and content. Examples of metadata include a timestamp and duration, the tokens and resolved monetary cost it incurred, and a link back to the exact agent version and session that produced it. The content is the substance of the interaction itself: prompts, responses, parameters. In WPP Research, we standardise on OpenTelemetry, so agents on any runtime emit the same shape of trace.
What is common is the format and query layer, not a single physical database. Traces can live in several stores: some separated by region for data residency, others kept inside a client’s own estate. A shared schema makes them behave as one virtual store, queryable as a whole. Every query still obeys the residency, segregation and privacy rules that govern the underlying data, and sensitive fields are redacted or kept local. A trace is complete enough to reconstruct what happened, without exposing what it shouldn’t.
Principle 6 — Continuous verification
Is the agent doing its job correctly, safely, and efficiently?
Every agent should be tested along two axes: functional (“does it do the job correctly?”) and non-functional (“is it fast, safe, robust, secure, reliable?”), in both development and production.
No agent should reach production without a recorded pass through a pre-production gate that captures both axes. The types of tests included in the gate are decided and continuously updated centrally at the org level.
Once live, agents should be continuously re-evaluated to catch regression or drift. All testing logs (e.g. testing platform, inputs, outputs, scores, findings) should persist to a queryable store, in which every verification run is also linked to the exact version of the verified agent (Principle 2) and the telemetry trace (Principle 5) that the run produced. This makes the question “did this change make the agent worse?” directly answerable and auditable.
Principle 7 — Usage and cost attribution
What is the agent costing, and to whom?
Every other principle looks at a single agent; this one turns that activity into an account. Principle 5 records the cost of each event and Principle 1 stamps it with an identity. This seventh principle rolls those events up along the dimensions a business actually cares about: by user, by team, by client, by organisation. So we can always answer who spent what, and on whose behalf.
This is what lets us bill and charge back accurately, catch runaway spend early, and see genuine adoption. Like telemetry, it operates at the the org level, because a spend-and-usage view is only useful if it spans the whole estate. It builds directly on the other principles: the per-event cost from telemetry (Principle 5), the enforced identity (Principle 1) that makes every event attributable, and the control points (Principle 4) where its budgets are enforced.
Limitations
It is important to acknowledge the limitations of our ongoing governance efforts.
Adoption is earned, not enforced. Applying governance principles across an organization as large and diverse as WPP is a real undertaking, especially now that spinning up a new agent has never been easier. The approach we suggest is not to police every build. It is to make compliant agents the most attractive option for every current and aspiring agent builder. For instance, an agent that complies with the seven suggested governance principles could enjoy benefits such as:
- Having its cloud/LLM cost fully or partially covered by the org, rather than by the specific team that built it.
- Gaining access to proprietary org data.
- Being surfaced to the org’s internal and client-facing apps.
Freshly built agent are likely to be in violation of at least some of the principles. However, as long as the path to compliance is well-documnted and properly motivated as a prerequisite for such benefits, the risks are contained without stifling innovation.
Exceptions. The seven principles are designed for agents built or fully controlled by members of an org like WPP, running on infrastructure that the org owns or controls. However, not every agent falls under this umbrella. For example:
- Personal coding assistants and other desktop-resident agents, typically embedded in user IDEs or running locally.
- Third-party SaaS with built-in agents, such as an agentic feature inside a product we subscribe to but don’t host.
- Consumer and public AI assistants used ad hoc, such as public endpoints and interfaces with no org instrumentation.
- Agents that run in environments not managed by the org, such as those operated by clients or partners.
For agents like these, the suggested approach is to document the applicability of each of principle, identify gaps and risks, and design policies to address them.
The principles reflect today, and today will change. They are shaped by the current AI landscape and the current reality of our organisation. Both will evolve. We treat these seven principles as a living standard, and we commit to evolving them as the landscapes also evolves.
-

Do you really need to ask? Identifying the surveys worth running
Every year, brand-tracking programmes field hundreds of surveys: the same battery of questions, asked across audiences and categories, wave after wave. It is the backbone of how brands understand where they stand. It is also expensive, and a surprising amount of it tells you what you could already have worked out.
That last part is the interesting bit. In a new study using WPP’s Brand Asset® Valuator (BAV) data, we turned a familiar research question inside out. Instead of asking “how accurate can our tracking be?”, we asked the question a marketing or finance director would ask: “how few surveys can we run and still have data we are prepared to act on?” The answer, on the programme we tested, is about two-thirds of them. This post explains how we got there, what the trade-off looks like, and what it means for the way you plan fieldwork.
The expensive habit nobody questions
Brand tracking tends to treat every market the same. Point the same instrument at every audience and category, run it, then run it again next wave. The bill adds up fast. Published figures for professionally run surveys land around $17 per sampled person and $41 per completed questionnaire, rising to $110 per complete for some designs. These are US academic figures, but the economics travel, and they climb every year as response rates fall and each completed interview takes more effort to win.
There is a quality tax on top of the money. The more often and the longer you survey people, the more they rush, guess and drop out, so over-surveying can quietly erode the very accuracy it is meant to buy.
Meanwhile, markets are not equal in how much fresh measurement they need. Brand perceptions are sticky in some places and volatile in others. Some of your surveys move a lot between waves. Some barely move at all, and their results could have been predicted from the rest of the programme with useful precision. Treating both kinds the same means spending money where the answer was already in the archive.
The key insight: accuracy is a line you set, cost is what you cut
Here is the shift in thinking that makes the rest work.
Most conversations about predicting survey results get stuck on one question: “is the prediction accurate enough?” Asked that way, the answer is always “it could be better, at a price,” and the conversation goes nowhere.
We flipped it. The brand team decides in advance the lowest accuracy it is willing to act on. Call it the good-enough line. Then the job is to find the cheapest fieldwork programme that clears that line, and to predict everything else. Accuracy is not the thing you chase. It is a floor you refuse to go under. Cost is the thing you minimise.

Figure 1: Set the bar, find the cheapest plan, and predict the rest.
Put that way, three moves follow naturally.
- Set the bar. For this study we set it at 0.75 on a standard accuracy measure (R², which runs from 0 for useless to 1 for perfect). That is a demonstration, not a recommendation. A tracker that feeds boardroom reviews can live with a lower bar than one that triggers spend market by market.
- Find the cheapest plan. A “survival of the fittest” search (a genetic algorithm) tries thousands of combinations of surveys and hands back a named list of the ones to keep fielding. Named matters. A percentage tells you how much to cut; a list tells you what.
- Predict the rest. A model trained on the surveys you did field fills in the ones you did not. Because every survey is described by the same handful of ingredients (audience, category, brand, attribute), the model learns how each ingredient shifts a brand’s scores and applies that to combinations it has never seen.
What each slice of budget buys you
The most useful thing the study produces is not a single number, it is a curve.

Figure 2: Prediction accuracy against the share of surveys fielded, showing the good-enough line at 0.75.
Read it left to right. Field 10% of your surveys and predictions are poor. Field 30% and they are already respectable. Field 60% and they are close to as good as they will ever get. After that, every extra survey buys less than the one before. On this programme, going from 63% of surveys to 100% of them lifts accuracy by under three points on a 100-point scale, and that final stretch is more than a third of the entire fieldwork budget.
This is the whole argument in one picture. The curve flattens, so the last few points of accuracy are the most expensive fieldwork you will ever buy. A programme that insists on the highest possible accuracy is paying its highest price for its smallest gain. Choosing to stop at the good-enough line is not a compromise on quality. It is the same logic you apply to every other budget you own.
The curve also lets you price any bar you like. Happy with 0.70? Field about 39% of the programme. Want 0.775? That will cost you about 91% of it. The conversation between research and finance can be had in one currency.
What we found
We tested this on the 2023 wave of BAV in the United States: 130 surveys (13 categories by 10 audiences), each scoring 10 brands on 46 attributes, nearly 60,000 data points in all. We held 26 surveys back as a hidden test set that the model never saw, so every accuracy figure below is on markets it had to predict cold.
- About a third of fieldwork is predictable. At the 0.75 bar, the cheapest plan fields roughly 63% of surveys and predicts the rest, a fieldwork reduction of about 37%. Fielding everything would reach only 0.78.
- The search returns a list, not just a number. On the training pool of 104 surveys, the genetic algorithm found a committable core of 81, a 22% reduction, with accuracy essentially unchanged (0.777 to 0.775). Being honest about this one: random selection can hit the bar with fewer surveys than the search found, because the search stops the moment it clears the line and does not keep pushing on cost. The search’s value is the named list you can act on, not a smaller number.
- The obvious shortcut does not work. We also tried filtering out the hardest-to-predict slices before compressing. On this data it made predictions worse at every threshold we tried. The model’s interaction features already capture what filtering would remove, so compression works at the level of whole surveys, which is also the level at which you actually buy fieldwork.
The trap: last year’s plan does not travel
Here is the finding that should change how you plan.
Take the surveys the search chose on 2023 and carry them, unchanged, into 2024 and 2025. Accuracy drops below the bar, to 0.73 and 0.69. Markets move, and a plan frozen on one wave goes stale on the next.
The fix is not to abandon the plan, it is to dilute it.

Figure 3: Accuracy rises as fresh surveys replace last year’s selections within a fixed 65-survey budget.
We held the budget fixed at 65 surveys and varied the mix. Spending all 65 on last year’s chosen surveys was the worst use of the money. Swapping just six of them for fresh surveys of the new wave lifted accuracy by five to six points. Swapping half was best, lifting it by eight points in both years at identical cost, and back above the bar. And the fresh surveys were picked at random. No clever selection needed; the model simply needs some measurement of the wave it is predicting.
So the deployment rule is simple. Under a fixed budget, reserve about half of it for fresh measurement of the wave you are planning, and let the model reconstruct the rest.
What this means if you run a tracker
Use it to plan, not to autopilot. Treat the predictions as a confident starting point that frees up budget, not a verdict that retires a market for good.
It is reallocation, not retreat. The point is not to survey less for its own sake. On these data, about a third of fieldwork spend could be redirected from predictable surveys towards the markets where fresh measurement demonstrably adds information, or simply saved.
Set the bar deliberately. The good-enough line is a management decision, not a modelling one. Decide it based on what the data is used for, and use the curve to see what it costs.
Refresh every wave. Do not carry last wave’s plan forward untouched. Mix in fresh surveys of the target wave, roughly half the budget on the evidence here.
Watch for shocks. A new entrant, a scandal, a viral moment can turn a stable market volatile overnight. The model cannot anticipate a structural break; only people watching the market can.
- It is reallocation, not retreat. The point is not to survey less for its own sake. It is to move a meaningful share of the budget — on these data, somewhere between a quarter and two-fifths of it — out of predictable slices and into volatile ones, where fresh measurement demonstrably adds information.
- Use savings strategically. Reduced fieldwork can mean lower cost, but it can also mean better coverage: expanding into new markets, adding new audiences, or going deeper where the evidence is most decision-relevant.
- Build in safeguards against drift. A slice that was predictable last year may not stay predictable forever. A new entrant, scandal, viral moment, campaign, or category disruption can change the market quickly. A deployed system should therefore include two safeguards. First, continue to field a small sample of otherwise “skippable” slices as a robustness check. Second, avoid using predicted values as if they were fresh observations in the next wave. In other words, do not let estimates compound into future estimates without being periodically re-anchored in real fieldwork.
- Use it to plan, not to autopilot. Treat the predictions as a confident starting point that frees up budget, not as a final verdict that retires a market for good.
The bottom line
Brand tracking does not have to be all-or-nothing, measure everything or fly blind. The data already in your archive can tell you, with useful precision, which surveys are worth re-fielding and which can be predicted from what you already know. And once you frame it as a budget decision rather than an accuracy contest, the answer to “why isn’t the prediction more accurate?” becomes obvious: because the last few points cost more than they are worth, and you chose not to buy them.
It is, in a sense, the mirror image of the synthetic-audiences debate. That conversation asks whether AI can replace the respondents. This one asks a more immediate and less risky question: of all the surveys you are about to run, which ones do you even need to? For roughly a third of them, the most honest answer is that you already have the data.
Sources and further reading
A curated subset of the work behind this post; full citations appear in the paper.
- Olson et al. (2024), “Examining variation in survey costs across surveys,” Sociological Methods & Research: the per-survey cost figures.
- Groves & Heeringa (2006), “Responsive design for household surveys,” Journal of the Royal Statistical Society A: survey cost as a constraint to be traded against accuracy.
- Bronnenberg, Dhar & Dubé (2009), “Brand history, geography, and the persistence of brand shares,” Journal of Political Economy: why some markets barely move.
- Prokhorenkova et al. (2018), “CatBoost: Unbiased boosting with categorical features,” NeurIPS: the recovery model.
- Yang et al. (2023), “Dataset pruning,” ICLR: why large fractions of training data are often redundant.
- Argyle et al. (2023), “Out of one, many: Using language models to simulate human samples,” Political Analysis: the synthetic-audiences counterpoint.
Disclaimer: This content was created with AI assistance. All research and conclusions are the work of WPP Research.
-

BAV Research Pod
The BAV Research Pod is a dedicated intelligence hub that leverages WPP’s BrandAsset® Valuator (BAV), the world’s largest, most enduring study of brand equity, to unlock advanced methodologies and strategic efficiencies. In its inaugural paper, the pod tackles the millions of dollars wasted annually on redundant brand-tracking surveys. By developing a predictive methodology that analyses historical BAV data, the team successfully simulated roughly half of the standard surveys with no loss in accuracy, even a year later. The result is a dynamic roadmap that tells brands exactly where consumer perception is moving, allowing them to eliminate redundant fieldwork and shift budgets to where measurement actually drives value.
-
The Power of Collaborative Semantic Embedding Spaces
One of the central challenges in modern marketing is no longer simply gaining access to more data, but gaining access to more usable intelligence, in more places, without forcing that data to cross boundaries it should not cross.
As identity signals continue to weaken, the industry needs a model that can preserve, and even enhance, the richness of consumer understanding while moving beyond reliance on limited individual-level identifiers. A shared embedding space for consumer intelligence offers a compelling path forward.
Embedding spaces
An embedding space turns any signal, whether text, image, video, behaviour or context, into a point whose location encodes its meaning, so related things naturally sit close together regardless of the form they arrived in. Each partner builds its own space from the patterns in its data. Aligning these independent spaces into a common frame is a hard technical problem, but it is one we know how to solve, allowing intelligence to flow between partners while the underlying data stays where it is.
A space like this can hold almost any kind of signal, from interests and behaviours to content, creative, outcomes, context, attention, geography and intent. And because meaning lives in position rather than in a shared key, these signals combine and reinforce one another without ever needing to match on an individual identifier. Advertisers, agencies, publishers and SSPs can each contribute what they know and draw on the collective picture, gaining sharper, more granular insight and stronger performance, even as traditional IDs fade away.
A shared space
This model becomes even more powerful when the same embedding space exists both centrally and at the edge, particularly on the publisher and SSP side. Central intelligence can provide strategic understanding, audience structure and planning logic, while edge intelligence applies that understanding at the point of impression, where decisions are actually made.
In this setup, publishers and SSPs do not need to send sensitive user-level data back to a central system, and central systems do not need to expose client or model intelligence in raw form. Instead, both operate within a shared semantic space, allowing impression-level signals to be interpreted with greater relevance and performance. In practice, this leads to better targeting, more effective contextual decisioning and stronger feedback loops, all without introducing the privacy risks associated with moving data across organizational boundaries.
A simple example.
An advertiser has learned what its most valuable customers look like, not as a list of individuals but as a region of the shared space defined by the interests, behaviours and contextual moments those customers have in common. That region is shared with a publisher’s edge model. When an impression becomes available, the publisher places the live context and on-site signals into the same space and simply measures how close they fall to the advertiser’s target region. A close match means the impression is a strong fit, so it can be prioritised and paired with the creative that best suits that region. No audience list changes hands and no user-level data leaves the publisher, yet the decision is as well informed as if both sides had pooled their data.
This is not a theoretical proposition for WPP. We have been developing these capabilities for many years and have products live in market today that share rich semantic consumer intelligence with publishers and SSPs, powering lossless activation with no sharing of consumers’ data.
In an agentic buying world, ensuring that the ground truth for both Buyer and Seller Agents can be trusted to be compatible and comparable increases trust and reduces inefficiency.
As the industry moves further beyond individual identifiers, this approach gives our clients and partners a durable foundation for understanding and reaching audiences, one that grows richer with every signal it encodes.
-

Evaluating Synthetic Audiences as Proxies for Human Users
Imagine running a focus group with ten thousand “people” overnight, for roughly the price of a few coffees. No recruiting, no scheduling, no drop-outs. Just ask your questions and get answers by morning.
That’s the promise of synthetic audiences: groups of AI agents, powered by the same large language models (LLMs) behind tools like ChatGPT, configured to stand in for human participants in surveys, experiments, and market research. The idea has exploded across marketing, psychology, economics, and design in just a couple of years.
In a recently published systematic review of over 100 related studies, we looked at what these AI “participants” can actually do, where they fall apart, and how to use them in practice. This blog post present a summary of the main findings. We also list some of the key references at the end of the post.
Why everyone is excited
The appeal becomes obvious the moment you look at the numbers. Studies that directly compared AI participants with human recruitment found cost differences of one to two orders of magnitude (that is, 10× to 100× cheaper per response). In one case, researchers generated tens of thousands of survey responses for roughly the API cost of a single query; in another, AI-generated open-ended answers for a user study cost a tiny fraction of what the equivalent human participants would have charged.
Beyond cost, there are practical superpowers:
- • No logistics headaches. An AI agent always answers, never gets bored, and never abandons your survey halfway through.
- • Experiments you otherwise couldn’t run. One team simulated a bank run using tens of thousands of AI-generated depositors across hundreds of demographic groups, a scenario you obviously can’t stage with real customers and real money. Others have built entire simulated societies: sandbox towns of AI “residents” that develop their own social dynamics, and large-scale models with over 10,000 agents interacting in urban and economic environments.
Where they actually shine
The honest headline from the research: synthetic audiences are good at capturing broad, aggregate patterns (the overall direction of an effect, rankings, and average tendencies), even when they miss the fine detail.
The more capable the model, the better this gets. Newer models have reproduced classic results from economics and social psychology, closely matched human behavior in trust games, and broadly agreed with human raters on things like brand perceptions. In other words, if your question is “which option do people lean toward?” an AI panel can often point you in the right direction, fast.
Where they break
The catch is that AI participants look convincing, and that’s exactly the trap. Fluent, human-sounding answers make it easy to assume there’s human-like thinking behind them. There usually isn’t. The review found four recurring failure modes, each backed by hard evidence.
1. They flatten the crowd. Real people disagree, surprise you, and sit at the extremes. AI participants tend to cluster around a bland “average.” One large replication study found an earlier model reproduced only 37.5% of a well-known set of psychology effects, versus a 50% rate for human samples. More tellingly, it kept giving near-identical answers to questions that normally produce lots of human variety. Other work documented a “hyper-accuracy” quirk, where AI participants give unrealistically precise, noise-free answers no real population would.
2. They carry hidden biases. Because they learn from internet-scale text, AI participants tend to sound disproportionately Western, English-speaking, wealthy, and politically skewed. One study comparing AI-simulated public opinion against a global survey found it performed far better for Western, English-speaking, high-income countries while systematically misrepresenting attitudes across gender, age, education, and class. Worse, when asked to portray specific groups, models often produce caricatures. Research on such groups found AI depictions that actual human participants judged not just inaccurate but actively harmful. The uncomfortable implication: synthetic audiences are least reliable precisely where fresh insight is most valuable, namely underrepresented and non-Western populations.
3. They’re twitchy. Reword a question slightly, reorder the options, or tweak the persona, and the answers can swing dramatically. Studies showed AI survey responses with unstable estimates and strong sensitivity to wording, persona, and even timing. In decision-making tasks, AI agents proved far more sensitive than humans to small “nudges”: a tiny change in framing produced outsized changes in behavior. That fragility makes results hard to reproduce.
4. They can fake reasoning. An answer that looks thoughtful may just be pattern-matching on things the model absorbed during training. Researchers showed that changing one variable in a study (say, a product’s price) can quietly shift the model’s unstated assumptions about everything else, contaminating any cause-and-effect conclusion. In another case, near-perfect economic “forecasts” turned out to be memorized history rather than genuine prediction. Add to this the finding that hallucinations are, mathematically, an unavoidable feature of these models, not a bug that will simply be patched away.
How researchers are fixing them
The field isn’t standing still. Three families of techniques are making synthetic audiences more trustworthy:
- • Better prompting. Giving the AI richer personas (detailed demographics, life-story backstories, even real interview transcripts) measurably improves how well it matches a target group. The gains are real but fragile, since the same prompt-sensitivity problem lurks underneath.
- • Fine-tuning. Training models on real human responses for a specific population improves accuracy, though it tends to compress diversity unless done carefully.
- • Human-AI hybrids, the most reliable approach. Instead of replacing people, you collect a small human sample and use it to calibrate and correct the AI’s output. A statistical framework called prediction-powered inference makes this rigorous, and studies report cutting the required number of human participants by 20% to 30% in some settings, while another approach reduced the human data needed by up to 80%, all without sacrificing statistical validity. You keep much of the speed and cost savings while keeping the results honest.
The promising road ahead
What makes this field genuinely exciting isn’t just today’s tools; it’s where the research is heading. The review highlights three directions that could turn synthetic audiences from a clever shortcut into a dependable scientific instrument.
1. Open, transparent models built for human behavior. Most studies so far rely on closed, proprietary systems whose training data and updates are secret, which makes results hard to audit or reproduce (a “silent” update can change behavior between studies). Open models change that. The standout example is Centaur, a model fine-tuned on over 10 million real human decisions across 160 psychology experiments. It outperformed traditional models at predicting how people actually behave and generalized to new tasks, offering a glimpse of purpose-built, inspectable “model populations” rather than borrowed chatbots.
2. Standards, benchmarks, and reproducibility. Right now, every team prompts differently, making studies hard to compare. The field is moving toward shared benchmarks and tools, such as standardized environments for running AI-based surveys and experiments, so researchers can tell whether an improvement truly generalizes or just got lucky on one prompt. This is the unglamorous infrastructure that turned modern machine learning into a cumulative science, and synthetic audiences need it too.
3. Looking inside the black box. The frontier is mechanistic interpretability: tools that trace which internal “circuits” of a model produce a given behavior. The payoff would be enormous: imagine pinpointing the exact components that trigger a demographic stereotype and dialing them down, or verifying that a persona prompt genuinely activates the right knowledge rather than just changing the writing style. This work is early and mostly aspirational today, but it points toward synthetic audiences we can actually diagnose and correct, rather than merely observe.
Best Practice
If you’re tempted to use AI participants, the research suggests:
- • Evaluate the fit. Research shows promising evidence for early exploration, pilots, brainstorming, broad directional questions, and well-represented (Western, English-speaking) populations. Still a poor fit for fine-grained individual differences, underrepresented or non-Western groups, and high-stakes domains like policy or healthcare.
- • Explore with AI, validate with Humans. Even a modest dose of real-world data, used to check the AI’s answers, dramatically improves fidelity.
The bottom line
Synthetic audiences are one of the most exciting new tools in research, but they’re a tool, not a replacement for people. They’re best understood as a fast, cheap way to explore ideas and generate hypotheses, with their results treated as a starting point to be validated, not a final verdict.
Used carefully, they can democratize research and let small teams ask big questions. Used carelessly, they produce confident-sounding answers that quietly reflect an AI’s blind spots rather than reality. The difference comes down to knowing what they can and can’t do, and increasingly on the open models, shared standards, and interpretability tools that will determine how far this technology can really go.
This post summarizes the systematic literature review “Synthetic Audiences as Proxies for Human Users,” published in IEEE Access, which analyzed 100 studies on LLM-based synthetic audiences following established review protocols. It’s written for a general audience; see the full paper for methods, evidence, and the complete set of 100 citations.
Read the full paper: Lappas & Filippas, “Synthetic Audiences as Proxies for Human Users: A Systematic Literature Review,” IEEE Access, 2026. https://ieeexplore.ieee.org/document/11563566/
Sources & further reading
The specific findings mentioned above come from the studies below (a curated subset of the 100 reviewed). Links are provided where a stable one exists.
Applications, scale, and cost
- Argyle et al. (2023), “Out of one, many: Using language models to simulate human samples,” Political Analysis. https://doi.org/10.1017/pan.2023.2
- Hämäläinen et al. (2023), “Evaluating large language models in generating synthetic HCI research data,” CHI. https://doi.org/10.1145/3544548.3580688
- Kazinnik (2024), “Bank run, interrupted: Modeling deposit withdrawals with generative AI,” SSRN. https://doi.org/10.2139/ssrn.4656722
- Park et al. (2023), “Generative agents: Interactive simulacra of human behavior,” UIST (the “Smallville” study). https://doi.org/10.1145/3586183.3606763
- Piao et al. (2025), “AgentSociety: Large-scale simulation of LLM-driven generative agents,” https://arxiv.org/abs/2502.08691
- Aher et al. (2023), “Using large language models to simulate multiple humans and replicate human subject studies,” ICML. https://proceedings.mlr.press/v202/aher23a.html
- Xie et al. (2024), “Can large language model agents simulate human trust behavior?,” NeurIPS. https://arxiv.org/abs/2402.04559
- Li et al. (2024), “Frontiers: Determining the validity of large language models for automated perceptual analysis,” Marketing Science. https://doi.org/10.1287/mksc.2023.0454
Limitations
- Park et al. (2024), “Diminished diversity-of-thought in a standard large language model,” Behavior Research Methods (the 37.5% vs 50% replication finding). https://doi.org/10.3758/s13428-023-02307-x
- Qu & Wang (2024), “Performance and biases of large language models in public opinion simulation,” Humanities and Social Sciences Communications. https://doi.org/10.1057/s41599-024-03609-x
- Cheng et al. (2023), “CoMPosT: Characterizing and evaluating caricature in LLM simulations,” https://arxiv.org/abs/2310.11501
- Wang et al. (2025), “Large language models that replace human participants can harmfully misportray and flatten identity groups,” Nature Machine Intelligence. https://doi.org/10.1038/s42256-025-00986-z
- Gadiraju et al. (2023), “‘I wouldn’t say offensive but…’: Disability-centered perspectives on large language models,” FAccT. https://doi.org/10.1145/3593013.3593989
- Bisbee et al. (2024), “Synthetic replacements for human survey data? The perils of large language models,” Political Analysis. https://doi.org/10.1017/pan.2024.3
- Cherep et al. (2025), “LLM agents are hypersensitive to nudges,” https://arxiv.org/abs/2505.11584
- Gui & Toubia (2025), “The challenge of using LLMs to simulate human behavior: A causal inference perspective,” https://arxiv.org/abs/2312.15524
- Xu et al. (2025), “Hallucination is inevitable: An innate limitation of large language models,” https://arxiv.org/abs/2401.11817
Improvement techniques and human-AI hybrids
- Moon et al. (2024), “Virtual personas for language models via an anthology of backstories,” https://arxiv.org/abs/2407.06576
- Suh et al. (2025), “Language model fine-tuning on scaled survey data (SubPOP),” https://arxiv.org/abs/2502.16761
- Angelopoulos et al. (2023), “Prediction-powered inference,” Science. https://doi.org/10.1126/science.adi6000
- De Bartolomeis et al. (2025), “Efficient randomized experiments using foundation models,” https://arxiv.org/abs/2502.04262 (the 20% to 30% figure).
- Wang et al. (2024), “Large language models for market research: A data-augmentation approach,” https://arxiv.org/abs/2412.19363 (the up-to-80% figure).
Future directions
- Binz et al. (2024), “Centaur: A foundation model of human cognition,” https://arxiv.org/abs/2410.20268
- Shapira et al. (2024), “GLEE: A unified framework and benchmark for language-based economic environments,” https://arxiv.org/abs/2410.05254
- Sharkey et al. (2025), “Open problems in mechanistic interpretability,” https://arxiv.org/abs/2501.16496
-

Learning through experience: teaching the viral Hermes agent to automate our work

Hermes What is Hermes ?
Hermes Agent is an open-source, self-hosted AI agent released by Nous Research, the lab behind the Hermes model family under an MIT license. Its main pitch is a built-in learning loop. Instead of resetting to zero every session, Hermes runs a post-execution review after each successful task, distills the steps that worked into a reusable, Markdown-defined “skill” and refines those skills the next time it hits a similar problem. It also keeps persistent memory across sessions, so it gradually builds a model of your projects and how you like things done, effectively learning through experience.
Unlike a copilot tethered to an IDE, Hermes is meant to live on a server and run unattended – a $5 VPS, a GPU box or serverless infra that costs almost nothing when idle. It talks to hundreds of LLMs through the OpenAI-compatible interface, can communicate via Telegram, Discord, Slack, WhatsApp, Signal, email and a CLI, supports natural-language cron for scheduled jobs, and can spin up subagents to parallelize work. It also ships with 40+ built-in skills out of the box.
What made it especially attractive to us is its ability to write and improve its own playbook. This ability made us think: could it learn how to do our job and fully automate the work we do at the Tech Pulse pod?
Our pod’s mission is to pick up state-of-the-art AI tools, experiment with them, and write an honest, evidence-backed assessment.
What makes our work tricky is the fact that every new tool we evaluate is different. We need to study it, figure out how to set it up, run it, and evaluate the results.
Our 6-step evaluation protocol
We begin by encoding our workflow into a strict, 6-step protocol that the agent can follow for every evaluation:
- Workspace initialization – spin up a clean, isolated project environment so each evaluation starts from a known state.
- Baseline replication – run the tool’s own “getting started” examples first, to confirm the headline claims reproduce before we push further.
- Rigorous verification – design and run additional autonomous tests that probe the tool under conditions its authors didn’t pick, rather than extrapolating from the happy path.
- Data synthesis & metrics control – measure the things that matter (recall, latency, accuracy) against ground truth data.
- Adversarial peer review – hand the findings to a separate agent powered by a powerful LLM (Claude Opus 4.8), to receive feedback and iterate.
- Finalization & delivery – format the assessment doc and ship it to the right channel (to our internal Notion knowledge base in our case).
The idea was simple: if these are the steps a human pod member walks through, can an agent walk through them unattended – and would the writeup at the end be any good?
The Skill system: how Hermes improves itself
The first time we ran Hermes on a real task, we walked it through the 6-step protocol by hand. Instead of just following along, Hermes wrote each step down as its own skill – a short Markdown file it can pull up at the start of any future run. So the protocol stopped being something we had to repeatedly provide as input; it became something the agent already knows.
Hermes used this knowledge to build a small set of skills covering everything we do: how to set up a clean workspace, how to test a new tool properly, and how to draft the write-up and run it past the reviewers. On top of those sits one master skill that holds the whole run together – it treats the 6 steps as a checklist and won’t let Hermes jump ahead before the previous step is actually done.
Impressively, when the LLM peer reviewer flagged something during an experiment (e.g. an unfair baseline, a missing caveat, a poorly designed experiment), Hermes would learn from the feedback and address the issue. It would also record its new learnings in the related skill files, so the next experiment started from a slightly stronger playbook.
That’s the self-improvement loop the Hermes pitch promises, and we actually watched it happen. The more experiments we run, the better the skills get, and the less we need to babysit Hermes for the next one. The underlying LLM that powers Hermes (Gemini 3.1 Pro) isn’t getting smarter – its playbook is, and Hermes is the one rewriting it.
Putting Hermes to the test
We gave Hermes a single, real assignment: take a brand-new open-source tool called Turbovec – which claims to store huge amounts of data in a tiny amount of memory and search it faster than the popular alternative – and find out whether those claims actually hold up.
We handed the agent the tool, its documentation, a bare cloud machine to work, and nothing else: no starter code, no template, no outline. Hermes had to decide what and how to test, run the experiments and, write the whole thing up on its own.
We reviewed Hermes’ output in the exact same way we would review a human colleague’s work, via three simple questions:
- Did it manage to set up the tool and run experiments? Yes! It wasn’t all smooth sailing. Hermes’s first pass used the wrong settings. One of the integrations that TurboVec advertised also didn’t work on the first try – a common challenge we face in the Tech Pulse Pod. However, rather than getting blocked by these stumbles, the agent noticed them, fixed them and left a clear trail of what went wrong and how it was corrected – exactly the kind of thing a rushed human reviewer might quietly skip over.
- Did it design a fair test? Yes! It first reproduced the tool authors’ own results, then set up an even-handed comparison against the leading alternative tool (FAISS), as a baseline. It was also careful enough to optimize the baseline tool’s configuration (rather than making it deliberately weak one), ensuring that the contest wasn’t rigged in Turbovec’s favor.
- Did it get the numbers right? Yes, with a bit of extra AI help. For inststance, its first attempt used a small data sample and then just assumed the results would scale up neatly. The Adversarial peer-review step (step 5 in our 6-step workflow) caught that this assumption was unsafe. Hermes accepted the criticism and re-ran the full-size test. The adversarial reviewer turned out to be right – using the small sample would have significantly skewed the results.
- Was the writeup appropriate? Yes, with a bit of extra AI help. Hermes’ original draft omitted some critical details and also inflated some of the findings. Thankfully, The LLM reviewer’s feedback in step 5 also ensured that claims got toned down to what the data actually supported.
So is it worth it?
Very promising. Left alone with a new GitHub repo and a blank machine, Hermes handled the mechanical, time-consuming work on its own: it downloaded the repo, installed everything, and ran some initial tests to make sure everything runs.
Even though it did stumble during the actual experimental design and assessment of the tool (TurboVec), the introduction of our second Reviewer agent was enough to address these issues and deliver, with no human in the loop.
Obviously this is just a single – albeit very promising – piece of evidence. We will keep pushing the limits of Hermes in the context of our pod’s work, with the intent to automate and scale-up our assessment efforts as much as possible.
Another angle we plan to explore is cost minimization. This first experiment showed the effectiveness of the iterative, dual-agent architecture:
- An affordable Gemini-powered Hermes to actually do the heavy lifting (open-ended, token-heavy)
- A more expensive Opus-powered Reviewer to review the report after each iteration and provide feedback (single-shot, token-lean)
A key question – and one that we keep facing in WPP Research – is: what is the cheapest LLM brain that we could use for each agent, while maintaining quality outcome?
-
Using DeepSeek v4 Pro as an Agentic Brain
1. The context
DeepSeek’s release pattern has been consistent: ship a model that posts frontier-comparable benchmark numbers at an order-of-magnitude lower price, then watch the other providers scramble. DeepSeek v4 Pro is the most aggressive instance of that pattern yet. Rather than putting it to the test via a standard off-the-shelf agentic benchmark, we focused on a more pertinent question: what happens when you actually deploy it as the LLM brain of an agent?
“Agentic ability” packs in at least three dimensions:
- Behaving like a competent professional in realistic settings. Stay in character. Stay focused on your objective. Communicate effectively with different types of stakeholders. Respect policies and constraints. Validate the quality of the information that you consume and produce. Adapt to changing circumstances.
- Long-horizon autonomous problem solving. Read a codebase, form a hypothesis, run an experiment, read the result, build on it. Repeat many times without a human in the loop.
- Cost-efficiency under sustained load. Being able to solve complex problems and succeed in real-world scenarios for $0.05 per task is a very different proposition from achieving the same outcomes for $1.50 per task, even if the success rate is identical.
To evaluate DeepSeek across all three dimensions, we picked two different agentic tools:
- VerifyAX (Conscium’s agent-evaluation platform). VerifyAX drops an agent into the kind of situation it would actually face once deployed: a realistic scenario populated by other characters (customers, interviewers, colleagues, adversaries), with an objective to achieve and rules to follow. Scenarios are automatically generated to exercise a wide panel of specific skills, from communication and safety to technical reasoning.
- Autoresearch (Andrej Karpathy’s open-source autonomous-research framework). It hands an agent a piece of code and a metric to improve, then steps back. The agent reads the code, makes one change, runs it, checks whether the metric improved, and decides whether to keep or undo the change. This continues for many iterations, with no human in the loop.
Cost (the third dimension) can be easily measured on both VerifyAX and Autoresearch, so we report it separately.
To benchmark DeepSeek v4 Pro against the rest of the frontier, we set up four agents, each powered by a different LLM: DeepSeek v4 Pro, GPT 5.5 (OpenAI), Claude Opus 4.6 (Anthropic), and Gemini 3.1 Pro (Google). The agent harness is identical across all four; only the LLM brain changes.
2. Experiment A: VerifyAX
Setup
We ran all four LLM-powered agents (powered by Opus 4.6, Gemini 3.1 Pro, DeepSeek v4 Pro, and GPT 5.5) through two VerifyAX-generated scenario suites:
- The harder suite — 25 multi-agent scenarios, in which the agent is given a fixed amount of time to navigate multiple stakeholders (manifested as VerifyAX-controlled agents), exchange information, use tools, and make decisions in order to achieve an objective. Each scenario is designed to test for specific skills.
- The easier suite — 25 one-on-one scenarios. Similar to the harder suite, but with only one stakeholder in each scenario.
Every run produces a pass/fail verdict per scenario plus a per-skill score (1–5) for every skill the scenario tests for. Across the 50 scenarios, the two suites collectively exercise 60+ distinct skills, drawn from a wider VerifyAX panel that covers safety, reasoning under pressure, information extraction, communication skills, and technical execution. A representative sample is shown in the table below:
Skill tag What it tests goal_injection_resistanceResists being hijacked into a new, conflicting mission mid-scenario. incremental_corruption_resistanceNotices when a chain of seemingly-harmless requests cumulatively breaks the rules. sycophancy_resistanceRefuses to validate a demonstrably false statement just because the counterparty insists on it. hate_speech_refusalRefuses to produce hate speech or discriminatory content. irreversible_action_cautionInsists on explicit informed consent before taking unrecoverable actions. ambiguity_clarificationPins down vague jargon or undefined terms before acting. contradiction_handlingCalls out inconsistencies in a counterparty’s statements tactfully, without triggering defensiveness. tradeoff_reasoningWeighs competing options and articulates the cost of each choice. data_hallucination_resistanceAsks for missing materials instead of inventing them when context is incomplete. long_horizon_thinkingPlans and sequences actions across many steps toward a distant goal. anger_deescalationAcknowledges emotion, defuses conflict, proposes concrete next steps. intermediate_mathSolves multi-step quantitative problems. advanced_web_researchAnswers complex questions requiring browsing, cross-referencing multiple sources, and synthesis. advanced_programmingSolves complex programming problems. Result
Harder suite (multi-agent):
Model Pass rate Avg skill grade (1–5) Agent cost $ / scenario Claude Opus 4.6 13/25 (52%) 4.63 $19.95 $0.80 Gemini 3.1 Pro 8/25 (32%) 4.42 $1.81 $0.07 DeepSeek v4 Pro 7/25 (28%) 4.04 $0.48 $0.02 GPT 5.5 7/25 (28%) 4.25 $1.79 $0.07 Easier suite (one-on-one):
Model Pass rate Avg skill grade (1–5) Agent cost $ / scenario Claude Opus 4.6 23/25 (92%) 4.83 $17.09 $0.68 Gemini 3.1 Pro 23/25 (92%) 4.71 $0.99 $0.04 GPT 5.5 20/25 (80%) 4.52 $1.19 $0.05 DeepSeek v4 Pro 19/25 (76%) 4.58 $0.30 $0.01 Three things jump out:
- Claude wins, comfortably. On the harder suite, 52% vs everyone else clustered at 28–32% — and the highest macro-averaged skill grade (4.63) of any model on either suite. On the easier suite, Claude and Gemini tie on pass rate at 92%, but Claude still edges Gemini on skill grade (4.83 vs 4.71).
- DeepSeek sits at the bottom on both suites. On the harder suite it ties GPT 5.5 at the bottom of the table (both 7/25, 28%). On the easier suite it’s the weakest of the four on pass rate (19/25, 76%), though its skill grade (4.58) actually edges GPT’s (4.52).
- The cost spread is startling. Claude’s per-scenario spend is ~40× DeepSeek’s on the harder suite and ~55× on the easier one. GPT 5.5 and Gemini sit in the same mid-range bracket (~$0.04–0.07/scenario); only Claude is in a different tier.
3. Experiment B: Autoresearch
Setup
Each agent starts with just two files:
train.py— a ~630-line script that trains a small language model from scratch. The starting model is small by today’s standards — 8.7M parameters, the kind of tiny transformer you’d find in an early GPT-2.program.md— a short prompt telling the agent what to do.
The agent then runs unattended, looping through these steps:
- Read the current training script and a log of everything it has tried before.
- Propose one specific code change.
- Train the modified model on GPU hardware for a fixed 5-minute budget.
- Show the trained model a chunk of text it has never seen and measure how well it predicts what comes next, character by character.
- Keep the change if the new score is better than the best so far; otherwise discard it and try something different next time.
- Repeat 50 times.
Result
Model % Improvement Cost Successful experiments GPT 5.5 6.83% $33.53 8 / 50 Gemini 3.1 Pro 6.12% $10.21 9 / 50 DeepSeek v4 Pro 6.11% $3.69 8 / 50 Claude Opus 4.6 6.04% $63.45 7 / 50 The improvement numbers cluster tightly. The cost numbers do not: Opus 4.6 cost ~17× more than DeepSeek v4 Pro for essentially the same outcome.
More interesting than the bottom line is the strategy fingerprint each model converged to. All four independently rediscovered the same single biggest win: cutting the training batch size in half (which trades smaller-per-step learning updates for more update steps inside the 5-minute budget — a good trade when the bottleneck is wall-clock, not data). After that they diverged:
- Opus 4.6 kept the network’s outer shape and redesigned the building blocks inside each layer — a more expressive math operation in every block (an activation function called SwiGLU, instead of ReLU²) and 50% more internal capacity per layer. Same outside, smarter inside.
- GPT 5.5 opted for a smaller, faster network (3 layers instead of 4, shorter context windows) so it could fit more training steps into the budget, with optimizer settings tuned to make those extra steps count.
- DeepSeek v4 Pro combined GPT’s move with the only attention-mechanism change that survived in any model’s final config: grouped-query attention (reusing key/value projections across heads to compress the attention block).
- Gemini 3.1 Pro left the network alone and changed how it was trained — same layers, same shape, same building blocks, but turned learning rates up and drove weight decay to zero. Every architectural change it tried, it reverted.
Those architectural moves had visible consequences for the final model size: Opus’s wider MLP made the model bigger than the 8.7M-parameter baseline, Gemini kept it at baseline size, GPT shrank it, and DeepSeek shrank it most — to 3.4M parameters, less than half the baseline.
Four genuinely distinct strategies landing within 0.8 percentage points of each other is itself an interesting result, suggesting there are several different ways to win at this task within the 5-minute budget, and that experimenting with different models can lead to distinct but equally promising paths. Increasing the rounds and the time budget per round can help explore these paths further.
The Bottom Line
On Autoresearch, where the LLM-powered agent is the only stakeholder and the loop is a tight code-edit / measure-result cycle, DeepSeek is tied with the top model on outcome (6.11% improvement, within 0.8 pp of GPT 5.5) at a fraction of the cost. On VerifyAX, where the agent has to survive multi-agent simulations that emulate real-world scenarios, it maintains its cost advantage but lands at the very bottom of the rankings in terms of skills and objective completion. This highlights something we already knew: the key is to pick the right tool (LLM brain) for the job. If you care about cost and expect your agent to work on a problem on its own for many iterations, then evidence suggests that DeepSeek is definitely worth a shot. However, if you expect your agent to operate in dynamic real-world environments with other stakeholders and complex constraints, then there are better options out there.
If you are looking for an LLM Brain that can perform in both contexts, Gemini is the most consistent of the four. Always in the top half, ties Claude on the easier VerifyAX suite at a fraction of the cost, and finishes Autoresearch essentially tied with DeepSeek. The pragmatic default if you don’t want to commit to either end of the cost/capability spectrum.
Our experiments are just two of many that could (and should) be run to evaluate a model as powerful and multi-faceted as a frontier LLM. They offer some real evidence about where each LLM lands in the contexts we texted, but there’s plenty more work to do to establish how it performs in other settings.
