From principles to practice: A governed multi-agent community for media campaigns

Written by

in

Agents are being built across WPP at remarkable pace and breadth. In every discipline, on every platform, by both technical and domain experts. This has already revolutionised the way we work. It has also surfaced a hard governance challenge that lurks underneath the excitement, and grows sharper as our agentic footprint expands. In a previous post, we set out the seven principles for agentic governance that every agent at WPP should satisfy. In this post, we show how these principles can work in practice, by applying them to diverse multi-agent community focused on media campaigns.

Click here for a video walkthrough of how we build and govern a community of multiple expert agents.

A community of expert agents

The community is made up of specialist agents, each focused on a different task in the media-campaign lifecycle. Crucially, each agent is an independent member of the community: it is built and maintained by a separate team, lives in its own repository, and is deployed independently from the others. The agents operate autonomously and communicate with one another entirely via A2A, the agent-to-agent protocol. There is no monolith behind the scenes, just a set of experts that discover and call on one another as the work demands (as seen in Figure 1).

The community at a glance: independent expert agents that coordinate over A2A.

The current members of the community are:

  • ▸Campaign designer. Receives a campaign brief, enriches it by consulting the other agents in the community, and uses the enriched brief to design a campaign. It is the agent most users talk to first.
  • ▸Brand Agent. Reports how consumers actually perceive a brand: how they view it, feel about it, and value it. It can answer questions like “Do high-income millennial consumers in Germany perceive Ford as a reliable brand?” It is powered by Tavily for analysing brand websites, Apify for inspecting social media, and WPP’s BAV platform, the largest and most comprehensive brand analytics platform in the world.
  • ▸Geo Agent. Provides demographic, economic, and other information at the zip-code level, answering questions such as “Which zip codes are most semantically similar to 10005?” or “Which zip code’s profile has changed the most in the past 6 months?” It is powered by the work we have done with Google via the Earth AI initiative, and specifically enabled through Google’s PDFM geo-embeddings.
  • ▸Zeitgeist Agent. Surfaces the trends and events shaping a given area, industry, and timeframe, using Google News and Trends. For example: “What relevant trends and events shaped the sportswear market in New York last month?”. It gives a campaign its current spatiotemporal context.
  • ▸Audience Agent. Takes a campaign brief plus the context gathered by the other specialists and recommends fitting target audiences, for example “Recommend 3 audiences for an Instagram campaign focused on awareness, with a $5,000 budget, targeting 90210, over September 1–15, 2026.” It leverages Google’s new Gemini Deep Research agent, provided to us through our Early Access Programme.
  • ▸Creative Agent. Generates images and videos from a prompt using Gemini and Veo, for example “Generate 2 tailored video concepts optimised for clicks on TikTok for L’Oréal skincare products.”
  • ▸Performance Agent. Evaluates a given campaign configuration and returns the expected media performance, drawing on the extensive campaign-performance modelling we have built in WPP Research and Open Intelligence. Its underlying model is trained on Open Intelligence data.
  • ▸Activation Agent. Simulates campaign activation and outcome measurement. Because the community is not permitted to publish and track live campaigns, this agent acts as a proxy for the market and lets us close the loop, responding to requests such as “Please activate the campaign” or “Get me the measured outcomes for the campaign.”

Together, these agents form a diverse expert community, which turns out to be the ideal environment for exercising every governance principle at once.

The community in action

Users interact with the community through a simple chat interface (Figure 2). Behind that single entry point, the community can handle everything from a quick lookup to the design of a full campaign, calling on whichever members it needs.

The single chat entry point to the whole community.

Figure 3 shows a simple request that can be covered a single agent. Asking “What information do you have about the brand?” routes to the Brand agent, which returns a summary of the brand’s characteristics grounded in BAV data:

A single-agent response: the Brand agent summarising a brand.

 

Figure 4 shows how this same request is received and viewed from the point of the view of the Brand Agent itself.

The same exchange from the Brand agent’s side in the playground.

A more complex request engages multiple agents. Asking the community to identify where the dimensions for which a brand’s perception is weaker than that of a competitor, then create an Instagram image for a US summer awareness campaign that addresses those weaknesses. calls the Brand agent to compare perception metrics and the Creative agent to generate an image that speaks directly to the gaps it found, as shown in Figures 5, 6 and 7.


Two agents in concert: the perception gaps the Brand agent found, and the image the Creative agent generated to address them.
Detailed report from Brand agent comparing two brands.
Tracing the two-agent exchange from each agent’s point of view.

A full brief engages the whole pipeline. Given a structured brief (brand, objective, budget, geo, timeframe, creative type, platform), the Campaign Designer picks up the task, coordinates with the community, and returns a task id so the user can track progress. A few minutes later, the community returns the optimal audience-and-creative combination for the brief, where “optimal” means maximising the expected media performance as scored by the Performance Predictor agent. From there, the user can activate the campaign through the Activation Agent and retrieve its (simulated) outcome metrics:

The community’s collective output: the optimal audience-and-creative combination for the brief.

Governance in practice

The same community, viewed through the lens of the seven principles, shows how each one is implemented today, and where we are still building.

1Registration and identity

Is the agent registered and cleared to run?

Every member of the community is enrolled in a single registry, implemented with Google’s Agent Registry (see Figure 9 and 10), a component of the Gemini Enterprise Agent Platform. Each agent is assigned a unique identifier, and its full profile is inspectable in one place. Adding a new agent is a matter of registering it and letting its details populate from its configuration, provided it is set up correctly for privileges and authentication.

Agent Registry with a list of enrolled agents, each with a unique identity and full profile.
View on single agent profile in Agent Registry.

2Traceable, versioned agent DNA

How is the agent built, and how has it evolved?

Every component that shapes an agent’s behaviour is stored and version-controlled. The primary component is the agent’s source code, which lives in a dedicated code repo. Figures 11-12 show the repo of the Zeitgeist agent, as an example. The repo stores the agent’s code, its agent card, its prompts templates.  Everything in one place, with every change recorded as a commit that carries a date and a name.

 

The Zeitgeist agent’s DNA in version control: agent card, prompts, tools, and the commit hash tied to each deployment.
View of the Zeitgeist’s agent card in the repository.

Even though the code in the repository is the first factor that can influence an agent’s behaviour, it is not the only one. There are many more. Principle 2 is about keeping track of all of the factors that together make up the agent’s full DNA for each deployed version. Other factors include:

  • ▸The exact deployed prompts that instruct the agent on how to behave. Prompts are often designed as flexible templates that are filled based on various settings from local .env files that are not tracked by the repo. For instance, a setting might inject different content or enforce a stricter output format for deployment in a production versus a testing environment.
  • ▸The LLM Brain. Which provider, which model, and the settings it runs with. We could hard-code a specific model with specific parameters inside the repository, but that is not best practice. Instead a central proxy, LiteLLM, routes different kinds of request from different agents so that cost and performance can be tuned in one place. We come back to this later in Principle 4, Traffic Control.
  • ▸The tools. The third party tools the agent calls. Their code sits in other repositories that we cannot see or monitor, so what gets recorded is a fingerprint of the tool the agent actually ended up holding, and for a tool on a remote server, what that server said it offered.
  • ▸The published agent card. The card file does indeed live in the repository, like the code. However, the card could also be quietly changed and published in a way that bypasses the repo, so we need to keep track of the published version at all times.
  • ▸The data sources. Which sources the agent is pointed at and what it may do with them. If one of them changes, or becomes compromised, the agent’s answers change with it.
  • ▸The dependencies. The third party libraries the agent imports and uses: numerical computation, memory management, anything else. If one of those gets updated, the agent’s behaviour can move without anybody touching the agent.
  • ▸The deployment info. The cloud project and region it went to, the environment, who ran the deploy, and when.

Figure 13 shows the dashboard that we have built to keep track of all of these components in a single unified view that allows us to easily detect and study any DNA change that happens in

Keeping track of the full agentic DNA.

3Access policy

What is the agent allowed to do, and what’s off-limits?

Today we enforce a simple policy that restricts each agent to a defined set of tools. However, the access-policy approach we’re working towards places a central policy layer between agents and the resources they want to use, so organisation-wide rules can be applied consistently rather than implemented separately by every agent team.

Each agent’s policy follows a policy-as-code paradigm in a central repository (see Figure 15), separate from the agent’s own codebase and versioned as it evolves. Before an agent reads data, calls a tool, or takes an action, a checkpoint inside the agent sends the request to this central rulebook, which returns a verdict of allow, deny, or escalate to a human.

Rules can depend on context such as the client, market, user, or budget: a routine audience request might be allowed, a health-related targeting might be escalated, and one targeting children denied. Moreover, launching a campaign can carry explicit authorisation rules, such as “Deny updating budget campaign that exceeds a 1000 GBP budget”.

We are currently experimenting with Open Policy Agent (OPA), an open-source, general-purpose policy engine. Figure 14 shows an example of a request that is rejected via the OPA engine.

 

Example of an order denied by OPA.
Police-as Code in the OPA engine

4Traffic control

How do we manage the agent’s traffic and resource use?

Suppose that this month’s spend report comes in at double of last month’s.

How could this happen? Which agent is responsible? What did it do that caused this spike? Did it send too many requests to other agents, to its own LLM brain, or to some tool out on the internet? Were those requests justified? And how do we stop it happening again?

To answer questions like that, the first thing to understand is that there are three roads in this multi-agent system that are purpose built to address the above concerns.

The first road is for model calls, an agent talking to its own LLM brain to make a decision. No agent in the community holds its own key to OpenAI or Google or anyone else. In our community, they all ask one service, the LiteLLM proxy, and it makes the call on their behalf. One doorway, so every model call is seen, priced and attributed to somebody.

The second road is for agents talking to each other. Agents do not call each other directly. Every one of those requests goes through an A2A gateway that we have built in WPP Research. Think of it as the reception desk of the agent community. Nothing reaches an agent without walking past it.

The third road is for tool calls, an agent reaching out to an MCP server to do something it cannot do itself. Traffic control for this road is something we are actively working on.

Let’s focus a bit more on the second road, the agent-to-agent reception desk ( A2A gateway). Every request goes through the same 8 steps.

▸One – Route: The request has the recipient agent’s name on it. The desk looks that name up in its register of who exists, and a name it does not recognise is turned away immediately.

▸Two – Authenticate. Every agent has its own key. The caller agent has to present the right key to get the next step.

▸Three – Throttle. Is this caller making too many requests? A caller stuck in a loop gets throttled here, rather than finding out later when the bill arrives.

▸Four – Budget. Every agent has a monthly cap. The desk adds up what this one has spent so far and compares. If this new call would take it over the line, it gets rejected.

▸Five – Hand over. Only now does the request reach the recipient agent. Credentials are swapped on the way through, so the caller’s private key doesn’t get exposed.

▸Six – Wait . The desk holds the line open and waits.

▸Seven – Receive.  The reply comes back and the reception desk sends it to the caller agent.

▸Eight – Log. A new log line is appended for every call. When, who called, which agent answered, how it ended, how long it took, what it cost.

Figures 16 and 17 illustrates the traffic-monitoring dashboard that we have have built for both LLM (Road 1) and A2A (Road 2) traffic. We are currently in the process of extending this for tool-based traffic (Road 3).

The A2A Gateway dashboard: call volume, routing and latency.
The A2A Gateway dashboard: call volume, and per-agent budgets, with a LiteLLM traffic cap in effect.

5Event-level telemetry

Can we trace every action and event in the agent’s lifetime?

We use Langfuse to log every interaction an agent has, together with latency, cost, inputs and outputs, and per-agent metadata. The payoff shows up most clearly when something breaks.

While preparing this community, we hit a bug where the Zeitgeist agent kept timing out (as seen in Figure 18). In Langfuse we could see the Campaign Designer calling the Zeitgeist agent and receiving an error, then follow the trace on the Zeitgeist side step by step: the marked error, the commit hash in its metadata pointing to the exact deployed version, and the stack trace naming the specific Python file and line. Following that to the code, we found a timeout set to an overly optimistic value. Because telemetry took us straight from the symptom to the offending line, we detected, diagnosed, and fixed it quickly.

Langfuse traces the timeout from symptom to root cause.

6Continuous verification

Is the agent doing its job correctly, safely, and efficiently?

We verify agents with a simulation-based tool called VerifyAx, which lets us onboard an agent, define testing scenarios, and get back metrics and reports on its performance. Two of the community’s agents are onboarded so far, with the rest in progress. In the Geo agent’s reports, for example, we can see both strong runs and cases where the agent slipped, such as falling for traps or drifting from expected behaviour, that Verify caught.

VerifyAx can test an agent before deployment, whenever it changes, and on a schedule once it is live, to catch regressions. Full continuous verification, though, needs more than simulation. We are actively working on monitoring agents during live, real-world interactions and detecting failure modes there. Google’s Model Armor is one of several tools we are evaluating for this.

VerifyAx reports for the Geo agent: a clean run alongside one where the agent slipped, and Verify caught it.

7Usage and cost attribution

What is the agent costing, and to whom?

The A2A Gateway and Langfuse together give us rich cost data at the event and agent level (see Lanfuse dashboards in Figures 20 and 21). That raw data is necessary but not yet sufficient: to fully answer who spent what, and on whose behalf, we still need to roll costs up along business-meaningful dimensions such as user, team, client, and organisation, with dashboards tailored to different roles. We are building those roll-ups, and plan to explore Google’s cost and usage dashboard templates as part of the work.

Cost data available today from the A2A Gateway, ahead of the business-level roll-ups we are building.
Cost data available today from Langfuse, ahead of the business-level roll-ups we are building.

What’s next

This community shows how fundamental governance principles can be implemented in practice. It is a big step forward, but the job is not done. We are actively extending the implementation across all seven principles. Our immediate priorities are to:

  • ▸Extend the policy engine we are developing for Principle 3
  • ▸Extend the community with non-GCP based agents, built on different SDKs and running on different runtimes, while ensuring that all principles still apply.
  • ▸Experiment with more advanced traffic-control protocols for Principle 4, focused on cutting LLM cost without sacrificing outcomes;
  • ▸Add multiple expert agents that cover the same tasks with different strengths and weaknesses, so we can study richer mechanisms for collaboration and reputation.

For the broader set of open problems we plan to tackle, see our research agenda for expert agent communities.

Authors

  • Ted co-leads WPP Research and serves as Head of Data Science at Satalia. He is an Assistant Professor in the Department of Marketing and Communication at the Athens University of Economics and Business. His research spans scalable algorithms for multimodal data, synthetic data generation, simulation-based verification for AI agents, and information diffusion and collective intelligence in expert networks.

  • Jael is a Data Scientist at Satalia, leveraging her physics background for a deep analytical foundation in complex systems analysis and modelling. Her experience, spanning foundational research in computational physics and a proven track record in data science consultancy, provides a unique perspective for architecting robust, scalable models in intricate environments. In Satalia’s Research Lab, she bridges scientific methodology with industrial innovation to address WPP’s most sophisticated data challenges. Her current research focuses on multimodal fusion models, aiming to improve campaign performance and pioneer state-of-the-art machine learning.

  • Eirini is a Data Scientist at Satalia with a multidisciplinary background in Management Science and Computer Science. She specialises in architecting end-to-end data science solutions, leveraging a deep technical toolkit to solve complex industrial challenges across diverse sectors. Known for bridging the gap between theoretical research and scalable application, she focuses on delivering high-impact models that translate abstract data patterns into actionable strategic intelligence.

    Her current research focuses on sophisticated campaign performance multimodal modelling and the development of data enrichment frameworks to maximise predictive accuracy.

  • Elektra Papazoglou is a machine learning engineer at Satalia, working in the Research Lab on NLP and LLM systems. Her previous work has focused on large-scale information extraction from unstructured data, with an emphasis on combining prompt engineering and efficient fine-tuning to build methods that are both performant and practical.
    She brings prior experience building production ML systems across recommendation, experimentation, and content understanding, bridging the gap between research and deployment.

More posts