Agents are being built across WPP at remarkable pace and breadth. In every discipline, on every platform, by both technical and domain experts. This has already revolutionised the way we work. It has also surfaced a hard governance challenge that lurks underneath the excitement, and grows sharper as our agentic footprint expands. In a previous post, we set out the seven principles for agentic governance that every agent at WPP should satisfy. In this post, we show how these principles can work in practice, by applying them to diverse multi-agent community focused on media campaigns.
Author: Eirini Kolimatsi
-
Are open-source embedding models good enough? A comparative study
Throughout the years, a very large number of embedding models have emerged, each one having different strengths and weaknesses. We set out to answer a practical question many data professionals face often:
Do premium embedding models significantly outperform free, open-source alternatives when predicting social media success?

Figure 1 In our latest research, we examined the performance of various embedding models across two distinct tasks. Our findings suggest that whilst premium models from industry leaders like OpenAI and Google technically produce the optimal results for most cases, the margin of victory is surprisingly narrow. Our main takeaway? For many predictive marketing tasks and Retrieval-Augmented Generation (RAG) tasks, free open alternatives are often good enough.
The methodology and model selection
To test the capabilities of different embeddings, we designed two evaluation tasks:
- Predictive downstream modelling: We generated embeddings from our data sources and fed them into a downstream LightGBM model to predict the final target (such as post popularity).
- Information retrieval: We compared the embeddings inside a RAG framework using a vector index to measure direct context retrieval performance.
We tested seven different embedding models to get a comprehensive view of the landscape:
Model Provider Base Architecture Category text-embedding-005Google Proprietary Premium (Paid API) text-embedding-3-largeOpenAI Proprietary Premium (Paid API) thenlper/gte-baseAlibaba BERT Open-Source BAAI/bge-base-en-v1.5BAAI BERT Open-Source all-mpnet-base-v2SBERT MPNet (Microsoft) Open-Source all-roberta-large-v1SBERT RoBERTa (Meta AI) Open-Source all-miniLM-L6-v2SBERT MiniLM (Microsoft) Open-Source Embedding models These open-source alternatives were primarily chosen to correlate with the embedding dimensions of the premium baselines, while specifically including a smaller (
all-miniLM-L6-v2) and a larger (all-roberta-large-v1) alternative to see how model size impacted performance.
Experiment 1: Predicting Instagram post popularity
For our first experiment, we used the public Instagram Influencer Dataset, which consists of various posts from different online influencers, to predict post popularity.
Since the raw data is highly visual, we first used Gemini to generate detailed textual descriptions of each post. We then converted these descriptions into embeddings using our seven chosen models, which were then used to train our downstream LightGBM model.
The results: OpenAI’s
text-embedding-3-largeproduced the best overall results with an average R² score of 0.475, closely followed by Google’stext-embedding-005at 0.470. However, the smallest model we tested (all-miniLM-L6-v2) still achieved a score of 0.440. This competitive showing from the open-source models is particularly impressive when you consider the potential “family alignment” advantage in the workflow, where descriptions generated by Gemini might naturally favour Google’s own embedding model. Despite premium models taking the lead, the R² performance gap between the best and worst models was a mere 0.035.
Experiment 2: The SMP challenge image dataset
To validate our initial findings, we applied the exact same framework to the image dataset from the Social Media Prediction (SMP) Challenge, which was instead used to predict popularity of Flickr posts.
The results: Once again, the premium models topped the charts, with both OpenAI and Google producing an identical R² score of 0.234. Just like our Instagram experiment, the weakest free alternative trailed by that same narrow margin, coming in at 0.196.
Experiment 3: Retrieval-augmented generation (RAG) performance
To see how these findings hold up beyond downstream regression tasks, we introduced a third experiment: a traditional Retrieval-Augmented Generation (RAG) evaluation.
At its core, a RAG framework acts as an open-book exam for an LLM. Instead of relying solely on its pre-trained internal knowledge, the system first searches an external database to retrieve the most relevant documents matching a user’s prompt. It then passes these documents alongside the question to the LLM, ensuring the final generated response is accurate, contextually grounded, and factual.
Using the industry-standard BEIR (SciFact) dataset, we indexed 5,180+ scientific documents and evaluated how effectively each embedding model could retrieve the exact context needed to answer 200 distinct queries.
We measured two key retrieval metrics:
- Hit rate: The percentage of queries where the correct document was successfully retrieved in the top 5 results.
- Mean reciprocal rank (MRR): A measure of where the correct document ranked (closer to 1 is better).
The results: This experiment provided an interesting twist. Whilst OpenAI’s premium
text-embedding-3-largeachieved the highest MRR (0.741), the open-sourcegte-basemodel proved remarkably competitive, securing a strong second-place MRR of 0.729 and comfortably beating Google’s premium offering (0.692).When it came to the overall Hit Rate, the open-source alternative stood completely shoulder-to-shoulder with the premium giants. Instead of a clear winner, we saw a three-way tie at 86% between
gte-base, OpenAI, and Google.
Results overview: Scores across multiple experiments
To ensure a robust evaluation across various datasets, we measured performance using the R² score averaged over multiple random training splits, allowing us to establish a reliable variance margin (±) for each model.
In practice, this means instead of training our LightGBM model just once on a single slice of data, we shuffled and split the dataset into different training and testing sets multiple times. Doing this ensures that a model’s high score wasn’t just a fluke resulting from a “lucky” data split. The resulting variance margin (±) acts like an error bar: a tighter margin indicates the model is highly stable, telling us exactly how consistent and reliable its predictions will be when exposed to entirely new data.
Model Size Category Exp 1: Instagram (R²) Exp 2: SMP Challenge (R²) Exp 3: RAG (MRR) Exp 3: RAG (Hit Rate) text-embedding-3-large (OpenAI)Premium / Baseline 0.475 ± 0.022 0.234 ± 0.024 0.741 0.86 text-embedding-005 (Google)Premium / Baseline 0.470 ± 0.025 0.234 ± 0.023 0.692 0.86 thenlper/gte-baseBaseline Match 0.458 ± 0.018 0.212 ± 0.027 0.729 0.86 BAAI/bge-base-en-v1.5Baseline Match 0.457 ± 0.020 0.217 ± 0.022 0.688 0.84 all-mpnet-base-v2Baseline Match 0.446 ± 0.021 0.216 ± 0.027 0.614 0.74 all-roberta-large-v1Larger Alternative 0.440 ± 0.021 0.196 ± 0.020 0.594 0.71 all-miniLM-L6-v2Smaller Alternative 0.439 ± 0.010 0.197 ± 0.015 0.596 0.75 Experiment results 
Experiment results The true cost of performance: Latency and pricing
Whilst the minimal gap in predictive accuracy alone makes a compelling case for utilising open-source alternatives, factoring in latency and execution costs makes the decision even clearer.
To illustrate this, we tracked the time and money spent to generate embeddings for each experiment. For reference, the dataset was of size 11,500 rows for Experiment 1, size 10,000 rows for Experiment 2 and 5,180 documents with 200 queries for Experiment 3.
When analysing costs, it is important to separate the costs into two main categories: the cost of calling each embedding model’s API and the cost of the underlying compute resources (such as a virtual machine or a local laptop) required to process the embeddings. For our experiments, we ran everything on a standard Virtual Machine (VM). We have excluded that infrastructure cost from this breakdown, as it fluctuates wildly depending on a developer’s specific deployment preferences and scaling needs.
Model Exp 1: Token Fees Exp 1: Latency (Seconds) Exp 2: Token Fees Exp 2: Latency (Seconds) Exp 3: Token Fees Exp 3: Latency (Seconds) text-embedding-3-large (OpenAI)$1.77 75 $0.51 41 $0.22 298 text-embedding-005 (Google)$1.29 78 $0.41 50 $0.17 241 all-roberta-large-v1Free 512 Free 438 Free 3,777 BAAI/bge-base-en-v1.5Free 344 Free 247 Free 2,229 all-mpnet-base-v2Free 282 Free 240 Free 1,990 thenlper/gte-baseFree 84 Free 60 Free 11,420 all-miniLM-L6-v2Free 43 Free 30 Free 251 Latency and costs The latency vs. cost trade-off: Who wins on speed?
Whilst open-source models eliminate recurring API costs entirely, a common concern is whether self-hosting sacrifices processing speed. To find out, we timed how long each model took to generate the embeddings across our experiments.
To keep the comparison fair, we ran all the open-source models on a standard Google Cloud Platform (GCP) virtual machine: an n1-standard-8 instance (8 vCPUs, 30 GB memory) equipped with a single NVIDIA T4 GPU.
The results revealed a highly competitive landscape with massive implications for production pipelines:
1. In small-to-medium tasks, light open source dominates
For standard datasets like Experiment 1 (Instagram) and Experiment 2 (SMP Challenge), the lightweight open-source models proved that you don’t need to pay for a premium API to get blazing-fast speeds:
- The absolute champion: The tiny
all-miniLM-L6-v2absolutely shredded the competition, completing Experiment 2 in a mere 30 seconds (and Experiment 1 in 43 seconds), beating Google’s premium API by up to 40% whilst costing $0 in token fees. - Premium APIs are fast, but costly: Google’s
text-embedding-005and OpenAI’stext-embedding-3-largeblazed through Experiment 2 in 50 and 41 seconds respectively, but carrying token bills of $0.41 and $0.51. Scaled across millions of rows, those micro-transactions add up quickly.
2. At scale (Experiment 3), the infrastructure tax emerges
When we scaled up to the larger RAG retrieval dataset in Experiment 3, the dynamics shifted heavily, highlighting the core trade-off of hosting your own models:
- Premium APIs pull ahead on massive batches: Google’s model completed the entire RAG dataset in just 241 seconds (costing $0.17), whilst OpenAI finished in 298 seconds ($0.22). Because their infrastructure is globally distributed and massively parallelised, they handle large batches effortlessly.
- The self-hosting bottleneck: Whilst
all-miniLM-L6-v2stayed nimble at 251 seconds, heavier open-source models struggled on our single-GPU setup. For instance,thenlper/gte-base, our RAG accuracy champion, took over 11,400 seconds (more than 3 hours) to complete the run on our T4 GPU.
The Takeaway: When budgeting for your pipeline, remember to separate API token costs from your underlying infrastructure costs (the VM compute time). Model latency serves as an excellent proxy for your infrastructure bill as the longer a self-hosted model runs, the longer your VM has to operate.
If you are running light-to-medium real-time tasks, optimised open-source models like
all-miniLM-L6-v2give you double savings: zero API fees and lower VM uptime. But if you’re processing massive, enterprise-scale batches and don’t want to invest in scaling a heavy local cluster of GPUs, paying a premium API fee might be the more cost-effective route.
Key findings and conclusion
Our experiments across predictive marketing tasks and search retrieval pipelines point to a nuanced conclusion:
- Premium edges out open-source in predictive modelling, but only just: Both OpenAI’s
text-embedding-3-largeand Google’stext-embedding-005consistently delivered the optimal results across our data splits. However, the difference in predictive performance between the premium giants and top open-source models remains minimal, highlighting that free alternatives are highly capable. - Open-source matches premium performance in RAG hit rates: For semantic search and context retrieval, the free open-source
gte-basemodel achieved an identical 86% Hit Rate to both OpenAI and Google, demonstrating that free options can deliver matching retrieval coverage. However, OpenAI’s premium offering did secure the crown for tighter search precision via the highest Mean Reciprocal Rank (MRR) of 0.741, withgte-basefollowing closely behind at 0.729. - The speed vs. cost trade-off has shifted: Highly optimised premium APIs have closed the speed gap, actually outpacing mid-sized open-source models like
gte-base. Whilst open-source still completely dominates on cost (being 100% free), the absolute speed crown belongs specifically to hyper-lightweight open-source models likeall-miniLM-L6-v2, which outran everything at 30 seconds. - Size isn’t everything: Interestingly, large open models like
all-roberta-large-v1underperformed across the board compared to optimised, smaller ones likegte-base.
Our study focused on specific, mainstream tasks, providing strong evidence that free and open-source models can often be enough to get the job done effectively for standard applications.
However, it is important to acknowledge that frontier embedding models, as well as massive open-weight models, still hold distinct advantages. For highly demanding use cases requiring massive context windows, multilingual support, or nuanced reasoning across obscure domains, state-of-the-art models consistently prove their worth. A quick look at popular industry benchmarks, such as Hugging Face’s Massive Text Embedding Benchmark (MTEB) leaderboard, clearly demonstrates this. The frontrunners at the very top of those charts are continually pushing the boundaries of what’s possible across dozens of highly specialised datasets.
Ultimately, depending on whether your priority is:
- Squeezing out the absolute highest KPI value and conquering highly complex edge cases
- Minimising cost and latency whilst maintaining competitive results
You now have a highly capable spectrum of both premium and open-source models to choose from.
Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.
-
From guesswork to foresight: How AI is predicting the future of marketing campaigns

Ever wonder why some advertisements seem to pop up exactly when you’re thinking about buying something, while others feel completely irrelevant? Or how a brand knows just the right message to share to get your attention? The answer lies in the evolving world of marketing campaigns, and increasingly, in the powerful capabilities of Artificial Intelligence (AI).
But what is a marketing campaign?
At its core, a marketing campaign is a carefully planned series of activities designed to achieve a specific goal for a business – whether that’s selling more products, building brand awareness, or encouraging people to sign up for a service. Think of it like launching a rocket: you need to choose the right destination (your objective), design a powerful engine (your creative message), select the perfect crew (your audience), and pick the best launchpad (your platform).
The process of creating and running these campaigns involves countless decisions, such as:
- Audience: Who are we trying to reach? What are their interests, demographics, and behaviours?
- Brand: What message do we want to convey about our brand? How does our brand resonate with the audience?
- Creative: What kind of ads should we run? (text, images, videos, headlines, calls to action).
- Objective: What’s the main goal? (e.g., getting clicks, making sales, increasing brand recognition).
- Platform: Where should we run these ads? (e.g., Facebook, Instagram, Google Search, TV, billboards).
Campaign design is a complex process shaped by multiple factors, such as creative genius, market insights, and domain expertise. Marketers launch campaigns, closely track their impact, and then adjust their approach in real time, refining messages or recalibrating target audiences. However, given the diverse and dynamic nature of consumer behavior, this iterative adaptation process can be taxing in terms of both budget and time. It’s like setting a rocket’s course: unforeseen atmospheric shifts can require significant mid-flight corrections, each consuming valuable resources.
This is where the big challenge lies: how do we predict if a campaign will be successful before we invest significant time and money into it?
Machine Learning: Your marketing crystal ball
This challenge is precisely where Machine Learning (ML) steps in. Simply put, Machine Learning is a branch of AI that allows computers to “learn” from data without being explicitly programmed. Instead of following a strict set of rules, ML algorithms analyze vast amounts of past information, identify hidden patterns and relationships, and then use those learnings to make predictions or decisions on new, unseen data.
In the context of marketing campaigns, ML becomes an incredibly powerful tool:
- Data powerhouse: Imagine collecting every detail from thousands of past marketing campaigns: who saw them, what the ads looked like, where they were shown, how much they cost, and crucially, what the final outcome was (e.g., how many clicks, sales, or sign-ups they generated). ML algorithms can digest this colossal amount of data in seconds.
- Pattern recognition: These algorithms don’t just store data; they look for correlations. Did campaigns with a specific type of image perform better with a certain age group? Does a particular headline style lead to more conversions on one platform versus another? ML can uncover these subtle yet powerful insights that human analysts might miss.
- Predictive power: Once trained, an ML model can take the proposed details of a new campaign (e.g., its target audience, creative idea, intended platform) and predict its likely outcome. It can estimate click-through rates, conversion probabilities, or even the potential return on investment (ROI) before a single dollar is spent.
The benefits are transformative: marketers can make data-driven decisions, allocate budgets more efficiently, target the most receptive audiences with precision, and ultimately, launch campaigns with a much higher probability of success. It’s like having a detailed weather forecast for your rocket launch, helping you choose the perfect day and trajectory.
The multimodal challenge: Mixing apples, oranges, and billboards
In reality, and contrary to what many might assume, a campaign isn’t just a neat row of numbers on a spreadsheet; it’s a vibrant, messy mix of text, images, locations, and abstract concepts like brand identity. This presents a fundamental challenge: how do we empower AI to not just process, but truly understand and effectively connect these inherently different types of information to form a holistic view? For instance, how can an AI understand the interplay between the nuanced visual cues of a video Ad with the detailed socio-economic data of a specific target audience in a specific location?
The “secret sauce” is a technology called embeddings. Think of an embedding as a universal translator. It takes complex information, like the “feeling” of a brand or the intent of a sentence, and turns it into a list of numbers that an algorithm can easily digest. However, every piece of the campaign puzzle requires a different translation strategy.
Translating the campaign puzzle
To build a complete picture, we process each element through a specialized lens:
- Audience, Platform, and Objective: We convert these categories into numerical “flavours.” This allows the AI to recognise the distinct profile of, for example, an Instagram awareness campaign versus a search engine lead-generation tactic.
- Brand identity: We leverage the fact that Large Language Models (LLMs) already possess a wealth of knowledge about established brands. By feeding the AI a rich, descriptive profile of a brand, we create a deep numerical representation of its identity. This task is so nuanced that it led to the birth of our Brand Perception Atlas Pod.
- Creative (Images): A picture may be worth a thousand words, but our models currently prefer numbers. To bridge this gap, we use AI to extract a highly detailed description of each image, which is then translated into data. We quickly discovered that the quality of these descriptions depends entirely on the instructions given to the AI. This led us to develop the Self-improving AI Agent.
- Geography: Location is more than just a pin on a map. To capture the true essence of a region, we use advanced models that go beyond coordinates. In detail, Google’s PDFM (Pre-trained Deep Foundation Models) Embeddings are able to capture the social, economic, and demographic fabric of an area, providing the AI with the “soul” of a location rather than just its name.
Where does the data come from?
Real-world marketing data is essential, but on its own it is not enough for AI research. At WPP, we combine rich, real-world data with carefully engineered synthetic data to build and evaluate models more effectively. Real data grounds our work in genuine market behaviour, complexity, and business context. Synthetic data adds something equally important: control. It allows us to create the specific conditions we need to properly challenge, probe, and improve our models.
This matters because many of the scenarios that determine whether a model is truly robust are rare, emerging, or simply absent from historical records until the moment they become a real problem. To prepare for that, we deliberately generate datasets that introduce edge cases, shifting patterns, variable data volumes, heterogeneity, sparsity, and data drift. In other words, we use synthetic data to stress-test models in ways that real data alone cannot support, so they are more resilient, reliable, and ready for the real world.
To address this, we built a Synthetic Data Generator. Think of this as a high-fidelity flight simulator for marketing. Instead of testing our models only on the limited “flights” we’ve taken in the past, this tool creates realistic, artificial campaign data. This allows us to:
- Train with precision: We can create scenarios that haven’t happened yet to see how the AI reacts.
- Test the limits: We can stress-test our models against extreme market conditions without any real-world risk.
- Ensure safety: We can evaluate performance using high-quality data that carries none of the privacy concerns of personal information.
- Hold the answer key: Because we generate this artificial data from scratch, we already know the exact outcome (the “ground truth”) of every scenario. It’s like giving our AI a test where we already hold the perfect answer key, allowing us to verify its predictions and recommendations.
By “conjuring” this artificial data, we ensure our models are battle-tested and ready for the complexities of the live market.
From data to decisions: Empowering the expert
We’ve explored the “ingredients” and the “recipe,” but what does this actually look like in the hands of a marketing expert? Our goal isn’t just to crunch numbers; it’s to provide actionable recommendations that make experts more efficient and their campaigns more successful.
Imagine a strategist coming to the platform with a specific mission:
“I’m launching a campaign for Brand B, targeting Audience A in Location X, with the objective of Increasing Awareness. What is the best platform and creative style to use?“
To answer a question like this, we need more than a search engine, we need a Predictive Engine, our “crystal ball”.
Before we can offer a recommendation, we must train a Machine Learning (ML) model to understand performance. We teach it to look at millions of historical and synthetic data points to predict an outcome: Is this specific combination of elements likely to be Good, Average, or Bad?
There isn’t just one way to build this crystal ball. In our research, we explore a spectrum of algorithms, including both traditional models and modern techniques. Each approach offers its own set of advantages: some prioritise speed, while others prioritise pinpoint accuracy. By testing across this variety, we ensure that when an expert asks for a recommendation, the answer is backed by the most robust mathematical thinking available today.
1. The reliable workhorse: LightGBM
We started with a classic, high-speed approach called LightGBM. Think of this as a highly efficient logic tree. It’s fast, dependable, and excellent at spotting clear patterns in structured data. It serves as our “baseline”, the standard we aim to beat.
2. The specialist team: Neural Networks
Next, we built a more sophisticated system, based on Neural Network architectures, that works like a well-organized corporation. We divided the AI into two stages:
- Specialized departments: Each type of data (like your brand identity or your creative images) is handled by its own “mini-expert” that decides which details are actually important.
- The executive board: Once the experts have done their work, a central “manager”, which is called MLP, looks at all the reports together to make the final call: Will this campaign succeed?
In this category, we have experimented with multiple different architectures and techniques. For example, one of our best models, before making a prediction, it mathematically groups elements that “belong” together. If a specific high-energy image consistently drives high success when paired with a young, active audience, the model learns to pull those winning pieces closer. This not only makes the model smarter but also helps us give you much better recommendations for future pairings.
3. The language experts: LLMs
Finally, we tested whether a standard Large Language Model (like the ones used for chatbots) could do the job on its own. Interestingly, we found that “out-of-the-box” AI isn’t naturally great at these specific marketing predictions. However, when we provide specialized training (process called “Fine-Tuning”) their performance skyrockets, as evidenced by our research: From hype to impact: Predicting campaign performance with fine-tuned LLMs.
The verdict: Measuring impact
To evaluate our models and determine how accurately they predict campaign performance, we must first establish a rigorous testing ground. This involves two key components: the diversity of our data and the precision of our metrics.
Datasets
To ensure our findings aren’t just a “lucky” outlier, we don’t rely on a single source of information. Instead, we test every model against three different versions of our synthetic datasets. By proving that our models can perform consistently across various simulated environments, we can be confident that their predictive power is both reliable and adaptable to real-world shifts.
While each of our datasets shares a consistent structure, we have intentionally varied their internal characteristics to put our models through a rigorous stress test. By using our Synthetic Data Generator, we can precisely control three key variables to create progressively more challenging environments:
- Volume: Testing how the models perform with both limited information and vast amounts of data.
- Balance: Adjusting the “label distribution”. For example, creating datasets where “average” results are far more common than clear successes or failures, to reflect the reality of a crowded market.
- Signal strength: Tuning how obvious or subtle the patterns are, which forces the models to work harder to find the winning combinations.
This approach ensures that our models aren’t just memorizing easy patterns, but are truly learning to find value in complex, “noisy” environments where the right answer isn’t always obvious.
Model performance
When it comes to measuring performance, we use a standard industry benchmark known as the F1 score, because simple accuracy can be a liar. Imagine you have a box of 100 fruits: 10 are apples and 90 are oranges. You build a robot to grab only the apples. If the robot sits still and does nothing, it is technically “90% accurate” because it correctly ignored the 90 oranges, but it’s a total failure at its job. The F1 score exposes this by balancing two hidden grades:
- Precision (the “quality” grade): When the robot grabs a fruit and says “Apple,” is it right? High precision means it never accidentally grabs an orange.
- Recall (the “completeness” grade): Did the robot find all 10 apples, or did it leave some behind? High recall means the robot is thorough and doesn’t miss any.
The F1 score is a single number that averages these two. Unlike a normal average, it “punishes” extreme failure. If your robot is perfectly accurate but misses every single apple, its F1 score will be 0. This gives us a much more honest picture of how well a model actually works in the real world. To circle back to our case, we use the F1 scores in two ways:
- The big picture: We report the Average F1 score across the entire dataset to show overall model health.
- Performance by category: We break down results into three specific classes: Negative, Average, and Positive.
This granular view is where the true business value lies. It allows us to ensure the model excels at the extremes, identifying the “Negative” combinations a marketer should avoid at all costs, and the “Positive” combinations that will truly drive results beyond the status quo.
To bring structure to our innovation, we developed a centralized model Leaderboard. This platform serves as the definitive “source of truth” for our research team, ensuring that every breakthrough is measured against the same rigorous standards. The Leaderboard allows team members to download standardized training and testing splits for any dataset (whether real or synthetic) and submit their results for comparison. By centralizing our findings in one place, we achieve several key advantages:
- True comparability: We can be certain that we are comparing equals across different algorithms and techniques.
- Accelerated testing: It allows us to quickly and safely iterate on new ideas without reinventing the wheel.
- Institutional knowledge: It creates a permanent record of our progress, ensuring that the best-performing models are always visible and ready to be deployed.
This structured environment is what allows us to move from individual experiments to a scalable, high-efficiency engine for marketing AI.
With our datasets defined and our Leaderboard in place, we put our models to the ultimate test. By measuring how each approach handled “Negative,” “Average,” and “Positive” campaign outcomes, we can clearly see which strategies offer the most reliable path to success.
Here is a glimpse of how our top models performed across the board:
Dataset Model Overall F1 Neg F1 Avg F1 Pos F1 Small and easy Tree-based 85.11% 86.4% 80.21% 88.73% Deep Learning (v1) 86.26% 89.81% 82.51% 86.44% Deep Learning (v2) 89.47% 91.25% 85.95% 91.22% Small and slightly noisy Tree-based 85.54% 87.48% 80.92% 88.21% Deep Learning (v1) 84.80% 89.41% 80.50% 84.50% Deep Learning (v2) 88.84% 91.13% 85.26% 90.14% Big and slightly noisy Tree-based 77.40% 76.40% 81.47% 74.34% Deep Learning (v1) 80.74% 80.88% 86.10% 75.25% Deep Learning (v2) 81.29% 82.25% 84.58% 77.03% F1 score Performance of Top Models (Tree-based, Deep Learning v1, and Deep Learning v2) Across Varied Datasets Analysing the Leaderboard: Reliability at scale
The results from our testing provide a clear picture of how these models handle real-world complexity.
Our tree-based model remains a formidable workhorse, maintaining an F1 score above 85%. Most importantly, it demonstrates high accuracy in identifying “Positive” and “Negative” outcomes. This means the model is exceptionally reliable at flagging the two things marketers care about most: which campaigns are likely to be massive successes and which ones are headed for failure. While performance naturally dips as we introduce more noise and scale into the datasets, its baseline remains impressively high.
While the classic models are strong, both of our Deep Learning approaches consistently take the lead. These models perform better because of their inherent capacity for “relational intelligence”, they can spot the subtle, complex connections that simpler, logic-based systems often miss.
As the datasets grow larger and the patterns become more “noisy,” this deep understanding becomes a critical advantage. Seeing these models maintain performance above the 80% mark, even in the most challenging scenarios, gives us the confidence that our AI can handle high complexity scenarios.
The privacy puzzle: Learning without sharing
While our research shows how powerful these models can be, a significant question remains: How do we build an elite AI that learns from everyone, without exposing anyone’s private data?
In the traditional world of marketing, building a “super-brain” meant pooling all client data into one giant, central database. In today’s world, that is a massive privacy red flag. We believe you shouldn’t have to choose between competitive intelligence and data security. To solve this, we utilize a cutting-edge approach called Federated Learning (FL).
Think of Federated Learning like a team of specialized doctors working in different hospitals. To find a cure for a new disease, they don’t send their private patient files to a central office, that would be a breach of trust. Instead, each doctor studies their own patients locally, they discover what works and what doesn’t and then they share only the “recipe for the cure” with their colleagues, never the patient’s identity.
In our ecosystem, each client trains the AI model locally on their own private data. The “lessons learned” are sent back to our main server, where they are combined to create a smarter, globally-informed model for everyone.
The result? You benefit from a model that has “seen” millions of scenarios, yet your private data never leaves your hands.
Discover more about Federated Learning by reading our post: Training together, sharing nothing: The promise of Federated Learning.
Looking to the Future: The Next Frontier
Our journey doesn’t end with a successful prediction. We are already exploring the next horizon of marketing intelligence, moving from understanding the past to actively designing the future. Here is what we are building next:
1. Infusing data with “common sense”
What if our models understood human psychology as well as they understand spreadsheets? We are exploring ways to inject the broad, contextual knowledge of LLMs directly into our training data. This gives our models a “common sense” layer, allowing them to understand the subtle cultural and social nuances that drive human behavior. Learn more about it in our post: The uncharted territory: Beyond the known data
2. AI building AI
We believe the best architect for a complex model might be the AI itself. Instead of manually designing every layer of a neural network, we are using advanced systems to automatically discover the ultimate model structure for marketing predictions. This process of “digital evolution”, which we delve into further in our AlphaEvolve article, ensures our tech is always one step ahead.
3. From predictions to proactive recommendations
We are currently building a tool that doesn’t just predict success, it suggests it. Imagine entering your brand and target audience and having the AI instantly recommend the perfect visual or message. We are perfecting this using two unique methods:
- The Matchmaker space: Utilizing the “relational map” we built in our earlier phases to instantly pair your audience with the creative assets they are most likely to love.
- The “hot or cold” optimisation: We treat our AI like a high-precision compass. If you have a brand and an audience but are missing the right “creative,” the system rapidly tests thousands of variations. It plays a high-speed game of “hot or cold” until it locks onto the highest possible performance score.
By moving from educated guesswork to advanced, multimodal AI, we are finally bridging the gap between creative intuition and measurable results. The rocket is fueled, the coordinates are set, and the launch sequence has officially begun.
Ready to explore the specifics? Read our full technical deep dive into Multimodal Fusion Models for a closer look at our methodology.
Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.
-
The uncharted territory: Beyond the known data
Our prior research has confirmed a fundamental truth: Machine Learning is exceptionally good at finding patterns in existing data. By analysing thousands of past campaigns, these models identify the threads of success and can predict outcomes for similar strategies with high reliability. This is an incredibly powerful tool for optimising what we already know.
However, the real-world often presents us with a different challenge: The Unknown.
- The data gap: While digital marketing generates vast amounts of information, it is rarely “clean” or perfectly integrated. Furthermore, when launching a new product or entering an entirely new market, historical data is often scarce or non-existent.
- The novelty gap: Traditional Machine Learning excels at spotting correlations, but it can struggle with the unprecedented. What happens when a novel creative concept emerges, or a sudden social trend shifts audience behaviour overnight? Because the model hasn’t “seen” these shifts in the past, it may lack the context to predict the future.
Bridging the gap: Can AI enhance our data?
At the heart of our latest research is a fundamental question:
What happens when we merge our proprietary data with the vast, world-level knowledge of a Large Language Model (LLM)?
While “adding AI” is a popular trend, real business value isn’t a given. We set out to discover if an LLM acts as a true force multiplier that fills in missing pieces, or if it simply repeats what we already know, or worse, introduces “noise” that clouds our judgment. To find the answer, we tested two distinct strategies to enhance our historical campaign data:
- Hybrid Graph creation: We build a digital “web” that connects our internal campaign facts with the LLM’s external context. This allows us to map out relationships between brands and audiences that our internal data alone might have missed.
- Active Learning: Think of this as a focused “tutoring” session. We use AI to identify the most confusing parts of our data. By putting an LLM in the loop to address these specific gaps, the model learns exactly where it can provide the most clarity.
By testing these methodologies, we are aiming to answer a critical, industry-defining question: Is the secret to superior performance simply more LLM integration, or does the true value still reside in the expert knowledge of marketing professionals that WPP has established over the years?
The path forward: Testing the synergy
To determine if AI-driven insights translate into real-world business value, we put two distinct methodologies to the test: Hybrid Graph creation and Active Learning. Each approach ensures that the LLM isn’t just a passive observer, but an active contributor to our strategy.
To test this synergy, we utilised a specialised export from WPP’s proprietary dataset, ensuring anonymity and privacy. This data captures the full lifecycle of a campaign, including information on audience characteristics, geographical location and platform-specific delivery settings. Crucially, each entry includes a definitive label indicating the campaign’s objective and its final outcome, allowing us to measure success with high precision. For a deeper dive into the architecture and specific variables of this dataset, please refer to the Campaign Intelligence Dataset Pod.
Strategy I: Hybrid Graph creation – beyond the spreadsheet
Instead of looking at data as a simple list, we treat it as a dynamic relationship map. Imagine this network as a constellation where every individual data point, whether it’s a demographic like ‘woman,’ a platform like ‘Facebook,’ or a region like ‘Spain’, becomes a node. These nodes are interconnected by lines that represent the strength and nature of their relationships.
To visualise this concept, here is an example of how this relationship network is mapped in our Hybrid Graph:

Hybrid Graph example: each node is an attribute from the real dataset. Nodes are positively connected (green) if they perform well together, negatively connected (red) otherwise. Not connected nodes indicate average or non existing relation. By combining our internal campaign history with the LLM’s broader world knowledge, we create a “Hybrid Graph”. With this multi-layered map, we expect to be able to see connections that traditional spreadsheets ignore and the process of building it unfolds in three strategic phases:
Phase 1: Establishing the ground truth
Before we ask an AI for help, we perform a “Deep Dive” into our historical data to separate coincidence from repeatable success through two rigorous tests:
- The consistency test (purity): We look for a clear “verdict”. For example, if a specific audience and brand pairing resulted in a positive outcome 85% of the time, we have a reliable pattern. The data is giving us a clear “Yes”.
- The volume test (cardinality): Consistency only matters if it happens often. For instance, if a positive connection appears consistently, we know it’s a statistically significant trend, not just a stroke of luck.
By filtering through these lenses, we identify the bedrock of our dataset.
Phase 2: The expert second opinion
Next, we turn to the “Emerging Patterns”: combinations where the data shows a clear leaning (like a 60% success rate), but the evidence isn’t yet overwhelming. Historically, these “maybe” scenarios might have been ignored. Now, we invite the LLM to act as a Strategic Consultant.
- The power of the upvote: When the LLM’s intuition aligns with our data’s hints, we gain a new level of confidence. For example, if a location and brand pairing resulted in a positive outcome 60% of the time, we suspect there is a pattern, but we need the LLM to confirm.
- Validation through synergy: By getting a “Yes” from the AI to back up our data, we move these patterns from the “maybe” pile into our active knowledge base.
Phase 3: Illuminating the “dark spots”
Finally, we shift our focus to the areas our data couldn’t reach, the “dark spots”. These are combinations that were excluded because our data was too noisy or the scenarios were entirely new.
We identify every combination where we currently lack confidence. However, this is a vast number of combinations, making it computationally infeasible to check all of them. That’s why we sample these gaps and ask the LLM for original insights based on its understanding of global markets. To give an example, the cases we’re targeting here look like: a specific audience and location pairing that is non-existent in the dataset, or a combination that results in a positive outcome half the time and a negative outcome the other half. Because of the lack of a clear pattern in the real data, we directly ask the LLM.
In short, this allows us, through the LLM, to clarify noisy data and predict outcomes for entirely new scenarios.
The challenge of illuminating our data “dark spots” led to a major breakthrough. Our first approach, the Hybrid Graph, essentially looks at low-signal campaign data and uses LLMs to make an educated guess to fill in the gaps. But this sparked a bigger, more strategic question: Instead of just guessing what’s in the dark, how can we actively hunt for the exact pieces of missing information that will make our predictions smarter? Out of millions of possible campaign combinations, how do we pinpoint the specific scenarios that will teach our model the most? This strategy of “smart hunting” forms the foundation of our second approach: Active Learning.
Phase 4: Combining everything together
The culmination of this research is a unified Hybrid Graph. By merging our proven history, our validated suspicions and our newly discovered insights, we create a living map of intelligence.
The result is a specialised dataset that is expected to offer the best of both worlds:
- The grounding of reality: Rooted in the hard facts of our actual campaign history.
- The foresight of AI: Enhanced by the vast, contextual knowledge of the LLM.
Strategy 2: Active Learning – solving the puzzle of uncertainty
Where the Hybrid Graph fills gaps with new insights, Active Learning focuses on a different truth: data isn’t always helpful if it’s redundant. To truly advance our models, we don’t need more of what we already know; we need clarity in the “grey areas” of our knowledge.
For example, imagine our data clearly shows that “TikTok campaigns” aimed at “Gen Z” are consistently successful, while “LinkedIn campaigns” aimed at “Millennials” usually underperform. But what happens if we want to run a “TikTok campaign” for “Millennials”? The model might be completely unsure if there is no clear pattern for that specific combination. Instead of analysing thousands more Gen Z campaigns we already understand, Active Learning specifically targets this exact missing combination. By resolving this one grey area, the model learns whether the platform or the audience age is the true driver of performance.
In the world of data, this uncertainty occurs when there isn’t a strong, consistent signal: when parts of a dataset tell conflicting stories, making it difficult to separate real patterns from mere noise.
In modern marketing, the number of possible combinations between audiences, brands and locations is astronomical; blindly analysing every single one would be incredibly slow, if not impossible. Instead, we use Active Learning as a strategic guide to identify the specific “pockets” of a dataset where our current models are struggling the most. It sifts through the records and picks only the most confusing, yet valuable, points for evaluation.
By focusing our efforts strictly on the areas where the model is most uncertain, we achieve two major goals:
- Maximised intelligence: We gain the most knowledge from the fewest possible data points.
- Operational speed: We bypass the “noise” of what we already know, allowing us to build high-performing models in a fraction of the time.
Ultimately, this approach turns a daunting, “infinite” dataset into a manageable, high-impact asset.
The LLM as our “oracle”
Identifying the most uncertain points in our data is only half the battle; the real value lies in what we do with them. Once we have selected these high-priority “grey areas,” we bring in the LLM to act as our oracle.
Using sophisticated prompting techniques, we present these uncertain points to the AI for a professional verdict. Our goal is to transform these pockets of doubt into certainty, backed by high-quality expert information.
By doing this, we effectively bridge the “information gap”. We aren’t just adding more data for the sake of volume; we are harvesting targeted knowledge. This process turns a previously unknown variable into a strategic asset, ensuring that our final model isn’t just a reflection of what we’ve seen before, but a fusion of our experience and the AI’s broader market expertise.
Two paths to higher intelligence
To find the most efficient way to “teach” our models, we experimented with multiple different strategies for choosing which questions to ask our LLM oracle. Below, we outline our core foundational technique and the more advanced method that has proven to be our most effective to date.
Approach 1: The broad search
This is a high-level “scouting” mission. We create a large pool of random potential campaign scenarios and ask our current model to predict how they would perform. We then identify the scenarios where the model is the most confused, the “shaky” predictions, and send those directly to the LLM oracle for a definitive answer. It’s a fast, effective way to shore up general weaknesses in our knowledge.
Approach 2: The targeted stress test (our top performer)
Our most successful approach is much more surgical. Instead of looking at random scenarios, we actively look for the “tipping points”, the exact moment a campaign shifts from being a success to a failure, or vice versa.
- Finding the edge: We take a known successful campaign and a known failure, then subtly blend their features to create a new, “borderline” scenario.
- Measuring confusion: We keep adjusting the features until a pre-trained auxiliary model (in this case, a tree-based one) flips its prediction. We then rank and select the scenarios where the outcome is most uncertain, ensuring we capture the most informative data points for our oracle to review.
- The expert verdict: We present these precise “tipping points” to the LLM oracle. By giving the AI specific examples of similar successes and failures as context, we get an incredibly high-quality label.
- Iterative learning: Once the LLM provides the answers for these “grey areas,” we integrate them into our official records. We then retrain our auxiliary model on this newly enriched dataset, making it instantly more precise. From there, the process begins again, creating a continuous loop that proactively hunts for and eliminates our model’s remaining blind spots.

The Active Learning Loop with the four main phases: training existing data, finding the tipping points, labelling them using an LLM, and finally adding them to existing data to start the loop over. By repeating this process, we don’t just add data; we specifically “fix” the model’s most significant blind spots. This iterative loop ensures that our final engine isn’t just bigger but it’s also significantly smarter.
Results
The balancing act: Extracting the final datasets
Building a Hybrid Graph is a delicate exercise in calibration. Our challenge was to find the perfect equilibrium: How much should we trust our internal data and how much “weight” should we give to the LLM’s external knowledge?
To test this, we generated several different graph versions, eventually selecting the largest and most robust one. This ensured our Synthetic Data Generator had a dense enough “knowledge web” to create high-quality, non-random datasets. To keep our findings clear, we kept environmental “noise” to a minimum, ensuring we were testing the core intelligence of the graph itself.
Similarly, when building the datasets to test our Active Learning strategies, we had to find the right blend of human experience and AI insight. After testing multiple configurations, we discovered our “Golden Ratio” was in the region of 80% Real-World Data and 20% LLM Knowledge. This 80/20 balance proved to be our most effective setting. It ensures the model remains firmly grounded in the proven reality of WPP’s historical success, while still allowing enough “AI intuition” to fill in the gaps and explore new strategic frontiers.
The reality check: Lessons from the data
To evaluate the results, we ran a “head-to-head” test. We trained one model using only real-world data and another one using our LLM-enhanced hybrid dataset. We then tested both against a “holdout” set of real campaign results.
Here are the results of our models, trained on the real dataset and tested against the holdout:
Model Overall F1 Neg F1 Avg F1 Pos F1 Tree-based 60% 54% 72% 55% Deep Learning 67% 60% 79% 63% Models’ Baseline Performance (trained and tested on real-world data) Building on the foundations of our previous research (From guesswork to foresight: How AI is predicting the future of marketing campaigns), we transitioned our models from a controlled synthetic environment to the complexities of 100% real-world campaign data.
Our standard models, which previously proved their strength in synthetic testing, delivered a highly competitive baseline. This “Reality Benchmark” set a high bar, while simultaneously identifying clear opportunities for our LLM-based techniques to add value.
The results revealed a clear trend: while the models excelled at identifying “Average” campaigns, they struggled to pinpoint the extreme “Positive” or “Negative” outliers. This is a common phenomenon in real-world marketing. Unlike our controlled synthetic environments, where we can perfectly balance the ratios, real-world data is heavily weighted toward “average” outcomes. Exceptional successes and disasters are rare, making them significantly harder for a model to learn and predict.
Within this context, Deep Learning (v2) emerged as our strongest baseline, achieving a solid 67% overall F1 score. The Tree-Based Approach performed slightly below the Deep Learning architecture, reinforcing our decision to move toward more “relational” neural networks to navigate the noise and imbalance of complex marketing datasets.
By establishing this 67% mark as our “Line in the Sand,” we can clearly measure the true impact of our Hybrid Graph and Active Learning interventions. Here is how our LLM-enhanced methodologies performed:
Method Variant Model Overall F1 Neg F1 Avg F1 Pos F1 Hybrid
Graph13k rows 60% density Tree-based 37% 51% 10% 50% Deep Learning 17% 11% 2% 37% 90k rows 60% density Tree-based 42% 57% 10% 60% Deep Learning 24% 14% 8% 35% Active Learning Broad Point Search
Real 80%, LLM 20%Deep Learning 66% 59% 77% 62% Targeted Point Search
Real 78%, LLM 21%Deep Learning 68% 60% 80% 63% Hybrid Graph & Active Learning Performance 
Performance comparison of the different models and experiment by metrics. The findings in the table above were unexpected, but deeply insightful: We noticed a significant drop in performance when the LLM was added to the loop for hybrid graph and almost no increase with active learning.
The hybrid graph challenge: A significant divergence
The most striking finding was the performance of the Hybrid Graph. Despite increasing the data volume to 95k rows, the scores dropped significantly, bottoming out at 17% to 42%.
This drop reveals a fundamental truth: generic LLMs are trained on public domain knowledge. They lack the specialised, proprietary marketing intelligence that WPP possesses. By weaving general AI “intuition” into a specialised graph, we introduced noise that actively diluted the high-quality signals of our real-world data.
Even with a denser graph, our models struggled to maintain a consistent F1 score. This proves that marketing success hinges on the niche, proprietary data unique to our field, information that simply isn’t available in the public sphere, rather than just a larger volume of generic data.
Active Learning: Reaching the efficiency frontier
In contrast, our Active Learning strategies, specifically the Targeted Point Search, successfully met the benchmark. Using our “Golden Ratio” (78% Real / 21% LLM), the Targeted Point Search achieved a 68% F1 score, slightly outperforming our best real-world baseline.
While our Targeted Point Search allowed us to maintain performance levels comparable to our 100% real-world baseline, we have to be honest: we expected a more significant leap. To justify a process of this complexity, the “performance lift” needs to be undeniable. This brings us to two critical, strategic questions:
- The quality risk: For such a marginal improvement in accuracy, is it worth introducing external AI “intuition” into our proprietary ecosystem when we cannot be 100% certain of its quality?
- The computational cost: Does the slight increase in predictive power justify the high computational expense and the mathematical difficulty of hunting for these “tipping points”?
In its current state, the answer is a cautious “No”. While the technology is fascinating, the results prove that our internal, high-fidelity data is already doing the heavy lifting. Introducing expensive, public-model “noise” for a 1% gain doesn’t just challenge our efficiency, it risks diluting the “Gold Standard” intelligence that WPP already possesses.
The strategic conclusion: Expert-led AI
Our research serves as a powerful reminder that AI is a force multiplier, not a replacement. The performance drop we saw with the Hybrid Graph Dataset underlines the immense competitive advantage of WPP’s proprietary data; generic models simply cannot replicate the “niche” intelligence we already possess.
While Active Learning was able to match our 67% baseline, “matching” the status quo is not enough to justify the hype or the computational cost.
The core insight: Data quality is the ultimate moat
This research proves a fundamental truth: Data quality is everything. A generic AI cannot replace the deep, specialised expertise of a marketing professional. The failure of the “public” LLM to improve our results demonstrates that the real path to success lies in keeping our experts in the loop. By using high-fidelity, professional strategy rather than general internet trends, we ensure our models are learning from the best in the business.
Moving forward: From baseline to breakthrough
To bridge the gap between “adequate” and “exceptional,” we have identified two clear technical paths to evolve this research:
- The fine-tuned oracle: Our current experiments used “off-the-shelf” LLMs. To truly elevate the results, the next logical step is to use fine-tuned models: AIs that have been specifically trained on WPP’s historical successes and internal playbooks. This transforms the oracle from a generalist into a marketing specialist.
- Real-world Active Learning: The ultimate validation of Active Learning isn’t a digital oracle; it’s the market itself. A Real-Time Loop can be implemented using Active Learning, to identify high-potential “blind spots,” launching those as live test campaigns and then feeding that real-world performance back into our models. This moves us from theoretical testing to real-world evolution.
Ready to explore the specifics? Read our full technical deep dive into Data Enrichment Pod for a closer look at our methodology.
Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.

