Abstract
Every marketing campaign leaves a trace.
Some of that trace is easy to see: impressions, clicks, spend, platform, objective, audience, geography. It appears in dashboards, weekly reports, post-campaign reviews and performance summaries. It tells teams what happened.
But campaign data can also contain another kind of signal. It lives in the relationships between decisions: the way a brand, audience, objective, platform and geography worked together; the way one setup made sense for one brand but not another; the way a campaign can be technically similar to something that worked before and still not be strategically right.
That hidden signal is what planners are often looking for.
Most marketing teams are already good at reading the past. The harder question comes earlier, when a campaign is still taking shape and budget has not yet been committed:
What should we try next?
This proof of concept from the Performance AI team at Open Intelligence explores that question. It builds on earlier work in the Campaign Performance Modelling Pod, but moves the problem closer to live planning: from synthetic benchmarking into observed campaign performance data, and from prediction alone into recommendation.
The ambition is simple to say and difficult to execute: help teams move from campaign reporting to campaign recommendation.
The Shift
Table 1 summarizes the shift from retrospective campaign reporting to forward-looking campaign recommendation:
| Campaign Reporting | Campaign Recommendation |
|---|---|
| Looks backward | Looks forward |
| Explains what happened | Suggests what to try next |
| Uses dashboards and post-campaign analysis | Supports planning before budget is committed |
| Compares historical outcomes | Ranks possible campaign setups |
| Answers “what worked?” | Answers “what is worth testing?” |
The Moment Before Launch
Think about the moment before a campaign goes live.
The brief is understood. The brand has a goal. There is a target audience in mind. The team has options: different objectives, platforms, markets and ways to complete the setup. Some combinations feel familiar. Others are plausible but untested. A few may be hiding in plain sight.
This is where experience matters. A planner brings commercial judgement, client context, category knowledge and creative instinct. But even an expert cannot manually inspect every possible campaign combination. The space becomes too large too quickly.
A campaign is not one choice. It is a set of connected choices, summarized in Table 2:
| Campaign Ingredient | Planning Question |
|---|---|
| Brand | Which brand context are we planning for? |
| Audience | Who are we trying to reach? |
| Objective | What do we want them to do? |
| Platform | Where should the campaign run? |
| Geography | In which market or region? |
Each answer changes the meaning of the others. An audience that performs well for one objective may be ordinary for another. A platform that works for one brand may be less effective for a different brand with a different history. A geography can look promising until it is paired with the wrong activation context.
This is why marketing performance is hard to compress into simple rules.
“This platform works” is rarely enough.
The better question is:
“This platform, for this brand, with this audience, in this context: does that combination look strong?”
That is a combination problem. And combination problems are exactly where machine learning can help, if designed carefully.
From Looking Back To Looking Forward
Campaign data has long played an important role in helping teams understand performance, learn from results, and refine their approach.
There is also an opportunity to extend that value further. Some of the most important decisions happen before there is a result to measure, which means bringing data into the process earlier can help teams make more informed choices about where and how to invest.
This proof of concept turns historical campaign data into something more forward-looking: not a replacement for human judgement, and not an automatic campaign machine, but a planning companion that can answer two practical questions.
The two tasks addressed by this work are summarized in Table 3:
| Task | Question |
|---|---|
| Prediction | Based on historical patterns, if we run this campaign setup, what click-through-rate (CTR) should we expect? |
| Recommendation | If part of the setup is fixed, which missing pieces look most promising? |
Click-through rate (CTR) serves as our example performance metric—a practical starting point given the available data. We use it across the dataset, including campaigns with objectives other than clicks, to explore whether the approach can support prediction and recommendation. This does not mean CTR measures success for every objective: stronger click-through performance does not necessarily indicate stronger awareness, engagement or video-view outcomes.
Prediction asks the model to estimate the likely performance of a complete setup.
Recommendation asks it to help complete an incomplete setup. For example, if the brand is fixed and the team is still exploring audience, objective, platform and geography, the system can suggest combinations that appear strong for that brand.
The important phrase is:
for that brand
Why Good Is Not Universal
Campaign performance is not a single global scoreboard.
Some brands naturally operate in higher-click environments. Others may have lower typical click-through rates but still run highly effective campaigns relative to their own history.
If we judge every campaign against one universal CTR scale, we risk confusing “high overall” with “strong for this brand.” That distinction matters for recommendation.
Table 4 contrasts a global view of campaign performance with the brand-relative approach used in this work:
| Global View | Brand-Relative View |
|---|---|
| Rewards campaigns with high raw CTR | Rewards campaigns that are strong for that brand |
| Can favor brands with naturally higher CTRs | Accounts for each brand’s historical baseline |
| Asks “what looked good anywhere?” | Asks “what looks good here?” |
So this proof of concept uses brand-relative CTR labels. Each campaign’s CTR is compared with the brand’s own historical CTR distribution.
The resulting brand-relative labels are defined in Table 5:
| Label | Meaning |
|---|---|
| Positive | Strong for that brand |
| Neutral | Typical for that brand |
| Negative | Weak for that brand |
This turns recommendation into a more practical planning task: not simply identifying what performed well somewhere, but identifying what looks promising in the specific context being planned.
Real Data, Real Messiness
Earlier research in this area used synthetic data to test whether multimodal models could learn campaign-performance patterns under controlled conditions. Synthetic data is powerful because it gives researchers a clean answer key. You can design the difficulty of the task, control the signal and stress-test models in repeatable ways.
This new proof of concept takes a different step. It uses observed campaign performance data.
The team started with more than 21 million daily campaign rows, covering the period from January 2023 to July 2026 across 75 anonymized brands, more than 1,000 audience representations, 7 objectives, and 42 platform representations. The data included 167 geography representations, with the United Kingdom present in each campaign row, sometimes alongside additional locations. After aggregation, filtering, and preparation, the data contained 20,067 distinct campaign setups.
That sounds tidy, but real campaign data is rarely tidy in practice.
Some combinations appear often, others barely appear at all. Some brands have lots of history, others have less. CTR is highly skewed, with many low values and a long tail of higher-performing cases.
This is part of the value of working with real data. It forces the model to face the same imbalance and sparsity that planning teams face. There is no perfect laboratory. There is only the work of extracting useful structure from imperfect evidence.
Brand identifiers were anonymized before modelling, so the system could learn brand-relative patterns without exposing identifiable brand names inside the modelling workflow.
Teaching The Model To Read A Campaign Setup
The model does not treat a campaign as a flat spreadsheet row. It learns from different kinds of campaign information and brings them together into a shared representation of the full setup.
A brand is different from an audience. An audience is different from a platform. A platform is different from a geography. These are not interchangeable fields, so the model gives each campaign ingredient its own representation before combining them.
The five model inputs and their roles are summarized in Table 6:
| Input | What The Model Learns |
|---|---|
| Brand | Brand-specific performance context |
| Audience | Audience behaviour patterns |
| Objective | Intent and optimization goal |
| Platform | Channel-specific behaviour |
| Geography | Market or regional context |
Together, these inputs help the model learn whether a campaign setup appears coherent, whether it resembles combinations that worked well for the brand, and whether it deserves attention during planning.
The same learned representation supports both prediction and recommendation, as shown in Table 7:
| Use Case | Output |
|---|---|
| Prediction | Estimated CTR for a complete setup |
| Recommendation | Ranked campaign completions for an incomplete setup |
The model is therefore not just producing a score. It is learning a map of campaign compatibility.
Prediction: Estimating The Likely Outcome
The first task was to test whether the model could predict CTR for campaign combinations it had not simply memorized.
For most readers, the exact definitions of the prediction metrics are less important than the direction of the result: the model learned meaningful signal from the campaign ingredients and their relationships.
It was not perfect. The model remained conservative in some high-CTR slices, especially for click-focused objectives. That matters because a production planning tool should not be overconfident when the performance distribution has a long tail. Calibration and monitoring would be needed before this kind of model became planner-facing.
Still, prediction is only half the story. The more interesting question is what happens when the model is asked not just to score a campaign, but to suggest one.
Recommendation: Finding Stronger Completions
A recommendation system starts with an incomplete thought.
A planner might know the brand, but still be deciding which audience, platform, objective or geography to prioritize. Or they might know the brand and objective, but still be choosing between possible activation contexts.
The model’s job is to complete the thought in a useful way.
Not randomly. Not by simply replaying a dashboard. And not by assuming that the highest global CTR is always the best answer.
Instead, the system asks which possible completions look strong for the brand and coherent inside the learned campaign-performance space.
This proof of concept tested the two recommendation settings summarized in Table 8:
| Setting | Meaning | Why It Matters |
|---|---|---|
| Closed vocabulary | Recommends from campaign setups already present in the historical dataset | Easier to evaluate because outcomes are known |
| Open vocabulary | Proposes mostly new combinations that have not appeared historically | Closer to the real planning opportunity |
Together, the two settings answer the different evaluation questions shown in Table 9:
| Closed Vocabulary | Open Vocabulary |
|---|---|
| Can the model rank known options better than random? | Can the model suggest new combinations that still look plausible and promising? |
Closed Vocabulary: Can It Rank Known Options?
Closed vocabulary recommendation is the cleanest test of ranking quality.
If the model recommends from historical rows, we can check whether its top suggestions were actually positive, neutral or negative for that brand. We can also compare it with random selection from the same candidate pool.
Across 57 valid brand-fixed queries, the model performed substantially better than random selection, as shown in Table 10:
| Metric | Model | Random from same candidate pools |
|---|---|---|
| Positive precision@10 | 76.2% | 50.2% |
| Non-negative precision@10 | 97.1% | 82.0% |
| Negative rate@10 | 2.9% | 18.0% |
| Known@10 | 100.0% | n/a |
In plain terms, the model surfaced more historically strong options and avoided many more historically weak ones.
Table 11 presents how performance changed according to how much of the campaign setup was already fixed:
| Query Type | Positive@10 | Positive Lift vs Random | Negative Reduction vs Random |
|---|---|---|---|
| Brand only | 93.0% | 3.085x | 28.3pp |
| Brand + 1 modality | 68.7% | 1.683x | 16.0pp |
| Brand + 2 modalities | 84.3% | 1.939x | 10.4pp |
| Brand + 3 modalities | 65.6% | 1.276x | 6.5pp |
The strongest result came when the planner had fixed only the brand. In that setting, the model had space to explore many possible completions, and it found much stronger options than random selection.
As more parts of the campaign were fixed, the search space became narrower. By the time brand plus three other ingredients were already specified, there was often only one meaningful decision left. In that case, the candidate pool itself was already favourable, so the model had less room to prove that it was adding ranking value.
Product lesson:
The recommender appears most valuable when it is used early enough in planning, while there is still meaningful room to shape the campaign.
Open Vocabulary: Can It Suggest Something New?
Closed vocabulary recommendation is measurable. Open vocabulary recommendation is more ambitious.
In open vocabulary mode, the system can propose combinations that mostly do not exist in the historical dataset. This is closer to how planners actually think. They are not always asking for the best old campaign. They are often asking for a new plan that is informed by old evidence.
The challenge is evaluation. If a recommended combination has never been run before, there is no historical label to check. We cannot say, from offline data alone, whether it definitely would have performed well.
The team therefore used the diagnostic signals summarized in Table 12:
| Diagnostic Question | What It Checks |
|---|---|
| Is the recommendation close to positive examples? | Whether it sits near strong historical patterns in the learned space |
| Is predicted CTR high for the brand? | Whether the setup looks strong relative to the brand’s own history |
| Is the setup already known? | Whether the model is copying old rows or proposing something new |
Overall, the open-vocabulary recommendations showed encouraging evidence, as reported in Table 13:
| Diagnostic | Result |
|---|---|
| Closest-to-positive distribution rate | 72.5% |
| Closest-to-non-negative distribution rate | 90.0% |
| Closest-to-negative distribution rate | 10.0% |
| Average predicted CTR percentile vs brand history | 87.2% |
| Known@10 | 5.0% |
The low Known@10 is important. It means the recommender was usually not just copying old rows. Most suggestions were genuinely new relative to the historical dataset.
The strongest open-vocabulary evidence appeared for broader queries: brand-only, brand plus one ingredient, and brand plus two ingredients. These are the situations where the model has enough freedom to search the space and identify combinations that look both coherent and high-potential.
When the query became too constrained, recommendation quality became less distinctive. Again, this is a planning lesson: the system is most useful when it can participate in shaping the campaign, not merely fill in the final blank.
A New Kind Of Planning Conversation
The most interesting version of this tool is not a black box that says “do this.”
It is a system that changes the planning conversation.
Today, a team might ask:
“How did campaigns like this perform before?”
With this kind of model, they can start asking:
New Planning Questions
What would make this setup stronger?
Which options should we avoid?
Is this recommendation strong for this brand, or just strong in general?
Are we exploring genuinely new combinations, or repeating what we already know?
If we fix the objective, what changes in the best platform and geography choices?
Those are different questions. They move data from the end of the process into the beginning. They make historical evidence active at the moment decisions are being formed.
The Human Role Becomes Even More Critical
It is tempting to describe systems like this as automation. But the better description is augmentation.
The model can search across many combinations quickly. It can learn patterns that are difficult to see in a dashboard. It can help prioritize options and identify setups that deserve a closer look.
But it does not know the whole brief. It does not understand every client constraint. It does not know the timing constraints, market nuance, creative ambition, brand safety, media availability or whether a recommendation is commercially appropriate in a specific moment.
That is why the planner remains central.
The model can widen the field of view. The human still decides what is meaningful.
In practice, the best version of this system would not simply output a single answer. It would support exploration. A planner could change the fixed ingredients, compare recommendation contexts, inspect confidence signals and decide whether a suggested combination is worth taking into expert review or live testing.
The purpose is not to remove judgement. It is to give judgement a sharper instrument.
What Makes This Different From The Earlier Work?
The earlier Campaign Performance Modelling Pod work asked whether multimodal AI could learn campaign-performance patterns in controlled synthetic environments. That was an important step because it let the team test models under known conditions.
This proof of concept asks a different question:
Can similar ideas work on observed campaign performance data, where the labels are messier, the distribution is uneven, and the recommendation problem is closer to real planning?
It also extends the ambition from prediction to recommendation. Instead of asking only “will this work?”, the system begins to ask “what should we try?”
That is a meaningful shift. Prediction is useful when the planner already has a complete idea. Recommendation is useful when the idea is still being formed.
What Comes Next
This is still a proof of concept. Several things would need to happen before a system like this could move beyond the current form.
The main steps required to move toward are outlined in Table 14:
| Next Step | Why It Matters |
|---|---|
| Expert review | Check whether recommendations make strategic sense |
| Controlled live testing | Validate performance beyond offline diagnostics |
| Better calibration | Improve reliability in high-CTR slices |
| Confidence thresholds | Help planners distinguish strong suggestions from weaker ones |
| Recommendation explanations | Show which campaign ingredients are driving the suggestion |
That last point matters. A recommendation without an explanation is hard to trust. A useful planning tool should not only say:
“This looks promising.”
It should help the planner understand the evidence behind the suggestion.
The direction, however, is clear.
Campaign data should not only describe the past. It should help teams design the future.
The more useful planning question is no longer just:
“What happened?”
It is:
“Given what we know, what is worth trying next?”
That is the shift this proof of concept begins to make: from reports that explain yesterday’s performance to recommendations that help shape tomorrow’s campaigns.
Read The Technical Report
Ready to explore the specifics? Read our full technical deep dive into our Technical Report for a closer look at our methodology.
Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.