Author: Andreas Mastakouris

  • From Campaign Reports to Campaign Recommendations

    Abstract

    Every marketing campaign leaves a trace.

    Some of that trace is easy to see: impressions, clicks, spend, platform, objective, audience, geography. It appears in dashboards, weekly reports, post-campaign reviews and performance summaries. It tells teams what happened.

    But campaign data can also contain another kind of signal. It lives in the relationships between decisions: the way a brand, audience, objective, platform and geography worked together; the way one setup made sense for one brand but not another; the way a campaign can be technically similar to something that worked before and still not be strategically right.

    That hidden signal is what planners are often looking for.

    Most marketing teams are already good at reading the past. The harder question comes earlier, when a campaign is still taking shape and budget has not yet been committed:

    What should we try next?

    This proof of concept from the Performance AI team at Open Intelligence explores that question. It builds on earlier work in the Campaign Performance Modelling Pod, but moves the problem closer to live planning: from synthetic benchmarking into observed campaign performance data, and from prediction alone into recommendation.

    The ambition is simple to say and difficult to execute: help teams move from campaign reporting to campaign recommendation.

    The Shift

    Table 1 summarizes the shift from retrospective campaign reporting to forward-looking campaign recommendation:

    Campaign ReportingCampaign Recommendation
    Looks backwardLooks forward
    Explains what happenedSuggests what to try next
    Uses dashboards and post-campaign analysisSupports planning before budget is committed
    Compares historical outcomesRanks possible campaign setups
    Answers “what worked?”Answers “what is worth testing?”
    Table 1. Comparison of campaign reporting and campaign recommendation.

    The Moment Before Launch

    Think about the moment before a campaign goes live.

    The brief is understood. The brand has a goal. There is a target audience in mind. The team has options: different objectives, platforms, markets and ways to complete the setup. Some combinations feel familiar. Others are plausible but untested. A few may be hiding in plain sight.

    This is where experience matters. A planner brings commercial judgement, client context, category knowledge and creative instinct. But even an expert cannot manually inspect every possible campaign combination. The space becomes too large too quickly.

    A campaign is not one choice. It is a set of connected choices, summarized in Table 2:

    Campaign IngredientPlanning Question
    BrandWhich brand context are we planning for?
    AudienceWho are we trying to reach?
    ObjectiveWhat do we want them to do?
    PlatformWhere should the campaign run?
    GeographyIn which market or region?
    Table 2. Campaign ingredients and their associated planning questions.

    Each answer changes the meaning of the others. An audience that performs well for one objective may be ordinary for another. A platform that works for one brand may be less effective for a different brand with a different history. A geography can look promising until it is paired with the wrong activation context.

    This is why marketing performance is hard to compress into simple rules.

    “This platform works” is rarely enough.

    The better question is:

    “This platform, for this brand, with this audience, in this context: does that combination look strong?”

    That is a combination problem. And combination problems are exactly where machine learning can help, if designed carefully.

    From Looking Back To Looking Forward

    Campaign data has long played an important role in helping teams understand performance, learn from results, and refine their approach.

    There is also an opportunity to extend that value further. Some of the most important decisions happen before there is a result to measure, which means bringing data into the process earlier can help teams make more informed choices about where and how to invest.

    This proof of concept turns historical campaign data into something more forward-looking: not a replacement for human judgement, and not an automatic campaign machine, but a planning companion that can answer two practical questions.

    The two tasks addressed by this work are summarized in Table 3:

    TaskQuestion
    PredictionBased on historical patterns, if we run this campaign setup, what click-through-rate (CTR) should we expect?
    RecommendationIf part of the setup is fixed, which missing pieces look most promising?
    Table 3. Prediction and recommendation tasks addressed by the model.

    Click-through rate (CTR) serves as our example performance metric—a practical starting point given the available data. We use it across the dataset, including campaigns with objectives other than clicks, to explore whether the approach can support prediction and recommendation. This does not mean CTR measures success for every objective: stronger click-through performance does not necessarily indicate stronger awareness, engagement or video-view outcomes.

    Prediction asks the model to estimate the likely performance of a complete setup.

    Recommendation asks it to help complete an incomplete setup. For example, if the brand is fixed and the team is still exploring audience, objective, platform and geography, the system can suggest combinations that appear strong for that brand.

    The important phrase is:

    for that brand

    Why Good Is Not Universal

    Campaign performance is not a single global scoreboard.

    Some brands naturally operate in higher-click environments. Others may have lower typical click-through rates but still run highly effective campaigns relative to their own history.

    If we judge every campaign against one universal CTR scale, we risk confusing “high overall” with “strong for this brand.” That distinction matters for recommendation.

    Table 4 contrasts a global view of campaign performance with the brand-relative approach used in this work:

    Global ViewBrand-Relative View
    Rewards campaigns with high raw CTRRewards campaigns that are strong for that brand
    Can favor brands with naturally higher CTRsAccounts for each brand’s historical baseline
    Asks “what looked good anywhere?”Asks “what looks good here?”
    Table 4. Comparison of global and brand-relative performance evaluation.

    So this proof of concept uses brand-relative CTR labels. Each campaign’s CTR is compared with the brand’s own historical CTR distribution.

    The resulting brand-relative labels are defined in Table 5:

    LabelMeaning
    PositiveStrong for that brand
    NeutralTypical for that brand
    NegativeWeak for that brand
    Table 5. Definitions of the brand-relative CTR labels.

    This turns recommendation into a more practical planning task: not simply identifying what performed well somewhere, but identifying what looks promising in the specific context being planned.

    Real Data, Real Messiness

    Earlier research in this area used synthetic data to test whether multimodal models could learn campaign-performance patterns under controlled conditions. Synthetic data is powerful because it gives researchers a clean answer key. You can design the difficulty of the task, control the signal and stress-test models in repeatable ways.

    This new proof of concept takes a different step. It uses observed campaign performance data.

    The team started with more than 21 million daily campaign rows, covering the period from January 2023 to July 2026 across 75 anonymized brands, more than 1,000 audience representations, 7 objectives, and 42 platform representations. The data included 167 geography representations, with the United Kingdom present in each campaign row, sometimes alongside additional locations. After aggregation, filtering, and preparation, the data contained 20,067 distinct campaign setups.

    That sounds tidy, but real campaign data is rarely tidy in practice.

    Some combinations appear often, others barely appear at all. Some brands have lots of history, others have less. CTR is highly skewed, with many low values and a long tail of higher-performing cases.

    This is part of the value of working with real data. It forces the model to face the same imbalance and sparsity that planning teams face. There is no perfect laboratory. There is only the work of extracting useful structure from imperfect evidence.

    Brand identifiers were anonymized before modelling, so the system could learn brand-relative patterns without exposing identifiable brand names inside the modelling workflow.

    Teaching The Model To Read A Campaign Setup

    The model does not treat a campaign as a flat spreadsheet row. It learns from different kinds of campaign information and brings them together into a shared representation of the full setup.

    A brand is different from an audience. An audience is different from a platform. A platform is different from a geography. These are not interchangeable fields, so the model gives each campaign ingredient its own representation before combining them.

    The five model inputs and their roles are summarized in Table 6:

    InputWhat The Model Learns
    BrandBrand-specific performance context
    AudienceAudience behaviour patterns
    ObjectiveIntent and optimization goal
    PlatformChannel-specific behaviour
    GeographyMarket or regional context
    Table 6. Campaign inputs and the information learned by the model.

    Together, these inputs help the model learn whether a campaign setup appears coherent, whether it resembles combinations that worked well for the brand, and whether it deserves attention during planning.

    The same learned representation supports both prediction and recommendation, as shown in Table 7:

    Use CaseOutput
    PredictionEstimated CTR for a complete setup
    RecommendationRanked campaign completions for an incomplete setup
    Table 7. Model outputs for prediction and recommendation.

    The model is therefore not just producing a score. It is learning a map of campaign compatibility.

    Prediction: Estimating The Likely Outcome

    The first task was to test whether the model could predict CTR for campaign combinations it had not simply memorized.

    For most readers, the exact definitions of the prediction metrics are less important than the direction of the result: the model learned meaningful signal from the campaign ingredients and their relationships.

    It was not perfect. The model remained conservative in some high-CTR slices, especially for click-focused objectives. That matters because a production planning tool should not be overconfident when the performance distribution has a long tail. Calibration and monitoring would be needed before this kind of model became planner-facing.

    Still, prediction is only half the story. The more interesting question is what happens when the model is asked not just to score a campaign, but to suggest one.

    Recommendation: Finding Stronger Completions

    A recommendation system starts with an incomplete thought.

    A planner might know the brand, but still be deciding which audience, platform, objective or geography to prioritize. Or they might know the brand and objective, but still be choosing between possible activation contexts.

    The model’s job is to complete the thought in a useful way.

    Not randomly. Not by simply replaying a dashboard. And not by assuming that the highest global CTR is always the best answer.

    Instead, the system asks which possible completions look strong for the brand and coherent inside the learned campaign-performance space.

    This proof of concept tested the two recommendation settings summarized in Table 8:

    SettingMeaningWhy It Matters
    Closed vocabularyRecommends from campaign setups already present in the historical datasetEasier to evaluate because outcomes are known
    Open vocabularyProposes mostly new combinations that have not appeared historicallyCloser to the real planning opportunity
    Table 8. Closed- and open-vocabulary recommendation settings.

    Together, the two settings answer the different evaluation questions shown in Table 9:

    Closed VocabularyOpen Vocabulary
    Can the model rank known options better than random?Can the model suggest new combinations that still look plausible and promising?
    Table 9. Evaluation questions for closed- and open-vocabulary recommendation.

    Closed Vocabulary: Can It Rank Known Options?

    Closed vocabulary recommendation is the cleanest test of ranking quality.

    If the model recommends from historical rows, we can check whether its top suggestions were actually positive, neutral or negative for that brand. We can also compare it with random selection from the same candidate pool.

    Across 57 valid brand-fixed queries, the model performed substantially better than random selection, as shown in Table 10:

    MetricModelRandom from same candidate pools
    Positive precision@1076.2%50.2%
    Non-negative precision@1097.1%82.0%
    Negative rate@102.9%18.0%
    Known@10100.0%n/a
    Table 10. Closed-vocabulary recommendation performance compared with random selection.

    In plain terms, the model surfaced more historically strong options and avoided many more historically weak ones.

    Table 11 presents how performance changed according to how much of the campaign setup was already fixed:

    Query TypePositive@10Positive Lift vs RandomNegative Reduction vs Random
    Brand only93.0%3.085x28.3pp
    Brand + 1 modality68.7%1.683x16.0pp
    Brand + 2 modalities84.3%1.939x10.4pp
    Brand + 3 modalities65.6%1.276x6.5pp
    Table 11. Closed-vocabulary recommendation performance by query type.

    The strongest result came when the planner had fixed only the brand. In that setting, the model had space to explore many possible completions, and it found much stronger options than random selection.

    As more parts of the campaign were fixed, the search space became narrower. By the time brand plus three other ingredients were already specified, there was often only one meaningful decision left. In that case, the candidate pool itself was already favourable, so the model had less room to prove that it was adding ranking value.

    Product lesson:

    The recommender appears most valuable when it is used early enough in planning, while there is still meaningful room to shape the campaign.

    Open Vocabulary: Can It Suggest Something New?

    Closed vocabulary recommendation is measurable. Open vocabulary recommendation is more ambitious.

    In open vocabulary mode, the system can propose combinations that mostly do not exist in the historical dataset. This is closer to how planners actually think. They are not always asking for the best old campaign. They are often asking for a new plan that is informed by old evidence.

    The challenge is evaluation. If a recommended combination has never been run before, there is no historical label to check. We cannot say, from offline data alone, whether it definitely would have performed well.

    The team therefore used the diagnostic signals summarized in Table 12:

    Diagnostic QuestionWhat It Checks
    Is the recommendation close to positive examples?Whether it sits near strong historical patterns in the learned space
    Is predicted CTR high for the brand?Whether the setup looks strong relative to the brand’s own history
    Is the setup already known?Whether the model is copying old rows or proposing something new
    Table 12. Diagnostic questions used to evaluate open-vocabulary recommendations.

    Overall, the open-vocabulary recommendations showed encouraging evidence, as reported in Table 13:

    DiagnosticResult
    Closest-to-positive distribution rate72.5%
    Closest-to-non-negative distribution rate90.0%
    Closest-to-negative distribution rate10.0%
    Average predicted CTR percentile vs brand history87.2%
    Known@105.0%
    Table 13. Overall open-vocabulary recommendation diagnostics.

    The low Known@10 is important. It means the recommender was usually not just copying old rows. Most suggestions were genuinely new relative to the historical dataset.

    The strongest open-vocabulary evidence appeared for broader queries: brand-only, brand plus one ingredient, and brand plus two ingredients. These are the situations where the model has enough freedom to search the space and identify combinations that look both coherent and high-potential.

    When the query became too constrained, recommendation quality became less distinctive. Again, this is a planning lesson: the system is most useful when it can participate in shaping the campaign, not merely fill in the final blank.

    A New Kind Of Planning Conversation

    The most interesting version of this tool is not a black box that says “do this.”

    It is a system that changes the planning conversation.

    Today, a team might ask:

    “How did campaigns like this perform before?”

    With this kind of model, they can start asking:

    New Planning Questions

    What would make this setup stronger?

    Which options should we avoid?

    Is this recommendation strong for this brand, or just strong in general?

    Are we exploring genuinely new combinations, or repeating what we already know?

    If we fix the objective, what changes in the best platform and geography choices?

    Those are different questions. They move data from the end of the process into the beginning. They make historical evidence active at the moment decisions are being formed.

    The Human Role Becomes Even More Critical

    It is tempting to describe systems like this as automation. But the better description is augmentation.

    The model can search across many combinations quickly. It can learn patterns that are difficult to see in a dashboard. It can help prioritize options and identify setups that deserve a closer look.

    But it does not know the whole brief. It does not understand every client constraint. It does not know the timing constraints, market nuance, creative ambition, brand safety, media availability or whether a recommendation is commercially appropriate in a specific moment.

    That is why the planner remains central.

    The model can widen the field of view. The human still decides what is meaningful.

    In practice, the best version of this system would not simply output a single answer. It would support exploration. A planner could change the fixed ingredients, compare recommendation contexts, inspect confidence signals and decide whether a suggested combination is worth taking into expert review or live testing.

    The purpose is not to remove judgement. It is to give judgement a sharper instrument.

    What Makes This Different From The Earlier Work?

    The earlier Campaign Performance Modelling Pod work asked whether multimodal AI could learn campaign-performance patterns in controlled synthetic environments. That was an important step because it let the team test models under known conditions.

    This proof of concept asks a different question:

    Can similar ideas work on observed campaign performance data, where the labels are messier, the distribution is uneven, and the recommendation problem is closer to real planning?

    It also extends the ambition from prediction to recommendation. Instead of asking only “will this work?”, the system begins to ask “what should we try?”

    That is a meaningful shift. Prediction is useful when the planner already has a complete idea. Recommendation is useful when the idea is still being formed.

    What Comes Next

    This is still a proof of concept. Several things would need to happen before a system like this could move beyond the current form.

    The main steps required to move toward are outlined in Table 14:

    Next StepWhy It Matters
    Expert reviewCheck whether recommendations make strategic sense
    Controlled live testingValidate performance beyond offline diagnostics
    Better calibrationImprove reliability in high-CTR slices
    Confidence thresholdsHelp planners distinguish strong suggestions from weaker ones
    Recommendation explanationsShow which campaign ingredients are driving the suggestion
    Table 14. Proposed next steps toward a production planning workflow.

    That last point matters. A recommendation without an explanation is hard to trust. A useful planning tool should not only say:

    “This looks promising.”

    It should help the planner understand the evidence behind the suggestion.

    The direction, however, is clear.

    Campaign data should not only describe the past. It should help teams design the future.

    The more useful planning question is no longer just:

    “What happened?”

    It is:

    “Given what we know, what is worth trying next?”

    That is the shift this proof of concept begins to make: from reports that explain yesterday’s performance to recommendations that help shape tomorrow’s campaigns.

    Read The Technical Report

    Ready to explore the specifics? Read our full technical deep dive into our Technical Report for a closer look at our methodology.

    Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.

  • Performance-Aware Campaign Recommendations from Real Marketing Data

    Abstract

    Marketing teams often need to decide which campaign setup is worth testing before any budget is spent. The hard part is that performance does not depend on a single variable. It depends on the interaction between multiple factors, such as the brand, its audience, the campaign objective, the platform, the geography, and many other contextual signals.

    This proof of concept from the Performance AI team at Open Intelligence (OI PoC) addresses this as a multimodal prediction and recommendation problem. We trained a single model that learns from historical campaign performance, predicts click-through rate (CTR), and uses the same learned embedding space to recommend campaign completions. The result is a system that can move from “what happened before?” to “what should we try next?” while staying grounded in real performance data.

    This work builds on the earlier Campaign Performance Modelling Pod, which explored how AI can support campaign performance prediction using synthetic campaign data (see the Technical Report and Blog Post, respectively). The OI PoC analysis takes the next step by applying the same broader modelling agenda to observed/real performance data, trying to also examine both tasks of CTR prediction and recommendation in order to answer: given a partially specified campaign setup, what completion is most likely to perform well?

    The prepared dataset contains 20,067 unique real campaign setups derived from 21,247,221 raw daily rows. The evaluation is designed as a warm-start compositional generalization problem: the model sees known modality values during training, but must generalize to unseen full combinations at test time. This is important because it mirrors the real planning use case more closely than simple row memorization.

    In our evaluation, the model reached a test RMSE of 0.0151, MAE of 0.0052 and R2 of 0.6001 on unseen campaign combinations. For closed-vocabulary brand-fixed recommendations, the system achieved 76.2% positive precision@10 versus 50.2% for random selection from the same candidate pools, while reducing negative recommendations from 18.1% to 2.9%. For open-vocabulary recommendation, the strongest evidence appears in broader query settings where the model can still identify geometrically coherent completions that rank high relative to each brand’s historical CTR range.

    1. Motivation

    Campaign planning is highly combinatorial. A planner is not making one decision in isolation, but several interacting decisions at once: which audience to target, which campaign objective to optimize for, which platform to use, and which geography to prioritize for a given brand. Each of these choices can materially change the outcome, and the effect of any one decision depends on the others around it. In practice, that means the space of possible campaign setups grows much faster than the amount of historical evidence available to support them.

    This follows naturally from the previous Campaign Performance Modelling work. The earlier synthetic-data analysis asked whether AI could help move campaign planning from intuition-led guesswork toward more systematic performance foresight. Here, we ask a complementary question: once real OI performance data is available, can we learn a performance-aware representation that supports both CTR prediction and recommendation?

    This is exactly the gap OI PoC addresses. We want to help answer two related questions before budget is committed:

    1. If a team is considering a particular campaign setup, what CTR should they expect?
    2. If part of the setup is already fixed, what completion is most likely to perform well?

    The challenge is that historical data is both rich and sparse at the same time. At the raw row level there is plenty of volume, but once performance is aggregated into distinct campaign setups, only a small fraction of the theoretically possible combinations has ever been observed. A useful model therefore cannot depend on exact lookup or memorization. It needs to learn reusable structure from known examples and then transfer that structure to unseen combinations.

    This creates several practical modelling challenges:

    1. Prediction models can accidentally memorize historical rows instead of learning reusable structure.
    2. Recommendation systems can look accurate when the available candidate pool is already easy, even if the model is not adding much ranking value.
    3. Performance labels are not globally comparable across brands, because some brands naturally operate at different CTR ranges than others.
    4. Recommendation quality depends not just on whether a setup looks broadly similar to past positives, but on whether its internal combination of brand, audience, objective, platform and geography is coherent.

    The goal of OI PoC was to build a model that could generalize across known campaign ingredients in new configurations. This is a warm-start compositional generalization problem: the model has usually seen the individual ingredients before, but not the exact recipe. That is why the project combines CTR prediction and recommendation in a shared multimodal framework. The prediction side estimates likely performance for unseen setups, while the recommendation side uses the same learned representation to identify strong or plausibly strong completions for partially specified queries.

    2. Dataset

    The raw source dataset contained 21,247,221 daily campaign rows. After aggregation by campaign setup and filtering low-signal rows, this reduced to 20,067 unique learnable combinations. Furthermore, brand values were anonymized before being used as modelling inputs. This preserves the ability to learn brand-specific performance behaviour while keeping the modelling dataset separated from directly identifiable brands.

    The five campaign modalities and their coverage are summarized in Table 1:

    ModalityCoverage in the prepared dataset
    Brand75 anonymized brands
    Audiencemore than 1,000 audience representations
    Objective7 objectives
    Platform42 platform representations
    Geography167 geography representations
    Table 1. Campaign modalities and coverage in the prepared dataset.

    The train, validation and test split used 17,056 training rows, 1,505 validation rows and 1,506 test rows. The test set was designed to avoid exact-combination leakage.

    The split diagnostics are reported in Table 2:

    Split diagnosticResult
    Exact 5-way test combinations seen in training0 / 1,506
    Truly new 5-way test combinations1,506 / 1,506
    Test brands seen in training71 / 71
    Test objectives seen in training7 / 7
    Test platforms seen in training39 / 39
    Test geographies seen in training117 / 117
    Table 2. Test-set leakage and warm-start diagnostics.

    This means the evaluation is not a pure cold-start task. It is a test of whether the model can recombine known entities into unseen campaign setups.

    3. Why CTR Needed Brand-Relative Labels

    Click-through rate (CTR) serves as our example performance metric—a practical starting point given the available data. We use it across the dataset, including campaigns with objectives other than clicks, to explore whether the approach can support prediction and recommendation. The positive, neutral and negative labels therefore describe brand-relative CTR, not success against each campaign’s stated objective. Stronger click-through performance does not necessarily indicate stronger awareness, engagement or video-view outcomes. The next step is to extend the approach to available metrics better suited to those objectives.

    CTR distribution statistics are shown in Table 3, proving that CTR was highly skewed:

    StatisticCTR
    Mean0.0106
    Median0.0027
    Skewness7.8710
    Kurtosis108.02
    Table 3. Summary statistics of the CTR distribution.

    A global definition of “positive” campaign performance would over-favour brands with naturally higher CTR and penalise brands with lower typical CTR. Instead, each brand’s own CTR distribution was split into rank-based thirds: negative, neutral and positive.

    This created labels that mean “good for this brand”, rather than “high CTR globally”. The anonymized brand identifiers are still consistent across rows, which means the model can learn within-brand CTR patterns without needing access to the original brand names.

    The final label distribution is reported in Table 4:

    LabelRows
    Negative5,984
    Neutral8,030
    Positive6,053
    Table 4. Brand-relative CTR label distribution.

    Rank-based splitting was used because many brands had zero-heavy CTR distributions. Standard quantile thresholds often collapsed to the same value, while rank-based labels preserved useful within-brand ordering.

    4. Model Approach

    OI PoC uses one shared model for both prediction and recommendation.

    Each modality starts from a 256-dimensional embedding and passes through its own projector:

    • brand projector
    • audience projector
    • objective projector
    • platform projector
    • geography projector

    The projected modality embeddings are pooled into a shared campaign representation. From there, the model optimizes two signals:

    • a regression head predicts transformed CTR;
    • a centered contrastive objective shapes the embedding geometry so that positive, neutral and negative campaign setups become easier to retrieve and rank.

    The shared model architecture is illustrated in Figure 1:

    Figure 1. Shared multimodal architecture for CTR prediction and campaign recommendation.

    The centered contrastive objective is important because it encodes a more useful recommendation bias than a standard label-only contrastive loss. In this setting, we do not simply want all positive examples to collapse toward one another in the embedding space. Two campaign setups can both be labelled positive while being strategically very different in terms of brand, audience, objective, platform and geography. Forcing all positives to become globally similar would risk bringing together irrelevant campaign setups that happen to share the same final label but not the same internal structure.

    Instead, the centered loss focuses on intra-sample consistency. For each projected modality embedding, the model compares it against the center formed by the other modalities within the same campaign setup. This encourages the representation to answer a more precise question: do the parts of this setup belong together in a performance-aware way? Positive tuples should exhibit stronger internal agreement, neutral tuples should sit closer to an intermediate structure, and negative tuples should show weaker alignment. That is much closer to the actual recommendation task, where the model needs to judge whether a partially specified setup can be completed coherently rather than whether it merely resembles some broad global “positive” cloud.

    This design also makes the loss batch-sensitive in a useful way. The implementation weights modality contributions using both global and batch-level information. Global weighting reflects the overall cardinality of each modality vocabulary, while batch weighting reflects the effective diversity of modality values present in the current batch. In practice, this helps prevent the contrastive term from over-responding either to globally large modalities such as audience or to accidentally repetitive mini-batches where one modality has very little variation. The current implementation uses gamma_global = 0.5 and gamma_batch = 0.5, which gives a balanced influence to long-run dataset structure and local batch composition.

    The centered loss is also deliberately moderated relative to the CTR objective. In training, the total loss combines Huber regression with the centered contrastive term using LAMBDA_CENTERED = 0.05, plus a small norm penalty inside the contrastive component. This keeps CTR prediction as the dominant optimization target while still giving the embedding space enough geometric structure to support retrieval and reranking. The result is a representation that is not purely semantic and not purely regressive: it is shaped to preserve campaign-setup compatibility in a way that is directly useful for recommendation.

    This design keeps deployment simple: one model artifact, one embedding space and one monitoring surface. It also makes the recommendation layer performance-aware rather than purely semantic.

    5. Evaluation Design

    The evaluation focused on three questions:

    1. Can the model predict CTR for unseen campaign combinations?
    2. Can the learned representation identify stronger recommendations than random selection from the same candidate pool?
    3. Can the model propose new combinations in open vocabulary settings without simply replaying historical rows?

    For recommendations, the system uses a two-stage retrieve-then-rerank pipeline, as summarized in Table 5:

    StageRole
    Geometry retrievalFind low-distance candidates in the learned embedding space
    CTR rerankingSort retrieved candidates by predicted CTR
    Top-k outputReturn the best campaign completions
    Table 5. Retrieve-then-rerank recommendation pipeline.

    This is important operationally because the embedding space provides a fast candidate filter, while the regression head gives the final ranking signal.

    6. Regression Results

    As reported in Table 6, on the 1,506-row test set the model captured a meaningful share of CTR variance for unseen combinations:

    MetricTest result
    RMSE0.0151
    MAE0.0052
    R20.6001
    Table 6. Regression performance on the test set.

    The model is useful but conservative in high-CTR slices. For LINK_CLICKS, the actual average CTR was 0.02370 while the predicted average CTR was 0.01938, a bias of -0.00432. This is expected given the heavy-tailed target distribution and should be monitored before production use.

    Permutation diagnostics are shown in Table 7, demonstrating that objective and platform were the strongest regression drivers:

    Modality permutedMean MAE increase
    Objective0.8490
    Platform0.6561
    Brand0.2018
    Audience0.1184
    Geography0.0342
    Table 7. Permutation importance by campaign modality.

    The geography result is likely influenced by the dataset composition, where the prepared rows are heavily UK-centered.

    7. Closed-Vocabulary Recommendation Results

    Closed-vocabulary recommendations are completions drawn from rows already present in the historical dataset. This lets us evaluate recommendation quality against known labels.

    In practical terms, “closed” means the recommender is choosing from campaign setups that have already been observed historically, rather than inventing entirely new combinations. That makes this the most directly measurable recommendation setting: every retrieved completion can be checked against an existing brand-relative label, so we know whether the model is surfacing historically positive, neutral or negative outcomes. Closed-vocabulary evaluation is therefore the clearest test of ranking quality. If the model is genuinely useful, it should place more positive combinations near the top of the list and push negative combinations down, even when the available candidate pool is mixed.

    We group recommendation queries by how much of the campaign setup is already fixed. “Brand only” fixes the brand and leaves audience, objective, platform and geography open for recommendation. “Brand + 1” fixes the brand plus one of those four ingredients; “Brand + 2” fixes the brand plus two; and “Brand + 3” fixes the brand plus three. For example, a query with a fixed brand and objective is a “Brand + 1” query: the model recommends the remaining audience, platform and geography.

    Overall closed-vocabulary results are reported in Table 8. Across 57 valid evaluated brand-fixed queries, OI PoC achieved:

    MetricModelRandom from same candidate pools
    Positive precision@1076.2%50.2%
    Non-negative precision@1097.1%82.0%
    Negative rate@102.9%18.0%
    Known@10100.0%n/a
    Table 8. Closed-vocabulary recommendation performance compared with random selection.

    The model produced a 1.92x positive lift versus random selection and reduced negative recommendations by 15.1 percentage points.

    Results by query type are presented in Table 9. The strongest result came from broader brand-fixed queries:

    Query typePositive@10Positive lift vs randomNegative reduction vs random
    Brand only93.0%3.085x28.3pp
    Brand + 1 modality68.7%1.683x16.0pp
    Brand + 2 modalities84.3%1.939x10.4pp
    Brand + 3 modalities65.6%1.276x6.5pp
    Table 9. Closed-vocabulary performance by query type.

    Figure 2 shows the positive recommendation rate by query type, while Figure 3 shows the corresponding negative recommendation rate. The former demonstrates how often the top recommendations fall into the historically positive class for each query type. This is the most intuitive measure of recommendation value: higher is better, because it means the model succeeds in surfacing combinations that were historically strong for the given brand. The latter depicts the opposite side of the same story. Lower is better, because it means the recommender is avoiding combinations that were historically weak.

    Figure 2. Positive precision@10 by closed-vocabulary query type.
    Figure 3. Negative rate@10 by closed-vocabulary query type.

    These plots are important because raw precision alone can be misleading if the candidate pool is already easy. That is why the comparison against random selection from the same candidate pool matters so much. The model is not just scoring well because it is choosing among already strong options. In brand-only queries, the historical candidate pool contains only about 30.2% positives, yet the recommender raises positive precision@10 to 93.0% and cuts negatives from 29.3% to 1.0%. Brand + 1 remains very strong, with 44.6% candidate positives rising to 68.7% positive@10 and negatives falling from 19.0% to 3.0%. Brand + 2 is still strong, lifting positives from 68.8% to 84.3% and lowering negatives from 14.7% to 4.3%.

    Brand + 3 needs a more careful reading. On the surface, 65.6% positive@10 and 2.7% negative@10 still look solid. But the candidate pool for this case is already extremely favourable, with 57.0% positives and only 9.2% negatives before the model ranks anything. In other words, once four parts of the setup are already fixed, there is often very little ranking difficulty left. That is why the lift versus random is close to flat in this setting. The model is still producing sensible recommendations, but the closed-vocabulary evidence no longer shows a strong ranking advantage over the baseline.

    The interpretation is therefore more nuanced than simply saying “the model works everywhere.” Closed-vocabulary results show clear ranking value when the query leaves enough room for the model to discriminate among many possible completions. That is most visible in brand-only, brand + 1 and still present in brand + 2 queries, and much less convincing in brand + 3 where the space is already heavily constrained. This is exactly the kind of pattern we would hope to see from a recommender that is useful in realistic planning scenarios rather than only in highly filtered cases.

    8. Open-Vocabulary Recommendation Results

    Open-vocabulary recommendations generate combinations that are mostly not present in the historical dataset. This is the more interesting planning use case, but also the harder one to validate offline.

    Because most open-vocabulary recommendations are unknown historically, the evaluation relies on geometric diagnostics and predicted CTR relative to each brand’s historical distribution.

    Two diagnostics are especially useful here. The first is centered cosine dissimilarity, which measures how well a recommended completion fits the learned geometry of a query. Lower dissimilarity means the retrieved combination sits closer to the model’s positive reference structure; higher dissimilarity suggests a weaker or more ambiguous fit. The second is predicted CTR percentile versus brand history, which asks where the model’s predicted CTR for a recommendation sits relative to that brand’s own historical CTR distribution. A percentile near 1.0 means the recommendation is predicted to perform near the top end of what that brand has historically achieved, while a percentile closer to 0.5 means the recommendation looks more typical than exceptional.

    Overall open-vocabulary diagnostics are summarized in Table 10:

    DiagnosticResult
    Closest-to-positive distribution rate72.5%
    Closest-to-non-negative distribution rate90.0%
    Closest-to-negative distribution rate10.0%
    Average predicted CTR percentile vs brand history87.2%
    Known@105.0%
    Table 10. Overall open-vocabulary recommendation diagnostics.

    As depicted in Figure 4, the centered cosine dissimilarity plot shows that the model is most confident in broad recommendation settings. Brand-only queries have the strongest geometric fit, followed by brand + 1 and brand + 2 queries. This is consistent with the retrieval results: when the query leaves enough freedom for the model to search the space, the learned geometry can still find combinations that resemble historically positive structures. Brand + 3 behaves differently, with substantially weaker positive-distribution alignment and a higher negative-distribution rate. This suggests that once the query becomes too constrained, the geometry has much less room to identify clearly superior completions.

    Figure 4. Centered cosine dissimilarity by open-vocabulary query type.

    The CTR percentile plot tells a complementary story (see Figure 5). For brand-only recommendations, the mean predicted CTR lands around the 98.1st percentile of the brand’s historical distribution. Brand + 1 remains similarly strong at 94.9%, and brand + 2 still sits high at 91.2%. These are not just ‘acceptable’ completions; they are recommendations whose predicted CTR places them near the upper end of each brand’s historical range. Brand + 3 drops to roughly the 64.7th percentile, which is still above median, but much less distinctive. In other words, the model still finds plausible completions in this more constrained setting, but it is no longer consistently surfacing outcomes that look clearly elite for the brand.

    Figure 5. Predicted CTR percentile relative to each brand’s historical distribution.

    Taken together, the two plots support the same interpretation. The open-vocabulary recommender is most credible when it has enough flexibility to discover new completions rather than simply fill in one missing slot in an already highly specified setup. In broad and moderately constrained contexts, the recommended combinations both align strongly with the positive geometry of the learned space and score highly against the brand’s own historical CTR range. In heavily constrained contexts, the model can still produce sensible outputs, but the diagnostic strength is noticeably weaker and the case for production use becomes more cautious.

    Low known@10 is expected and desirable in this setting: it means the recommender is proposing new combinations rather than simply replaying historical rows. The overall known@10 rate is only 5.0%, confirming that most recommended combinations are genuinely new relative to the historical dataset.

    Overall, the strongest open-vocabulary evidence appears for brand-only, brand + 1 and brand + 2 queries. Brand + 3 degrades materially because the query is already highly constrained and leaves less room for the model to improve the completion.

    9. What This Enables

    OI PoC turns historical campaign data into a planning tool with three practical uses:

    1. CTR prediction for unseen combinations of known campaign entities.
    2. Brand-relative recommendation, where “positive” means strong for that brand rather than globally high CTR.
    3. Discovery of new campaign completions that are plausible in the learned performance-aware embedding space.

    The most defensible claim from the current results is:

    Among sufficiently supported queries, the model improves positive recommendation rate and sharply reduces negatives versus random selection, especially for brand-only, brand + 1 and brand + 2 recommendation contexts.

    What’s next?

    There are several directions we’re excited to explore from here. One is understanding when the model is confident enough in a recommendation to make it useful, combining signals such as retrieval distance, predicted CTR and the diversity of the candidate pool.

    We’re also interested in making recommendations easier to interpret: which modalities contribute most to predicted CTR, and what makes a particular campaign completion rank highly?

    Another direction is improving calibration across high-CTR segments and objectives such as LINK_CLICKS, while continuing to validate open-vocabulary recommendations with expert review.

    Ultimately, the most interesting question is how these recommendations perform beyond offline evaluation. A controlled online experiment comparing model-recommended campaign completions with existing planning approaches would be a natural next step.

    There’s more to explore here — stay tuned.