{"id":1760,"date":"2026-09-30T14:56:29","date_gmt":"2026-09-30T14:56:29","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=1760"},"modified":"2026-09-30T15:25:55","modified_gmt":"2026-09-30T15:25:55","slug":"performance-aware-campaign-recommendations-from-real-marketing-data","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=performance-aware-campaign-recommendations-from-real-marketing-data","title":{"rendered":"Performance-Aware Campaign Recommendations from Real Marketing Data"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Abstract<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Marketing teams often need to decide which campaign setup is worth testing before any budget is spent. The hard part is that performance does not depend on a single variable. It depends on the interaction between multiple factors, such as the brand, its audience, the campaign objective, the platform, the geography, and many other contextual signals.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This proof of concept from the Performance AI team at Open Intelligence (OI PoC) addresses this as a multimodal prediction and recommendation problem. We trained a single model that learns from historical campaign performance, predicts click-through rate (CTR), and uses the same learned embedding space to recommend campaign completions. The result is a system that can move from \u201cwhat happened before?\u201d to \u201cwhat should we try next?\u201d while staying grounded in <strong>real performance data<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This work builds on the earlier <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">Campaign Performance Modelling Pod<\/a>, which explored how AI can support campaign performance prediction using synthetic campaign data (see the <a href=\"https:\/\/research.wpp.com\/reports\/campaign-performance-modelling-pod-technical-walkthrough\">Technical Report<\/a> and <a href=\"https:\/\/research.wpp.com\/blog\/from-guesswork-to-foresight-how-ai-is-predicting-the-future-of-marketing-campaigns\">Blog Post<\/a>, respectively). The OI PoC analysis takes the next step by applying the same broader modelling agenda to <strong>observed\/real performance data<\/strong>, trying to also examine both tasks of CTR prediction and recommendation in order to answer: given a partially specified campaign setup, what completion is most likely to perform well?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The prepared dataset contains 20,067 unique real campaign setups derived from 21,247,221 raw daily rows. The evaluation is designed as a warm-start compositional generalization problem: the model sees known modality values during training, but must generalize to unseen full combinations at test time. This is important because it mirrors the real planning use case more closely than simple row memorization.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In our evaluation, the model reached a test RMSE of 0.0151, MAE of 0.0052 and R2 of 0.6001 on unseen campaign combinations. For closed-vocabulary brand-fixed recommendations, the system achieved 76.2% positive precision@10 versus 50.2% for random selection from the same candidate pools, while reducing negative recommendations from 18.1% to 2.9%. For open-vocabulary recommendation, the strongest evidence appears in broader query settings where the model can still identify geometrically coherent completions that rank high relative to each brand\u2019s historical CTR range.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">1. Motivation<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Campaign planning is highly combinatorial. A planner is not making one decision in isolation, but several interacting decisions at once: which audience to target, which campaign objective to optimize for, which platform to use, and which geography to prioritize for a given brand. Each of these choices can materially change the outcome, and the effect of any one decision depends on the others around it. In practice, that means the space of possible campaign setups grows much faster than the amount of historical evidence available to support them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This follows naturally from the previous <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">Campaign Performance Modelling<\/a> work. The earlier synthetic-data analysis asked whether AI could help move campaign planning from intuition-led guesswork toward more systematic performance foresight. Here, we ask a complementary question: once real OI performance data is available, can we learn a performance-aware representation that supports both CTR prediction and recommendation?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is exactly the gap OI PoC addresses. We want to help answer two related questions before budget is committed:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>If a team is considering a particular campaign setup, what CTR should they expect?<\/li>\n\n\n\n<li>If part of the setup is already fixed, what completion is most likely to perform well?<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The challenge is that historical data is both rich and sparse at the same time. At the raw row level there is plenty of volume, but once performance is aggregated into distinct campaign setups, only a small fraction of the theoretically possible combinations has ever been observed. A useful model therefore cannot depend on exact lookup or memorization. It needs to learn reusable structure from known examples and then transfer that structure to unseen combinations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This creates several practical modelling challenges:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Prediction models can accidentally memorize historical rows instead of learning reusable structure.<\/li>\n\n\n\n<li>Recommendation systems can look accurate when the available candidate pool is already easy, even if the model is not adding much ranking value.<\/li>\n\n\n\n<li>Performance labels are not globally comparable across brands, because some brands naturally operate at different CTR ranges than others.<\/li>\n\n\n\n<li>Recommendation quality depends not just on whether a setup looks broadly similar to past positives, but on whether its internal combination of brand, audience, objective, platform and geography is coherent.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The goal of OI PoC was to build a model that could generalize across known campaign ingredients in new configurations. This is a warm-start compositional generalization problem: the model has usually seen the individual ingredients before, but not the exact recipe. That is why the project combines CTR prediction and recommendation in a shared multimodal framework. The prediction side estimates likely performance for unseen setups, while the recommendation side uses the same learned representation to identify strong or plausibly strong completions for partially specified queries.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">2. Dataset<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The raw source dataset contained 21,247,221 daily campaign rows. After aggregation by campaign setup and filtering low-signal rows, this reduced to 20,067 unique learnable combinations. Furthermore, brand values were anonymized before being used as modelling inputs. This preserves the ability to learn brand-specific performance behaviour while keeping the modelling dataset separated from directly identifiable brands.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The five campaign modalities and their coverage are summarized in Table 1:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Modality<\/th><th>Coverage in the prepared dataset<\/th><\/tr><\/thead><tbody><tr><td>Brand<\/td><td>75 anonymized brands<\/td><\/tr><tr><td>Audience<\/td><td>more than 1,000 audience representations<\/td><\/tr><tr><td>Objective<\/td><td>7 objectives<\/td><\/tr><tr><td>Platform<\/td><td>42 platform representations<\/td><\/tr><tr><td>Geography<\/td><td>167 geography representations<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 1. Campaign modalities and coverage in the prepared dataset.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The train, validation and test split used 17,056 training rows, 1,505 validation rows and 1,506 test rows. The test set was designed to avoid exact-combination leakage. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The split diagnostics are reported in Table 2:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Split diagnostic<\/th><th>Result<\/th><\/tr><\/thead><tbody><tr><td>Exact 5-way test combinations seen in training<\/td><td>0 \/ 1,506<\/td><\/tr><tr><td>Truly new 5-way test combinations<\/td><td>1,506 \/ 1,506<\/td><\/tr><tr><td>Test brands seen in training<\/td><td>71 \/ 71<\/td><\/tr><tr><td>Test objectives seen in training<\/td><td>7 \/ 7<\/td><\/tr><tr><td>Test platforms seen in training<\/td><td>39 \/ 39<\/td><\/tr><tr><td>Test geographies seen in training<\/td><td>117 \/ 117<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 2. Test-set leakage and warm-start diagnostics.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This means the evaluation is not a pure cold-start task. It is a test of whether the model can recombine known entities into unseen campaign setups.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">3. Why CTR Needed Brand-Relative Labels<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Click-through rate (CTR) serves as our example performance metric\u2014a practical starting point given the available data. We use it across the dataset, including campaigns with objectives other than clicks, to explore whether the approach can support prediction and recommendation. The positive, neutral and negative labels therefore describe brand-relative CTR, not success against each campaign\u2019s stated objective. Stronger click-through performance does not necessarily indicate stronger awareness, engagement or video-view outcomes. The next step is to extend the approach to available metrics better suited to those objectives.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">CTR distribution statistics are shown in Table 3, proving that CTR was highly skewed:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Statistic<\/th><th>CTR<\/th><\/tr><\/thead><tbody><tr><td>Mean<\/td><td>0.0106<\/td><\/tr><tr><td>Median<\/td><td>0.0027<\/td><\/tr><tr><td>Skewness<\/td><td>7.8710<\/td><\/tr><tr><td>Kurtosis<\/td><td>108.02<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 3. Summary statistics of the CTR distribution.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">A global definition of \u201cpositive\u201d campaign performance would over-favour brands with naturally higher CTR and penalise brands with lower typical CTR. Instead, each brand\u2019s own CTR distribution was split into rank-based thirds: negative, neutral and positive.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This created labels that mean \u201cgood for this brand\u201d, rather than \u201chigh CTR globally\u201d. The anonymized brand identifiers are still consistent across rows, which means the model can learn within-brand CTR patterns without needing access to the original brand names.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The final label distribution is reported in Table 4:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Label<\/th><th>Rows<\/th><\/tr><\/thead><tbody><tr><td>Negative<\/td><td>5,984<\/td><\/tr><tr><td>Neutral<\/td><td>8,030<\/td><\/tr><tr><td>Positive<\/td><td>6,053<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 4. Brand-relative CTR label distribution.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Rank-based splitting was used because many brands had zero-heavy CTR distributions. Standard quantile thresholds often collapsed to the same value, while rank-based labels preserved useful within-brand ordering.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">4. Model Approach<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">OI PoC uses one shared model for both prediction and recommendation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Each modality starts from a 256-dimensional embedding and passes through its own projector:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>brand projector<\/li>\n\n\n\n<li>audience projector<\/li>\n\n\n\n<li>objective projector<\/li>\n\n\n\n<li>platform projector<\/li>\n\n\n\n<li>geography projector<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The projected modality embeddings are pooled into a shared campaign representation. From there, the model optimizes two signals:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>a regression head predicts transformed CTR;<\/li>\n\n\n\n<li>a centered contrastive objective shapes the embedding geometry so that positive, neutral and negative campaign setups become easier to retrieve and rank.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The shared model architecture is illustrated in Figure 1:<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"583\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1024x583.jpg\" alt=\"\" class=\"wp-image-1771\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1024x583.jpg 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-300x171.jpg 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-768x437.jpg 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1536x874.jpg 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2.jpg 1708w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 1. Shared multimodal architecture for CTR prediction and campaign recommendation.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The centered contrastive objective is important because it encodes a more useful recommendation bias than a standard label-only contrastive loss. In this setting, we do not simply want all positive examples to collapse toward one another in the embedding space. Two campaign setups can both be labelled positive while being strategically very different in terms of brand, audience, objective, platform and geography. Forcing all positives to become globally similar would risk bringing together irrelevant campaign setups that happen to share the same final label but not the same internal structure.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instead, the centered loss focuses on intra-sample consistency. For each projected modality embedding, the model compares it against the center formed by the other modalities within the same campaign setup. This encourages the representation to answer a more precise question: do the parts of this setup belong together in a performance-aware way? Positive tuples should exhibit stronger internal agreement, neutral tuples should sit closer to an intermediate structure, and negative tuples should show weaker alignment. That is much closer to the actual recommendation task, where the model needs to judge whether a partially specified setup can be completed coherently rather than whether it merely resembles some broad global \u201cpositive\u201d cloud.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This design also makes the loss batch-sensitive in a useful way. The implementation weights modality contributions using both global and batch-level information. Global weighting reflects the overall cardinality of each modality vocabulary, while batch weighting reflects the effective diversity of modality values present in the current batch. In practice, this helps prevent the contrastive term from over-responding either to globally large modalities such as audience or to accidentally repetitive mini-batches where one modality has very little variation. The current implementation uses <code>gamma_global = 0.5<\/code> and <code>gamma_batch = 0.5<\/code>, which gives a balanced influence to long-run dataset structure and local batch composition.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The centered loss is also deliberately moderated relative to the CTR objective. In training, the total loss combines Huber regression with the centered contrastive term using <code>LAMBDA_CENTERED = 0.05<\/code>, plus a small norm penalty inside the contrastive component. This keeps CTR prediction as the dominant optimization target while still giving the embedding space enough geometric structure to support retrieval and reranking. The result is a representation that is not purely semantic and not purely regressive: it is shaped to preserve campaign-setup compatibility in a way that is directly useful for recommendation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This design keeps deployment simple: one model artifact, one embedding space and one monitoring surface. It also makes the recommendation layer performance-aware rather than purely semantic.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">5. Evaluation Design<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The evaluation focused on three questions:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Can the model predict CTR for unseen campaign combinations?<\/li>\n\n\n\n<li>Can the learned representation identify stronger recommendations than random selection from the same candidate pool?<\/li>\n\n\n\n<li>Can the model propose new combinations in open vocabulary settings without simply replaying historical rows?<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">For recommendations, the system uses a two-stage retrieve-then-rerank pipeline, as summarized in Table 5:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Stage<\/th><th>Role<\/th><\/tr><\/thead><tbody><tr><td>Geometry retrieval<\/td><td>Find low-distance candidates in the learned embedding space<\/td><\/tr><tr><td>CTR reranking<\/td><td>Sort retrieved candidates by predicted CTR<\/td><\/tr><tr><td>Top-k output<\/td><td>Return the best campaign completions<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 5. Retrieve-then-rerank recommendation pipeline.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This is important operationally because the embedding space provides a fast candidate filter, while the regression head gives the final ranking signal.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">6. Regression Results<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">As reported in Table 6, on the 1,506-row test set the model captured a meaningful share of CTR variance for unseen combinations:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Metric<\/th><th>Test result<\/th><\/tr><\/thead><tbody><tr><td>RMSE<\/td><td>0.0151<\/td><\/tr><tr><td>MAE<\/td><td>0.0052<\/td><\/tr><tr><td>R2<\/td><td>0.6001<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 6. Regression performance on the test set.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The model is useful but conservative in high-CTR slices. For <code>LINK_CLICKS<\/code>, the actual average CTR was 0.02370 while the predicted average CTR was 0.01938, a bias of -0.00432. This is expected given the heavy-tailed target distribution and should be monitored before production use.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Permutation diagnostics are shown in Table 7, demonstrating that objective and platform were the strongest regression drivers:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Modality permuted<\/th><th>Mean MAE increase<\/th><\/tr><\/thead><tbody><tr><td>Objective<\/td><td>0.8490<\/td><\/tr><tr><td>Platform<\/td><td>0.6561<\/td><\/tr><tr><td>Brand<\/td><td>0.2018<\/td><\/tr><tr><td>Audience<\/td><td>0.1184<\/td><\/tr><tr><td>Geography<\/td><td>0.0342<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 7. Permutation importance by campaign modality.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The geography result is likely influenced by the dataset composition, where the prepared rows are heavily UK-centered.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">7. Closed-Vocabulary Recommendation Results<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Closed-vocabulary recommendations are completions drawn from rows already present in the historical dataset. This lets us evaluate recommendation quality against known labels.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In practical terms, \u201cclosed\u201d means the recommender is choosing from campaign setups that have already been observed historically, rather than inventing entirely new combinations. That makes this the most directly measurable recommendation setting: every retrieved completion can be checked against an existing brand-relative label, so we know whether the model is surfacing historically positive, neutral or negative outcomes. Closed-vocabulary evaluation is therefore the clearest test of ranking quality. If the model is genuinely useful, it should place more positive combinations near the top of the list and push negative combinations down, even when the available candidate pool is mixed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We group recommendation queries by how much of the campaign setup is already fixed. \u201cBrand only\u201d fixes the brand and leaves audience, objective, platform and geography open for recommendation. \u201cBrand + 1\u201d fixes the brand plus one of those four ingredients; \u201cBrand + 2\u201d fixes the brand plus two; and \u201cBrand + 3\u201d fixes the brand plus three. For example, a query with a fixed brand and objective is a \u201cBrand + 1\u201d query: the model recommends the remaining audience, platform and geography.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall closed-vocabulary results are reported in Table 8. Across 57 valid evaluated brand-fixed queries, OI PoC achieved:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Metric<\/th><th>Model<\/th><th>Random from same candidate pools<\/th><\/tr><\/thead><tbody><tr><td>Positive precision@10<\/td><td>76.2%<\/td><td>50.2%<\/td><\/tr><tr><td>Non-negative precision@10<\/td><td>97.1%<\/td><td>82.0%<\/td><\/tr><tr><td>Negative rate@10<\/td><td>2.9%<\/td><td>18.0%<\/td><\/tr><tr><td>Known@10<\/td><td>100.0%<\/td><td>n\/a<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 8. Closed-vocabulary recommendation performance compared with random selection.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The model produced a 1.92x positive lift versus random selection and reduced negative recommendations by 15.1 percentage points.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Results by query type are presented in Table 9. The strongest result came from broader brand-fixed queries:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Query type<\/th><th>Positive@10<\/th><th>Positive lift vs random<\/th><th>Negative reduction vs random<\/th><\/tr><\/thead><tbody><tr><td>Brand only<\/td><td>93.0%<\/td><td>3.085x<\/td><td>28.3pp<\/td><\/tr><tr><td>Brand + 1 modality<\/td><td>68.7%<\/td><td>1.683x<\/td><td>16.0pp<\/td><\/tr><tr><td>Brand + 2 modalities<\/td><td>84.3%<\/td><td>1.939x<\/td><td>10.4pp<\/td><\/tr><tr><td>Brand + 3 modalities<\/td><td>65.6%<\/td><td>1.276x<\/td><td>6.5pp<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 9. Closed-vocabulary performance by query type.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 2 shows the positive recommendation rate by query type, while Figure 3 shows the corresponding negative recommendation rate. The former demonstrates how often the top recommendations fall into the historically positive class for each query type. This is the most intuitive measure of recommendation value: higher is better, because it means the model <em>succeeds<\/em> in surfacing combinations that were historically strong for the given brand. The latter depicts the opposite side of the same story. Lower is better, because it means the recommender is avoiding combinations that were historically weak.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"558\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1024x558.png\" alt=\"\" class=\"wp-image-1950\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1024x558.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-300x163.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-767x418.png 767w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-2048x1116.png 2048w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1536x837.png 1536w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 2. Positive precision@10 by closed-vocabulary query type.<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"558\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1024x558.png\" alt=\"\" class=\"wp-image-1951\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1024x558.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-300x163.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-768x418.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1536x836.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-2048x1115.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 3. Negative rate@10 by closed-vocabulary query type.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">These plots are important because raw precision alone can be misleading if the candidate pool is already easy. That is why the comparison against random selection from the same candidate pool matters so much. The model is not just scoring well because it is choosing among already strong options. In brand-only queries, the historical candidate pool contains only about 30.2% positives, yet the recommender raises positive precision@10 to 93.0% and cuts negatives from 29.3% to 1.0%. Brand + 1 remains very strong, with 44.6% candidate positives rising to 68.7% positive@10 and negatives falling from 19.0% to 3.0%. Brand + 2 is still strong, lifting positives from 68.8% to 84.3% and lowering negatives from 14.7% to 4.3%.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Brand + 3 needs a more careful reading. On the surface, 65.6% positive@10 and 2.7% negative@10 still look solid. But the candidate pool for this case is already extremely favourable, with 57.0% positives and only 9.2% negatives before the model ranks anything. In other words, once four parts of the setup are already fixed, there is often very little ranking difficulty left. That is why the lift versus random is close to flat in this setting. The model is still producing sensible recommendations, but the closed-vocabulary evidence no longer shows a strong ranking advantage over the baseline.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The interpretation is therefore more nuanced than simply saying \u201cthe model works everywhere.\u201d Closed-vocabulary results show clear ranking value when the query leaves enough room for the model to discriminate among many possible completions. That is most visible in brand-only, brand + 1 and still present in brand + 2 queries, and much less convincing in brand + 3 where the space is already heavily constrained. This is exactly the kind of pattern we would hope to see from a recommender that is useful in realistic planning scenarios rather than only in highly filtered cases.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">8. Open-Vocabulary Recommendation Results<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Open-vocabulary recommendations generate combinations that are mostly not present in the historical dataset. This is the more interesting planning use case, but also the harder one to validate offline.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because most open-vocabulary recommendations are unknown historically, the evaluation relies on geometric diagnostics and predicted CTR relative to each brand\u2019s historical distribution.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two diagnostics are especially useful here. The first is centered cosine dissimilarity, which measures how well a recommended completion fits the learned geometry of a query. Lower dissimilarity means the retrieved combination sits closer to the model\u2019s positive reference structure; higher dissimilarity suggests a weaker or more ambiguous fit. The second is predicted CTR percentile versus brand history, which asks where the model\u2019s predicted CTR for a recommendation sits relative to that brand\u2019s own historical CTR distribution. A percentile near 1.0 means the recommendation is predicted to perform near the top end of what that brand has historically achieved, while a percentile closer to 0.5 means the recommendation looks more typical than exceptional.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall open-vocabulary diagnostics are summarized in Table 10:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Diagnostic<\/th><th>Result<\/th><\/tr><\/thead><tbody><tr><td>Closest-to-positive distribution rate<\/td><td>72.5%<\/td><\/tr><tr><td>Closest-to-non-negative distribution rate<\/td><td>90.0%<\/td><\/tr><tr><td>Closest-to-negative distribution rate<\/td><td>10.0%<\/td><\/tr><tr><td>Average predicted CTR percentile vs brand history<\/td><td>87.2%<\/td><\/tr><tr><td>Known@10<\/td><td>5.0%<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 10. Overall open-vocabulary recommendation diagnostics.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">As depicted in Figure 4, the centered cosine dissimilarity plot shows that the model is most confident in broad recommendation settings. Brand-only queries have the strongest geometric fit, followed by brand + 1 and brand + 2 queries. This is consistent with the retrieval results: when the query leaves enough freedom for the model to search the space, the learned geometry can still find combinations that resemble historically positive structures. Brand + 3 behaves differently, with substantially weaker positive-distribution alignment and a higher negative-distribution rate. This suggests that once the query becomes too constrained, the geometry has much less room to identify clearly superior completions.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"517\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-1024x517.png\" alt=\"\" class=\"wp-image-1952\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-1024x517.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-300x152.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-768x388.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case.png 1506w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 4. Centered cosine dissimilarity by open-vocabulary query type.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The CTR percentile plot tells a complementary story (see Figure 5). For brand-only recommendations, the mean predicted CTR lands around the 98.1st percentile of the brand\u2019s historical distribution. Brand + 1 remains similarly strong at 94.9%, and brand + 2 still sits high at 91.2%. These are not just \u2018acceptable\u2019 completions; they are recommendations whose predicted CTR places them near the upper end of each brand\u2019s historical range. Brand + 3 drops to roughly the 64.7th percentile, which is still above median, but much less distinctive. In other words, the model still finds plausible completions in this more constrained setting, but it is no longer consistently surfacing outcomes that look clearly elite for the brand.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"523\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-1024x523.png\" alt=\"\" class=\"wp-image-1953\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-1024x523.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-300x153.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-766x391.png 766w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr.png 1491w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 5. Predicted CTR percentile relative to each brand\u2019s historical distribution.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Taken together, the two plots support the same interpretation. The open-vocabulary recommender is most credible when it has enough flexibility to discover new completions rather than simply fill in one missing slot in an already highly specified setup. In broad and moderately constrained contexts, the recommended combinations both align strongly with the positive geometry of the learned space and score highly against the brand\u2019s own historical CTR range. In heavily constrained contexts, the model can still produce sensible outputs, but the diagnostic strength is noticeably weaker and the case for production use becomes more cautious.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Low known@10 is expected and desirable in this setting: it means the recommender is proposing new combinations rather than simply replaying historical rows. The overall known@10 rate is only 5.0%, confirming that most recommended combinations are genuinely new relative to the historical dataset.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overall, the strongest open-vocabulary evidence appears for brand-only, brand + 1 and brand + 2 queries. Brand + 3 degrades materially because the query is already highly constrained and leaves less room for the model to improve the completion.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">9. What This Enables<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">OI PoC turns historical campaign data into a planning tool with three practical uses:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>CTR prediction for unseen combinations of known campaign entities.<\/li>\n\n\n\n<li>Brand-relative recommendation, where \u201cpositive\u201d means strong for that brand rather than globally high CTR.<\/li>\n\n\n\n<li>Discovery of new campaign completions that are plausible in the learned performance-aware embedding space.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">The most defensible claim from the current results is:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">Among sufficiently supported queries, the model improves positive recommendation rate and sharply reduces negatives versus random selection, especially for brand-only, brand + 1 and brand + 2 recommendation contexts.<\/p>\n<\/blockquote>\n\n\n\n<h2 class=\"wp-block-heading\">What&#8217;s next?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">There are several directions we\u2019re excited to explore from here. One is understanding when the model is confident enough in a recommendation to make it useful, combining signals such as retrieval distance, predicted CTR and the diversity of the candidate pool.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We\u2019re also interested in making recommendations easier to interpret: which modalities contribute most to predicted CTR, and what makes a particular campaign completion rank highly?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Another direction is improving calibration across high-CTR segments and objectives such as <code>LINK_CLICKS<\/code>, while continuing to validate open-vocabulary recommendations with expert review.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ultimately, the most interesting question is how these recommendations perform beyond offline evaluation. A controlled online experiment comparing model-recommended campaign completions with existing planning approaches would be a natural next step.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">There\u2019s more to explore here \u2014 stay tuned.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Abstract Marketing teams often need to decide which campaign setup is worth testing before any budget is spent. The hard part is that performance does not depend on a single variable. It depends on the interaction between multiple factors, such as the brand, its audience, the campaign objective, the platform, the geography, and many other [&hellip;]<\/p>\n","protected":false},"author":35,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":"{\"authors\":[53],\"author_categories\":{\"53\":\"1\"},\"fallback_author_user\":\"4\",\"ppma_author_box_select\":\"\",\"selected_authors\":[{\"id\":53,\"display_name\":\"Andreas Mastakouris\",\"is_guest\":0,\"category_id\":\"1\"}]}"},"tags":[],"content_types":[{"id":51,"name":"Technical Report","slug":"technical-walkthrough"}],"ppma_author":[{"id":35,"display_name":"Andreas Mastakouris","first_name":"Andreas","last_name":"Mastakouris","nickname":"andreas.mastakouris","user_nicename":"andreas-mastakouris","user_email":"andreas.mastakouris@satalia.com","biographical_info":"Andreas is a Data Scientist at Open Intelligence, with a background in Mechanical Engineering and a M.Sc. in Data Science and Machine Learning. His work focuses on developing and deploying advanced machine learning solutions, with expertise spanning deep learning, generative AI, probabilistic modeling, and variational inference. He has experience applying state-of-the-art AI techniques to real-world industry challenges, including multimodal learning, foundation models, and large-scale predictive systems.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/image.png","job_title":"Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":null}],"class_list":["post-1760","research_feed","type-research_feed","status-publish","hentry","content_type-technical-walkthrough"],"acf":{"content":"<p><!-- wp:paragraph --><\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">Abstract<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Marketing teams often need to decide which campaign setup is worth testing before any budget is spent. The hard part is that performance does not depend on a single variable. It depends on the interaction between multiple factors, such as the brand, its audience, the campaign objective, the platform, the geography, and many other contextual signals.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This proof of concept from the Performance AI team at Open Intelligence (OI PoC) addresses this as a multimodal prediction and recommendation problem. We trained a single model that learns from historical campaign performance, predicts click-through rate (CTR), and uses the same learned embedding space to recommend campaign completions. The result is a system that can move from \u201cwhat happened before?\u201d to \u201cwhat should we try next?\u201d while staying grounded in <strong>real performance data<\/strong>.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This work builds on the earlier <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">Campaign Performance Modelling Pod<\/a>, which explored how AI can support campaign performance prediction using synthetic campaign data (see the <a href=\"https:\/\/research.wpp.com\/reports\/campaign-performance-modelling-pod-technical-walkthrough\">Technical Report<\/a> and <a href=\"https:\/\/research.wpp.com\/blog\/from-guesswork-to-foresight-how-ai-is-predicting-the-future-of-marketing-campaigns\">Blog Post<\/a>, respectively). The OI PoC analysis takes the next step by applying the same broader modelling agenda to <strong>observed\/real performance data<\/strong>, trying to also examine both tasks of CTR prediction and recommendation in order to answer: given a partially specified campaign setup, what completion is most likely to perform well?<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The prepared dataset contains 20,067 unique real campaign setups derived from 21,247,221 raw daily rows. The evaluation is designed as a warm-start compositional generalization problem: the model sees known modality values during training, but must generalize to unseen full combinations at test time. This is important because it mirrors the real planning use case more closely than simple row memorization.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>In our evaluation, the model reached a test RMSE of 0.0151, MAE of 0.0052 and R2 of 0.6001 on unseen campaign combinations. For closed-vocabulary brand-fixed recommendations, the system achieved 76.2% positive precision@10 versus 50.2% for random selection from the same candidate pools, while reducing negative recommendations from 18.1% to 2.9%. For open-vocabulary recommendation, the strongest evidence appears in broader query settings where the model can still identify geometrically coherent completions that rank high relative to each brand\u2019s historical CTR range.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">1. Motivation<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Campaign planning is highly combinatorial. A planner is not making one decision in isolation, but several interacting decisions at once: which audience to target, which campaign objective to optimize for, which platform to use, and which geography to prioritize for a given brand. Each of these choices can materially change the outcome, and the effect of any one decision depends on the others around it. In practice, that means the space of possible campaign setups grows much faster than the amount of historical evidence available to support them.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This follows naturally from the previous <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">Campaign Performance Modelling<\/a> work. The earlier synthetic-data analysis asked whether AI could help move campaign planning from intuition-led guesswork toward more systematic performance foresight. Here, we ask a complementary question: once real OI performance data is available, can we learn a performance-aware representation that supports both CTR prediction and recommendation?<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This is exactly the gap OI PoC addresses. We want to help answer two related questions before budget is committed:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:list {\"ordered\":true} --><\/p>\n<ol class=\"wp-block-list\"><!-- wp:list-item --><\/p>\n<li>If a team is considering a particular campaign setup, what CTR should they expect?<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>If part of the setup is already fixed, what completion is most likely to perform well?<\/li>\n<p><!-- \/wp:list-item --><\/ol>\n<p><!-- \/wp:list --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The challenge is that historical data is both rich and sparse at the same time. At the raw row level there is plenty of volume, but once performance is aggregated into distinct campaign setups, only a small fraction of the theoretically possible combinations has ever been observed. A useful model therefore cannot depend on exact lookup or memorization. It needs to learn reusable structure from known examples and then transfer that structure to unseen combinations.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This creates several practical modelling challenges:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:list {\"ordered\":true} --><\/p>\n<ol class=\"wp-block-list\"><!-- wp:list-item --><\/p>\n<li>Prediction models can accidentally memorize historical rows instead of learning reusable structure.<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Recommendation systems can look accurate when the available candidate pool is already easy, even if the model is not adding much ranking value.<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Performance labels are not globally comparable across brands, because some brands naturally operate at different CTR ranges than others.<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Recommendation quality depends not just on whether a setup looks broadly similar to past positives, but on whether its internal combination of brand, audience, objective, platform and geography is coherent.<\/li>\n<p><!-- \/wp:list-item --><\/ol>\n<p><!-- \/wp:list --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The goal of OI PoC was to build a model that could generalize across known campaign ingredients in new configurations. This is a warm-start compositional generalization problem: the model has usually seen the individual ingredients before, but not the exact recipe. That is why the project combines CTR prediction and recommendation in a shared multimodal framework. The prediction side estimates likely performance for unseen setups, while the recommendation side uses the same learned representation to identify strong or plausibly strong completions for partially specified queries.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">2. Dataset<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The raw source dataset contained 21,247,221 daily campaign rows. After aggregation by campaign setup and filtering low-signal rows, this reduced to 20,067 unique learnable combinations. Furthermore, brand values were anonymized before being used as modelling inputs. This preserves the ability to learn brand-specific performance behaviour while keeping the modelling dataset separated from directly identifiable brands.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The five campaign modalities and their coverage are summarized in Table 1:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Modality<\/th>\n<th>Coverage in the prepared dataset<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Brand<\/td>\n<td>75 anonymized brands<\/td>\n<\/tr>\n<tr>\n<td>Audience<\/td>\n<td>more than 1,000 audience representations<\/td>\n<\/tr>\n<tr>\n<td>Objective<\/td>\n<td>7 objectives<\/td>\n<\/tr>\n<tr>\n<td>Platform<\/td>\n<td>42 platform representations<\/td>\n<\/tr>\n<tr>\n<td>Geography<\/td>\n<td>167 geography representations<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 1. Campaign modalities and coverage in the prepared dataset.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The train, validation and test split used 17,056 training rows, 1,505 validation rows and 1,506 test rows. The test set was designed to avoid exact-combination leakage. <\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The split diagnostics are reported in Table 2:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Split diagnostic<\/th>\n<th>Result<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Exact 5-way test combinations seen in training<\/td>\n<td>0 \/ 1,506<\/td>\n<\/tr>\n<tr>\n<td>Truly new 5-way test combinations<\/td>\n<td>1,506 \/ 1,506<\/td>\n<\/tr>\n<tr>\n<td>Test brands seen in training<\/td>\n<td>71 \/ 71<\/td>\n<\/tr>\n<tr>\n<td>Test objectives seen in training<\/td>\n<td>7 \/ 7<\/td>\n<\/tr>\n<tr>\n<td>Test platforms seen in training<\/td>\n<td>39 \/ 39<\/td>\n<\/tr>\n<tr>\n<td>Test geographies seen in training<\/td>\n<td>117 \/ 117<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 2. Test-set leakage and warm-start diagnostics.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This means the evaluation is not a pure cold-start task. It is a test of whether the model can recombine known entities into unseen campaign setups.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">3. Why CTR Needed Brand-Relative Labels<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Click-through rate (CTR) serves as our example performance metric\u2014a practical starting point given the available data. We use it across the dataset, including campaigns with objectives other than clicks, to explore whether the approach can support prediction and recommendation. The positive, neutral and negative labels therefore describe brand-relative CTR, not success against each campaign\u2019s stated objective. Stronger click-through performance does not necessarily indicate stronger awareness, engagement or video-view outcomes. The next step is to extend the approach to available metrics better suited to those objectives.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>CTR distribution statistics are shown in Table 3, proving that CTR was highly skewed:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Statistic<\/th>\n<th>CTR<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Mean<\/td>\n<td>0.0106<\/td>\n<\/tr>\n<tr>\n<td>Median<\/td>\n<td>0.0027<\/td>\n<\/tr>\n<tr>\n<td>Skewness<\/td>\n<td>7.8710<\/td>\n<\/tr>\n<tr>\n<td>Kurtosis<\/td>\n<td>108.02<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 3. Summary statistics of the CTR distribution.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>A global definition of \u201cpositive\u201d campaign performance would over-favour brands with naturally higher CTR and penalise brands with lower typical CTR. Instead, each brand\u2019s own CTR distribution was split into rank-based thirds: negative, neutral and positive.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This created labels that mean \u201cgood for this brand\u201d, rather than \u201chigh CTR globally\u201d. The anonymized brand identifiers are still consistent across rows, which means the model can learn within-brand CTR patterns without needing access to the original brand names.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The final label distribution is reported in Table 4:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Label<\/th>\n<th>Rows<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Negative<\/td>\n<td>5,984<\/td>\n<\/tr>\n<tr>\n<td>Neutral<\/td>\n<td>8,030<\/td>\n<\/tr>\n<tr>\n<td>Positive<\/td>\n<td>6,053<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 4. Brand-relative CTR label distribution.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Rank-based splitting was used because many brands had zero-heavy CTR distributions. Standard quantile thresholds often collapsed to the same value, while rank-based labels preserved useful within-brand ordering.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">4. Model Approach<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>OI PoC uses one shared model for both prediction and recommendation.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Each modality starts from a 256-dimensional embedding and passes through its own projector:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:list --><\/p>\n<ul class=\"wp-block-list\"><!-- wp:list-item --><\/p>\n<li>brand projector<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>audience projector<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>objective projector<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>platform projector<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>geography projector<\/li>\n<p><!-- \/wp:list-item --><\/ul>\n<p><!-- \/wp:list --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The projected modality embeddings are pooled into a shared campaign representation. From there, the model optimizes two signals:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:list --><\/p>\n<ul class=\"wp-block-list\"><!-- wp:list-item --><\/p>\n<li>a regression head predicts transformed CTR;<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>a centered contrastive objective shapes the embedding geometry so that positive, neutral and negative campaign setups become easier to retrieve and rank.<\/li>\n<p><!-- \/wp:list-item --><\/ul>\n<p><!-- \/wp:list --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The shared model architecture is illustrated in Figure 1:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:image {\"id\":1771,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} --><\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"583\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1024x583.jpg\" alt=\"\" class=\"wp-image-1771\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1024x583.jpg 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-300x171.jpg 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-768x437.jpg 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1536x874.jpg 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2.jpg 1708w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 1. Shared multimodal architecture for CTR prediction and campaign recommendation.<\/figcaption><\/figure>\n<p><!-- \/wp:image --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The centered contrastive objective is important because it encodes a more useful recommendation bias than a standard label-only contrastive loss. In this setting, we do not simply want all positive examples to collapse toward one another in the embedding space. Two campaign setups can both be labelled positive while being strategically very different in terms of brand, audience, objective, platform and geography. Forcing all positives to become globally similar would risk bringing together irrelevant campaign setups that happen to share the same final label but not the same internal structure.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Instead, the centered loss focuses on intra-sample consistency. For each projected modality embedding, the model compares it against the center formed by the other modalities within the same campaign setup. This encourages the representation to answer a more precise question: do the parts of this setup belong together in a performance-aware way? Positive tuples should exhibit stronger internal agreement, neutral tuples should sit closer to an intermediate structure, and negative tuples should show weaker alignment. That is much closer to the actual recommendation task, where the model needs to judge whether a partially specified setup can be completed coherently rather than whether it merely resembles some broad global \u201cpositive\u201d cloud.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This design also makes the loss batch-sensitive in a useful way. The implementation weights modality contributions using both global and batch-level information. Global weighting reflects the overall cardinality of each modality vocabulary, while batch weighting reflects the effective diversity of modality values present in the current batch. In practice, this helps prevent the contrastive term from over-responding either to globally large modalities such as audience or to accidentally repetitive mini-batches where one modality has very little variation. The current implementation uses <code>gamma_global = 0.5<\/code> and <code>gamma_batch = 0.5<\/code>, which gives a balanced influence to long-run dataset structure and local batch composition.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The centered loss is also deliberately moderated relative to the CTR objective. In training, the total loss combines Huber regression with the centered contrastive term using <code>LAMBDA_CENTERED = 0.05<\/code>, plus a small norm penalty inside the contrastive component. This keeps CTR prediction as the dominant optimization target while still giving the embedding space enough geometric structure to support retrieval and reranking. The result is a representation that is not purely semantic and not purely regressive: it is shaped to preserve campaign-setup compatibility in a way that is directly useful for recommendation.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This design keeps deployment simple: one model artifact, one embedding space and one monitoring surface. It also makes the recommendation layer performance-aware rather than purely semantic.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">5. Evaluation Design<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The evaluation focused on three questions:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:list {\"ordered\":true} --><\/p>\n<ol class=\"wp-block-list\"><!-- wp:list-item --><\/p>\n<li>Can the model predict CTR for unseen campaign combinations?<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Can the learned representation identify stronger recommendations than random selection from the same candidate pool?<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Can the model propose new combinations in open vocabulary settings without simply replaying historical rows?<\/li>\n<p><!-- \/wp:list-item --><\/ol>\n<p><!-- \/wp:list --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>For recommendations, the system uses a two-stage retrieve-then-rerank pipeline, as summarized in Table 5:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Stage<\/th>\n<th>Role<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Geometry retrieval<\/td>\n<td>Find low-distance candidates in the learned embedding space<\/td>\n<\/tr>\n<tr>\n<td>CTR reranking<\/td>\n<td>Sort retrieved candidates by predicted CTR<\/td>\n<\/tr>\n<tr>\n<td>Top-k output<\/td>\n<td>Return the best campaign completions<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 5. Retrieve-then-rerank recommendation pipeline.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>This is important operationally because the embedding space provides a fast candidate filter, while the regression head gives the final ranking signal.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">6. Regression Results<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>As reported in Table 6, on the 1,506-row test set the model captured a meaningful share of CTR variance for unseen combinations:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Metric<\/th>\n<th>Test result<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>RMSE<\/td>\n<td>0.0151<\/td>\n<\/tr>\n<tr>\n<td>MAE<\/td>\n<td>0.0052<\/td>\n<\/tr>\n<tr>\n<td>R2<\/td>\n<td>0.6001<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 6. Regression performance on the test set.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The model is useful but conservative in high-CTR slices. For <code>LINK_CLICKS<\/code>, the actual average CTR was 0.02370 while the predicted average CTR was 0.01938, a bias of -0.00432. This is expected given the heavy-tailed target distribution and should be monitored before production use.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Permutation diagnostics are shown in Table 7, demonstrating that objective and platform were the strongest regression drivers:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Modality permuted<\/th>\n<th>Mean MAE increase<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Objective<\/td>\n<td>0.8490<\/td>\n<\/tr>\n<tr>\n<td>Platform<\/td>\n<td>0.6561<\/td>\n<\/tr>\n<tr>\n<td>Brand<\/td>\n<td>0.2018<\/td>\n<\/tr>\n<tr>\n<td>Audience<\/td>\n<td>0.1184<\/td>\n<\/tr>\n<tr>\n<td>Geography<\/td>\n<td>0.0342<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 7. Permutation importance by campaign modality.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The geography result is likely influenced by the dataset composition, where the prepared rows are heavily UK-centered.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">7. Closed-Vocabulary Recommendation Results<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Closed-vocabulary recommendations are completions drawn from rows already present in the historical dataset. This lets us evaluate recommendation quality against known labels.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>In practical terms, \u201cclosed\u201d means the recommender is choosing from campaign setups that have already been observed historically, rather than inventing entirely new combinations. That makes this the most directly measurable recommendation setting: every retrieved completion can be checked against an existing brand-relative label, so we know whether the model is surfacing historically positive, neutral or negative outcomes. Closed-vocabulary evaluation is therefore the clearest test of ranking quality. If the model is genuinely useful, it should place more positive combinations near the top of the list and push negative combinations down, even when the available candidate pool is mixed.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>We group recommendation queries by how much of the campaign setup is already fixed. \u201cBrand only\u201d fixes the brand and leaves audience, objective, platform and geography open for recommendation. \u201cBrand + 1\u201d fixes the brand plus one of those four ingredients; \u201cBrand + 2\u201d fixes the brand plus two; and \u201cBrand + 3\u201d fixes the brand plus three. For example, a query with a fixed brand and objective is a \u201cBrand + 1\u201d query: the model recommends the remaining audience, platform and geography.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Overall closed-vocabulary results are reported in Table 8. Across 57 valid evaluated brand-fixed queries, OI PoC achieved:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Metric<\/th>\n<th>Model<\/th>\n<th>Random from same candidate pools<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Positive precision@10<\/td>\n<td>76.2%<\/td>\n<td>50.2%<\/td>\n<\/tr>\n<tr>\n<td>Non-negative precision@10<\/td>\n<td>97.1%<\/td>\n<td>82.0%<\/td>\n<\/tr>\n<tr>\n<td>Negative rate@10<\/td>\n<td>2.9%<\/td>\n<td>18.0%<\/td>\n<\/tr>\n<tr>\n<td>Known@10<\/td>\n<td>100.0%<\/td>\n<td>n\/a<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 8. Closed-vocabulary recommendation performance compared with random selection.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The model produced a 1.92x positive lift versus random selection and reduced negative recommendations by 15.1 percentage points.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Results by query type are presented in Table 9. The strongest result came from broader brand-fixed queries:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Query type<\/th>\n<th>Positive@10<\/th>\n<th>Positive lift vs random<\/th>\n<th>Negative reduction vs random<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Brand only<\/td>\n<td>93.0%<\/td>\n<td>3.085x<\/td>\n<td>28.3pp<\/td>\n<\/tr>\n<tr>\n<td>Brand + 1 modality<\/td>\n<td>68.7%<\/td>\n<td>1.683x<\/td>\n<td>16.0pp<\/td>\n<\/tr>\n<tr>\n<td>Brand + 2 modalities<\/td>\n<td>84.3%<\/td>\n<td>1.939x<\/td>\n<td>10.4pp<\/td>\n<\/tr>\n<tr>\n<td>Brand + 3 modalities<\/td>\n<td>65.6%<\/td>\n<td>1.276x<\/td>\n<td>6.5pp<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 9. Closed-vocabulary performance by query type.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Figure 2 shows the positive recommendation rate by query type, while Figure 3 shows the corresponding negative recommendation rate. The former demonstrates how often the top recommendations fall into the historically positive class for each query type. This is the most intuitive measure of recommendation value: higher is better, because it means the model <em>succeeds<\/em> in surfacing combinations that were historically strong for the given brand. The latter depicts the opposite side of the same story. Lower is better, because it means the recommender is avoiding combinations that were historically weak.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:image {\"id\":1950,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} --><\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"558\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1024x558.png\" alt=\"\" class=\"wp-image-1950\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1024x558.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-300x163.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-767x418.png 767w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-2048x1116.png 2048w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1536x837.png 1536w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 2. Positive precision@10 by closed-vocabulary query type.<\/figcaption><\/figure>\n<p><!-- \/wp:image --><\/p>\n<p><!-- wp:image {\"id\":1951,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} --><\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"558\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1024x558.png\" alt=\"\" class=\"wp-image-1951\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1024x558.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-300x163.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-768x418.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1536x836.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-2048x1115.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 3. Negative rate@10 by closed-vocabulary query type.<\/figcaption><\/figure>\n<p><!-- \/wp:image --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>These plots are important because raw precision alone can be misleading if the candidate pool is already easy. That is why the comparison against random selection from the same candidate pool matters so much. The model is not just scoring well because it is choosing among already strong options. In brand-only queries, the historical candidate pool contains only about 30.2% positives, yet the recommender raises positive precision@10 to 93.0% and cuts negatives from 29.3% to 1.0%. Brand + 1 remains very strong, with 44.6% candidate positives rising to 68.7% positive@10 and negatives falling from 19.0% to 3.0%. Brand + 2 is still strong, lifting positives from 68.8% to 84.3% and lowering negatives from 14.7% to 4.3%.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Brand + 3 needs a more careful reading. On the surface, 65.6% positive@10 and 2.7% negative@10 still look solid. But the candidate pool for this case is already extremely favourable, with 57.0% positives and only 9.2% negatives before the model ranks anything. In other words, once four parts of the setup are already fixed, there is often very little ranking difficulty left. That is why the lift versus random is close to flat in this setting. The model is still producing sensible recommendations, but the closed-vocabulary evidence no longer shows a strong ranking advantage over the baseline.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The interpretation is therefore more nuanced than simply saying \u201cthe model works everywhere.\u201d Closed-vocabulary results show clear ranking value when the query leaves enough room for the model to discriminate among many possible completions. That is most visible in brand-only, brand + 1 and still present in brand + 2 queries, and much less convincing in brand + 3 where the space is already heavily constrained. This is exactly the kind of pattern we would hope to see from a recommender that is useful in realistic planning scenarios rather than only in highly filtered cases.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">8. Open-Vocabulary Recommendation Results<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Open-vocabulary recommendations generate combinations that are mostly not present in the historical dataset. This is the more interesting planning use case, but also the harder one to validate offline.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Because most open-vocabulary recommendations are unknown historically, the evaluation relies on geometric diagnostics and predicted CTR relative to each brand\u2019s historical distribution.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Two diagnostics are especially useful here. The first is centered cosine dissimilarity, which measures how well a recommended completion fits the learned geometry of a query. Lower dissimilarity means the retrieved combination sits closer to the model\u2019s positive reference structure; higher dissimilarity suggests a weaker or more ambiguous fit. The second is predicted CTR percentile versus brand history, which asks where the model\u2019s predicted CTR for a recommendation sits relative to that brand\u2019s own historical CTR distribution. A percentile near 1.0 means the recommendation is predicted to perform near the top end of what that brand has historically achieved, while a percentile closer to 0.5 means the recommendation looks more typical than exceptional.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Overall open-vocabulary diagnostics are summarized in Table 10:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:table --><\/p>\n<figure class=\"wp-block-table\">\n<table class=\"has-fixed-layout\">\n<thead>\n<tr>\n<th>Diagnostic<\/th>\n<th>Result<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Closest-to-positive distribution rate<\/td>\n<td>72.5%<\/td>\n<\/tr>\n<tr>\n<td>Closest-to-non-negative distribution rate<\/td>\n<td>90.0%<\/td>\n<\/tr>\n<tr>\n<td>Closest-to-negative distribution rate<\/td>\n<td>10.0%<\/td>\n<\/tr>\n<tr>\n<td>Average predicted CTR percentile vs brand history<\/td>\n<td>87.2%<\/td>\n<\/tr>\n<tr>\n<td>Known@10<\/td>\n<td>5.0%<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"wp-element-caption\">Table 10. Overall open-vocabulary recommendation diagnostics.<\/figcaption><\/figure>\n<p><!-- \/wp:table --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>As depicted in Figure 4, the centered cosine dissimilarity plot shows that the model is most confident in broad recommendation settings. Brand-only queries have the strongest geometric fit, followed by brand + 1 and brand + 2 queries. This is consistent with the retrieval results: when the query leaves enough freedom for the model to search the space, the learned geometry can still find combinations that resemble historically positive structures. Brand + 3 behaves differently, with substantially weaker positive-distribution alignment and a higher negative-distribution rate. This suggests that once the query becomes too constrained, the geometry has much less room to identify clearly superior completions.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:image {\"id\":1952,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} --><\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"517\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-1024x517.png\" alt=\"\" class=\"wp-image-1952\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-1024x517.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-300x152.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-768x388.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case.png 1506w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 4. Centered cosine dissimilarity by open-vocabulary query type.<\/figcaption><\/figure>\n<p><!-- \/wp:image --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The CTR percentile plot tells a complementary story (see Figure 5). For brand-only recommendations, the mean predicted CTR lands around the 98.1st percentile of the brand\u2019s historical distribution. Brand + 1 remains similarly strong at 94.9%, and brand + 2 still sits high at 91.2%. These are not just \u2018acceptable\u2019 completions; they are recommendations whose predicted CTR places them near the upper end of each brand\u2019s historical range. Brand + 3 drops to roughly the 64.7th percentile, which is still above median, but much less distinctive. In other words, the model still finds plausible completions in this more constrained setting, but it is no longer consistently surfacing outcomes that look clearly elite for the brand.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:image {\"id\":1953,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} --><\/p>\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"523\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-1024x523.png\" alt=\"\" class=\"wp-image-1953\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-1024x523.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-300x153.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-766x391.png 766w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr.png 1491w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 5. Predicted CTR percentile relative to each brand\u2019s historical distribution.<\/figcaption><\/figure>\n<p><!-- \/wp:image --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Taken together, the two plots support the same interpretation. The open-vocabulary recommender is most credible when it has enough flexibility to discover new completions rather than simply fill in one missing slot in an already highly specified setup. In broad and moderately constrained contexts, the recommended combinations both align strongly with the positive geometry of the learned space and score highly against the brand\u2019s own historical CTR range. In heavily constrained contexts, the model can still produce sensible outputs, but the diagnostic strength is noticeably weaker and the case for production use becomes more cautious.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Low known@10 is expected and desirable in this setting: it means the recommender is proposing new combinations rather than simply replaying historical rows. The overall known@10 rate is only 5.0%, confirming that most recommended combinations are genuinely new relative to the historical dataset.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Overall, the strongest open-vocabulary evidence appears for brand-only, brand + 1 and brand + 2 queries. Brand + 3 degrades materially because the query is already highly constrained and leaves less room for the model to improve the completion.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">9. What This Enables<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>OI PoC turns historical campaign data into a planning tool with three practical uses:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:list {\"ordered\":true} --><\/p>\n<ol class=\"wp-block-list\"><!-- wp:list-item --><\/p>\n<li>CTR prediction for unseen combinations of known campaign entities.<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Brand-relative recommendation, where \u201cpositive\u201d means strong for that brand rather than globally high CTR.<\/li>\n<p><!-- \/wp:list-item --><\/p>\n<p><!-- wp:list-item --><\/p>\n<li>Discovery of new campaign completions that are plausible in the learned performance-aware embedding space.<\/li>\n<p><!-- \/wp:list-item --><\/ol>\n<p><!-- \/wp:list --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>The most defensible claim from the current results is:<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:quote --><\/p>\n<blockquote class=\"wp-block-quote\"><p><!-- wp:paragraph --><\/p>\n<p>Among sufficiently supported queries, the model improves positive recommendation rate and sharply reduces negatives versus random selection, especially for brand-only, brand + 1 and brand + 2 recommendation contexts.<\/p>\n<p><!-- \/wp:paragraph --><\/p><\/blockquote>\n<p><!-- \/wp:quote --><\/p>\n<p><!-- wp:heading --><\/p>\n<h2 class=\"wp-block-heading\">What&#8217;s next?<\/h2>\n<p><!-- \/wp:heading --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>There are several directions we\u2019re excited to explore from here. One is understanding when the model is confident enough in a recommendation to make it useful, combining signals such as retrieval distance, predicted CTR and the diversity of the candidate pool.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>We\u2019re also interested in making recommendations easier to interpret: which modalities contribute most to predicted CTR, and what makes a particular campaign completion rank highly?<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Another direction is improving calibration across high-CTR segments and objectives such as <code>LINK_CLICKS<\/code>, while continuing to validate open-vocabulary recommendations with expert review.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>Ultimately, the most interesting question is how these recommendations perform beyond offline evaluation. A controlled online experiment comparing model-recommended campaign completions with existing planning approaches would be a natural next step.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n<p><!-- wp:paragraph --><\/p>\n<p>There\u2019s more to explore here \u2014 stay tuned.<\/p>\n<p><!-- \/wp:paragraph --><\/p>\n","content_quarter":"Q3 2026","related_pods":[1956]},"research_categories":[],"raw_acf":{"content":"<!-- wp:paragraph -->\n<p><\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">Abstract<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>Marketing teams often need to decide which campaign setup is worth testing before any budget is spent. The hard part is that performance does not depend on a single variable. It depends on the interaction between multiple factors, such as the brand, its audience, the campaign objective, the platform, the geography, and many other contextual signals.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This proof of concept from the Performance AI team at Open Intelligence (OI PoC) addresses this as a multimodal prediction and recommendation problem. We trained a single model that learns from historical campaign performance, predicts click-through rate (CTR), and uses the same learned embedding space to recommend campaign completions. The result is a system that can move from \u201cwhat happened before?\u201d to \u201cwhat should we try next?\u201d while staying grounded in <strong>real performance data<\/strong>.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This work builds on the earlier <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">Campaign Performance Modelling Pod<\/a>, which explored how AI can support campaign performance prediction using synthetic campaign data (see the <a href=\"https:\/\/research.wpp.com\/reports\/campaign-performance-modelling-pod-technical-walkthrough\">Technical Report<\/a> and <a href=\"https:\/\/research.wpp.com\/blog\/from-guesswork-to-foresight-how-ai-is-predicting-the-future-of-marketing-campaigns\">Blog Post<\/a>, respectively). The OI PoC analysis takes the next step by applying the same broader modelling agenda to <strong>observed\/real performance data<\/strong>, trying to also examine both tasks of CTR prediction and recommendation in order to answer: given a partially specified campaign setup, what completion is most likely to perform well?<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>The prepared dataset contains 20,067 unique real campaign setups derived from 21,247,221 raw daily rows. The evaluation is designed as a warm-start compositional generalization problem: the model sees known modality values during training, but must generalize to unseen full combinations at test time. This is important because it mirrors the real planning use case more closely than simple row memorization.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>In our evaluation, the model reached a test RMSE of 0.0151, MAE of 0.0052 and R2 of 0.6001 on unseen campaign combinations. For closed-vocabulary brand-fixed recommendations, the system achieved 76.2% positive precision@10 versus 50.2% for random selection from the same candidate pools, while reducing negative recommendations from 18.1% to 2.9%. For open-vocabulary recommendation, the strongest evidence appears in broader query settings where the model can still identify geometrically coherent completions that rank high relative to each brand\u2019s historical CTR range.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">1. Motivation<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>Campaign planning is highly combinatorial. A planner is not making one decision in isolation, but several interacting decisions at once: which audience to target, which campaign objective to optimize for, which platform to use, and which geography to prioritize for a given brand. Each of these choices can materially change the outcome, and the effect of any one decision depends on the others around it. In practice, that means the space of possible campaign setups grows much faster than the amount of historical evidence available to support them.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This follows naturally from the previous <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">Campaign Performance Modelling<\/a> work. The earlier synthetic-data analysis asked whether AI could help move campaign planning from intuition-led guesswork toward more systematic performance foresight. Here, we ask a complementary question: once real OI performance data is available, can we learn a performance-aware representation that supports both CTR prediction and recommendation?<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This is exactly the gap OI PoC addresses. We want to help answer two related questions before budget is committed:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:list {\"ordered\":true} -->\n<ol class=\"wp-block-list\"><!-- wp:list-item -->\n<li>If a team is considering a particular campaign setup, what CTR should they expect?<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>If part of the setup is already fixed, what completion is most likely to perform well?<\/li>\n<!-- \/wp:list-item --><\/ol>\n<!-- \/wp:list -->\n\n<!-- wp:paragraph -->\n<p>The challenge is that historical data is both rich and sparse at the same time. At the raw row level there is plenty of volume, but once performance is aggregated into distinct campaign setups, only a small fraction of the theoretically possible combinations has ever been observed. A useful model therefore cannot depend on exact lookup or memorization. It needs to learn reusable structure from known examples and then transfer that structure to unseen combinations.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This creates several practical modelling challenges:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:list {\"ordered\":true} -->\n<ol class=\"wp-block-list\"><!-- wp:list-item -->\n<li>Prediction models can accidentally memorize historical rows instead of learning reusable structure.<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Recommendation systems can look accurate when the available candidate pool is already easy, even if the model is not adding much ranking value.<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Performance labels are not globally comparable across brands, because some brands naturally operate at different CTR ranges than others.<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Recommendation quality depends not just on whether a setup looks broadly similar to past positives, but on whether its internal combination of brand, audience, objective, platform and geography is coherent.<\/li>\n<!-- \/wp:list-item --><\/ol>\n<!-- \/wp:list -->\n\n<!-- wp:paragraph -->\n<p>The goal of OI PoC was to build a model that could generalize across known campaign ingredients in new configurations. This is a warm-start compositional generalization problem: the model has usually seen the individual ingredients before, but not the exact recipe. That is why the project combines CTR prediction and recommendation in a shared multimodal framework. The prediction side estimates likely performance for unseen setups, while the recommendation side uses the same learned representation to identify strong or plausibly strong completions for partially specified queries.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">2. Dataset<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>The raw source dataset contained 21,247,221 daily campaign rows. After aggregation by campaign setup and filtering low-signal rows, this reduced to 20,067 unique learnable combinations. Furthermore, brand values were anonymized before being used as modelling inputs. This preserves the ability to learn brand-specific performance behaviour while keeping the modelling dataset separated from directly identifiable brands.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>The five campaign modalities and their coverage are summarized in Table 1:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Modality<\/th><th>Coverage in the prepared dataset<\/th><\/tr><\/thead><tbody><tr><td>Brand<\/td><td>75 anonymized brands<\/td><\/tr><tr><td>Audience<\/td><td>more than 1,000 audience representations<\/td><\/tr><tr><td>Objective<\/td><td>7 objectives<\/td><\/tr><tr><td>Platform<\/td><td>42 platform representations<\/td><\/tr><tr><td>Geography<\/td><td>167 geography representations<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 1. Campaign modalities and coverage in the prepared dataset.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>The train, validation and test split used 17,056 training rows, 1,505 validation rows and 1,506 test rows. The test set was designed to avoid exact-combination leakage. <\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>The split diagnostics are reported in Table 2:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Split diagnostic<\/th><th>Result<\/th><\/tr><\/thead><tbody><tr><td>Exact 5-way test combinations seen in training<\/td><td>0 \/ 1,506<\/td><\/tr><tr><td>Truly new 5-way test combinations<\/td><td>1,506 \/ 1,506<\/td><\/tr><tr><td>Test brands seen in training<\/td><td>71 \/ 71<\/td><\/tr><tr><td>Test objectives seen in training<\/td><td>7 \/ 7<\/td><\/tr><tr><td>Test platforms seen in training<\/td><td>39 \/ 39<\/td><\/tr><tr><td>Test geographies seen in training<\/td><td>117 \/ 117<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 2. Test-set leakage and warm-start diagnostics.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>This means the evaluation is not a pure cold-start task. It is a test of whether the model can recombine known entities into unseen campaign setups.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">3. Why CTR Needed Brand-Relative Labels<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>Click-through rate (CTR) serves as our example performance metric\u2014a practical starting point given the available data. We use it across the dataset, including campaigns with objectives other than clicks, to explore whether the approach can support prediction and recommendation. The positive, neutral and negative labels therefore describe brand-relative CTR, not success against each campaign\u2019s stated objective. Stronger click-through performance does not necessarily indicate stronger awareness, engagement or video-view outcomes. The next step is to extend the approach to available metrics better suited to those objectives.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>CTR distribution statistics are shown in Table 3, proving that CTR was highly skewed:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Statistic<\/th><th>CTR<\/th><\/tr><\/thead><tbody><tr><td>Mean<\/td><td>0.0106<\/td><\/tr><tr><td>Median<\/td><td>0.0027<\/td><\/tr><tr><td>Skewness<\/td><td>7.8710<\/td><\/tr><tr><td>Kurtosis<\/td><td>108.02<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 3. Summary statistics of the CTR distribution.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>A global definition of \u201cpositive\u201d campaign performance would over-favour brands with naturally higher CTR and penalise brands with lower typical CTR. Instead, each brand\u2019s own CTR distribution was split into rank-based thirds: negative, neutral and positive.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This created labels that mean \u201cgood for this brand\u201d, rather than \u201chigh CTR globally\u201d. The anonymized brand identifiers are still consistent across rows, which means the model can learn within-brand CTR patterns without needing access to the original brand names.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>The final label distribution is reported in Table 4:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Label<\/th><th>Rows<\/th><\/tr><\/thead><tbody><tr><td>Negative<\/td><td>5,984<\/td><\/tr><tr><td>Neutral<\/td><td>8,030<\/td><\/tr><tr><td>Positive<\/td><td>6,053<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 4. Brand-relative CTR label distribution.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>Rank-based splitting was used because many brands had zero-heavy CTR distributions. Standard quantile thresholds often collapsed to the same value, while rank-based labels preserved useful within-brand ordering.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">4. Model Approach<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>OI PoC uses one shared model for both prediction and recommendation.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Each modality starts from a 256-dimensional embedding and passes through its own projector:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:list -->\n<ul class=\"wp-block-list\"><!-- wp:list-item -->\n<li>brand projector<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>audience projector<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>objective projector<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>platform projector<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>geography projector<\/li>\n<!-- \/wp:list-item --><\/ul>\n<!-- \/wp:list -->\n\n<!-- wp:paragraph -->\n<p>The projected modality embeddings are pooled into a shared campaign representation. From there, the model optimizes two signals:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:list -->\n<ul class=\"wp-block-list\"><!-- wp:list-item -->\n<li>a regression head predicts transformed CTR;<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>a centered contrastive objective shapes the embedding geometry so that positive, neutral and negative campaign setups become easier to retrieve and rank.<\/li>\n<!-- \/wp:list-item --><\/ul>\n<!-- \/wp:list -->\n\n<!-- wp:paragraph -->\n<p>The shared model architecture is illustrated in Figure 1:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:image {\"id\":1771,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} -->\n<figure class=\"wp-block-image size-large\"><img src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/model-architecture-2-1024x583.jpg\" alt=\"\" class=\"wp-image-1771\"\/><figcaption class=\"wp-element-caption\">Figure 1. Shared multimodal architecture for CTR prediction and campaign recommendation.<\/figcaption><\/figure>\n<!-- \/wp:image -->\n\n<!-- wp:paragraph -->\n<p>The centered contrastive objective is important because it encodes a more useful recommendation bias than a standard label-only contrastive loss. In this setting, we do not simply want all positive examples to collapse toward one another in the embedding space. Two campaign setups can both be labelled positive while being strategically very different in terms of brand, audience, objective, platform and geography. Forcing all positives to become globally similar would risk bringing together irrelevant campaign setups that happen to share the same final label but not the same internal structure.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Instead, the centered loss focuses on intra-sample consistency. For each projected modality embedding, the model compares it against the center formed by the other modalities within the same campaign setup. This encourages the representation to answer a more precise question: do the parts of this setup belong together in a performance-aware way? Positive tuples should exhibit stronger internal agreement, neutral tuples should sit closer to an intermediate structure, and negative tuples should show weaker alignment. That is much closer to the actual recommendation task, where the model needs to judge whether a partially specified setup can be completed coherently rather than whether it merely resembles some broad global \u201cpositive\u201d cloud.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This design also makes the loss batch-sensitive in a useful way. The implementation weights modality contributions using both global and batch-level information. Global weighting reflects the overall cardinality of each modality vocabulary, while batch weighting reflects the effective diversity of modality values present in the current batch. In practice, this helps prevent the contrastive term from over-responding either to globally large modalities such as audience or to accidentally repetitive mini-batches where one modality has very little variation. The current implementation uses <code>gamma_global = 0.5<\/code> and <code>gamma_batch = 0.5<\/code>, which gives a balanced influence to long-run dataset structure and local batch composition.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>The centered loss is also deliberately moderated relative to the CTR objective. In training, the total loss combines Huber regression with the centered contrastive term using <code>LAMBDA_CENTERED = 0.05<\/code>, plus a small norm penalty inside the contrastive component. This keeps CTR prediction as the dominant optimization target while still giving the embedding space enough geometric structure to support retrieval and reranking. The result is a representation that is not purely semantic and not purely regressive: it is shaped to preserve campaign-setup compatibility in a way that is directly useful for recommendation.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>This design keeps deployment simple: one model artifact, one embedding space and one monitoring surface. It also makes the recommendation layer performance-aware rather than purely semantic.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">5. Evaluation Design<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>The evaluation focused on three questions:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:list {\"ordered\":true} -->\n<ol class=\"wp-block-list\"><!-- wp:list-item -->\n<li>Can the model predict CTR for unseen campaign combinations?<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Can the learned representation identify stronger recommendations than random selection from the same candidate pool?<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Can the model propose new combinations in open vocabulary settings without simply replaying historical rows?<\/li>\n<!-- \/wp:list-item --><\/ol>\n<!-- \/wp:list -->\n\n<!-- wp:paragraph -->\n<p>For recommendations, the system uses a two-stage retrieve-then-rerank pipeline, as summarized in Table 5:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Stage<\/th><th>Role<\/th><\/tr><\/thead><tbody><tr><td>Geometry retrieval<\/td><td>Find low-distance candidates in the learned embedding space<\/td><\/tr><tr><td>CTR reranking<\/td><td>Sort retrieved candidates by predicted CTR<\/td><\/tr><tr><td>Top-k output<\/td><td>Return the best campaign completions<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 5. Retrieve-then-rerank recommendation pipeline.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>This is important operationally because the embedding space provides a fast candidate filter, while the regression head gives the final ranking signal.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">6. Regression Results<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>As reported in Table 6, on the 1,506-row test set the model captured a meaningful share of CTR variance for unseen combinations:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Metric<\/th><th>Test result<\/th><\/tr><\/thead><tbody><tr><td>RMSE<\/td><td>0.0151<\/td><\/tr><tr><td>MAE<\/td><td>0.0052<\/td><\/tr><tr><td>R2<\/td><td>0.6001<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 6. Regression performance on the test set.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>The model is useful but conservative in high-CTR slices. For <code>LINK_CLICKS<\/code>, the actual average CTR was 0.02370 while the predicted average CTR was 0.01938, a bias of -0.00432. This is expected given the heavy-tailed target distribution and should be monitored before production use.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Permutation diagnostics are shown in Table 7, demonstrating that objective and platform were the strongest regression drivers:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Modality permuted<\/th><th>Mean MAE increase<\/th><\/tr><\/thead><tbody><tr><td>Objective<\/td><td>0.8490<\/td><\/tr><tr><td>Platform<\/td><td>0.6561<\/td><\/tr><tr><td>Brand<\/td><td>0.2018<\/td><\/tr><tr><td>Audience<\/td><td>0.1184<\/td><\/tr><tr><td>Geography<\/td><td>0.0342<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 7. Permutation importance by campaign modality.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>The geography result is likely influenced by the dataset composition, where the prepared rows are heavily UK-centered.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">7. Closed-Vocabulary Recommendation Results<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>Closed-vocabulary recommendations are completions drawn from rows already present in the historical dataset. This lets us evaluate recommendation quality against known labels.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>In practical terms, \u201cclosed\u201d means the recommender is choosing from campaign setups that have already been observed historically, rather than inventing entirely new combinations. That makes this the most directly measurable recommendation setting: every retrieved completion can be checked against an existing brand-relative label, so we know whether the model is surfacing historically positive, neutral or negative outcomes. Closed-vocabulary evaluation is therefore the clearest test of ranking quality. If the model is genuinely useful, it should place more positive combinations near the top of the list and push negative combinations down, even when the available candidate pool is mixed.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>We group recommendation queries by how much of the campaign setup is already fixed. \u201cBrand only\u201d fixes the brand and leaves audience, objective, platform and geography open for recommendation. \u201cBrand + 1\u201d fixes the brand plus one of those four ingredients; \u201cBrand + 2\u201d fixes the brand plus two; and \u201cBrand + 3\u201d fixes the brand plus three. For example, a query with a fixed brand and objective is a \u201cBrand + 1\u201d query: the model recommends the remaining audience, platform and geography.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Overall closed-vocabulary results are reported in Table 8. Across 57 valid evaluated brand-fixed queries, OI PoC achieved:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Metric<\/th><th>Model<\/th><th>Random from same candidate pools<\/th><\/tr><\/thead><tbody><tr><td>Positive precision@10<\/td><td>76.2%<\/td><td>50.2%<\/td><\/tr><tr><td>Non-negative precision@10<\/td><td>97.1%<\/td><td>82.0%<\/td><\/tr><tr><td>Negative rate@10<\/td><td>2.9%<\/td><td>18.0%<\/td><\/tr><tr><td>Known@10<\/td><td>100.0%<\/td><td>n\/a<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 8. Closed-vocabulary recommendation performance compared with random selection.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>The model produced a 1.92x positive lift versus random selection and reduced negative recommendations by 15.1 percentage points.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Results by query type are presented in Table 9. The strongest result came from broader brand-fixed queries:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Query type<\/th><th>Positive@10<\/th><th>Positive lift vs random<\/th><th>Negative reduction vs random<\/th><\/tr><\/thead><tbody><tr><td>Brand only<\/td><td>93.0%<\/td><td>3.085x<\/td><td>28.3pp<\/td><\/tr><tr><td>Brand + 1 modality<\/td><td>68.7%<\/td><td>1.683x<\/td><td>16.0pp<\/td><\/tr><tr><td>Brand + 2 modalities<\/td><td>84.3%<\/td><td>1.939x<\/td><td>10.4pp<\/td><\/tr><tr><td>Brand + 3 modalities<\/td><td>65.6%<\/td><td>1.276x<\/td><td>6.5pp<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 9. Closed-vocabulary performance by query type.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>Figure 2 shows the positive recommendation rate by query type, while Figure 3 shows the corresponding negative recommendation rate. The former demonstrates how often the top recommendations fall into the historically positive class for each query type. This is the most intuitive measure of recommendation value: higher is better, because it means the model <em>succeeds<\/em> in surfacing combinations that were historically strong for the given brand. The latter depicts the opposite side of the same story. Lower is better, because it means the recommender is avoiding combinations that were historically weak.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:image {\"id\":1950,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} -->\n<figure class=\"wp-block-image size-large\"><img src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_positive_vs_random-1024x558.png\" alt=\"\" class=\"wp-image-1950\"\/><figcaption class=\"wp-element-caption\">Figure 2. Positive precision@10 by closed-vocabulary query type.<\/figcaption><\/figure>\n<!-- \/wp:image -->\n\n<!-- wp:image {\"id\":1951,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} -->\n<figure class=\"wp-block-image size-large\"><img src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/closed_vocab_summary_brand_negative_vs_random-1024x558.png\" alt=\"\" class=\"wp-image-1951\"\/><figcaption class=\"wp-element-caption\">Figure 3. Negative rate@10 by closed-vocabulary query type.<\/figcaption><\/figure>\n<!-- \/wp:image -->\n\n<!-- wp:paragraph -->\n<p>These plots are important because raw precision alone can be misleading if the candidate pool is already easy. That is why the comparison against random selection from the same candidate pool matters so much. The model is not just scoring well because it is choosing among already strong options. In brand-only queries, the historical candidate pool contains only about 30.2% positives, yet the recommender raises positive precision@10 to 93.0% and cuts negatives from 29.3% to 1.0%. Brand + 1 remains very strong, with 44.6% candidate positives rising to 68.7% positive@10 and negatives falling from 19.0% to 3.0%. Brand + 2 is still strong, lifting positives from 68.8% to 84.3% and lowering negatives from 14.7% to 4.3%.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Brand + 3 needs a more careful reading. On the surface, 65.6% positive@10 and 2.7% negative@10 still look solid. But the candidate pool for this case is already extremely favourable, with 57.0% positives and only 9.2% negatives before the model ranks anything. In other words, once four parts of the setup are already fixed, there is often very little ranking difficulty left. That is why the lift versus random is close to flat in this setting. The model is still producing sensible recommendations, but the closed-vocabulary evidence no longer shows a strong ranking advantage over the baseline.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>The interpretation is therefore more nuanced than simply saying \u201cthe model works everywhere.\u201d Closed-vocabulary results show clear ranking value when the query leaves enough room for the model to discriminate among many possible completions. That is most visible in brand-only, brand + 1 and still present in brand + 2 queries, and much less convincing in brand + 3 where the space is already heavily constrained. This is exactly the kind of pattern we would hope to see from a recommender that is useful in realistic planning scenarios rather than only in highly filtered cases.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">8. Open-Vocabulary Recommendation Results<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>Open-vocabulary recommendations generate combinations that are mostly not present in the historical dataset. This is the more interesting planning use case, but also the harder one to validate offline.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Because most open-vocabulary recommendations are unknown historically, the evaluation relies on geometric diagnostics and predicted CTR relative to each brand\u2019s historical distribution.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Two diagnostics are especially useful here. The first is centered cosine dissimilarity, which measures how well a recommended completion fits the learned geometry of a query. Lower dissimilarity means the retrieved combination sits closer to the model\u2019s positive reference structure; higher dissimilarity suggests a weaker or more ambiguous fit. The second is predicted CTR percentile versus brand history, which asks where the model\u2019s predicted CTR for a recommendation sits relative to that brand\u2019s own historical CTR distribution. A percentile near 1.0 means the recommendation is predicted to perform near the top end of what that brand has historically achieved, while a percentile closer to 0.5 means the recommendation looks more typical than exceptional.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Overall open-vocabulary diagnostics are summarized in Table 10:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:table -->\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Diagnostic<\/th><th>Result<\/th><\/tr><\/thead><tbody><tr><td>Closest-to-positive distribution rate<\/td><td>72.5%<\/td><\/tr><tr><td>Closest-to-non-negative distribution rate<\/td><td>90.0%<\/td><\/tr><tr><td>Closest-to-negative distribution rate<\/td><td>10.0%<\/td><\/tr><tr><td>Average predicted CTR percentile vs brand history<\/td><td>87.2%<\/td><\/tr><tr><td>Known@10<\/td><td>5.0%<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Table 10. Overall open-vocabulary recommendation diagnostics.<\/figcaption><\/figure>\n<!-- \/wp:table -->\n\n<!-- wp:paragraph -->\n<p>As depicted in Figure 4, the centered cosine dissimilarity plot shows that the model is most confident in broad recommendation settings. Brand-only queries have the strongest geometric fit, followed by brand + 1 and brand + 2 queries. This is consistent with the retrieval results: when the query leaves enough freedom for the model to search the space, the learned geometry can still find combinations that resemble historically positive structures. Brand + 3 behaves differently, with substantially weaker positive-distribution alignment and a higher negative-distribution rate. This suggests that once the query becomes too constrained, the geometry has much less room to identify clearly superior completions.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:image {\"id\":1952,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} -->\n<figure class=\"wp-block-image size-large\"><img src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/distance_distribution_by_case-1024x517.png\" alt=\"\" class=\"wp-image-1952\"\/><figcaption class=\"wp-element-caption\">Figure 4. Centered cosine dissimilarity by open-vocabulary query type.<\/figcaption><\/figure>\n<!-- \/wp:image -->\n\n<!-- wp:paragraph -->\n<p>The CTR percentile plot tells a complementary story (see Figure 5). For brand-only recommendations, the mean predicted CTR lands around the 98.1st percentile of the brand\u2019s historical distribution. Brand + 1 remains similarly strong at 94.9%, and brand + 2 still sits high at 91.2%. These are not just \u2018acceptable\u2019 completions; they are recommendations whose predicted CTR places them near the upper end of each brand\u2019s historical range. Brand + 3 drops to roughly the 64.7th percentile, which is still above median, but much less distinctive. In other words, the model still finds plausible completions in this more constrained setting, but it is no longer consistently surfacing outcomes that look clearly elite for the brand.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:image {\"id\":1953,\"sizeSlug\":\"large\",\"linkDestination\":\"none\"} -->\n<figure class=\"wp-block-image size-large\"><img src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/09\/predicted_vs_brand_ctr-1024x523.png\" alt=\"\" class=\"wp-image-1953\"\/><figcaption class=\"wp-element-caption\">Figure 5. Predicted CTR percentile relative to each brand\u2019s historical distribution.<\/figcaption><\/figure>\n<!-- \/wp:image -->\n\n<!-- wp:paragraph -->\n<p>Taken together, the two plots support the same interpretation. The open-vocabulary recommender is most credible when it has enough flexibility to discover new completions rather than simply fill in one missing slot in an already highly specified setup. In broad and moderately constrained contexts, the recommended combinations both align strongly with the positive geometry of the learned space and score highly against the brand\u2019s own historical CTR range. In heavily constrained contexts, the model can still produce sensible outputs, but the diagnostic strength is noticeably weaker and the case for production use becomes more cautious.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Low known@10 is expected and desirable in this setting: it means the recommender is proposing new combinations rather than simply replaying historical rows. The overall known@10 rate is only 5.0%, confirming that most recommended combinations are genuinely new relative to the historical dataset.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Overall, the strongest open-vocabulary evidence appears for brand-only, brand + 1 and brand + 2 queries. Brand + 3 degrades materially because the query is already highly constrained and leaves less room for the model to improve the completion.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">9. What This Enables<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>OI PoC turns historical campaign data into a planning tool with three practical uses:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:list {\"ordered\":true} -->\n<ol class=\"wp-block-list\"><!-- wp:list-item -->\n<li>CTR prediction for unseen combinations of known campaign entities.<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Brand-relative recommendation, where \u201cpositive\u201d means strong for that brand rather than globally high CTR.<\/li>\n<!-- \/wp:list-item -->\n\n<!-- wp:list-item -->\n<li>Discovery of new campaign completions that are plausible in the learned performance-aware embedding space.<\/li>\n<!-- \/wp:list-item --><\/ol>\n<!-- \/wp:list -->\n\n<!-- wp:paragraph -->\n<p>The most defensible claim from the current results is:<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:quote -->\n<blockquote class=\"wp-block-quote\"><!-- wp:paragraph -->\n<p>Among sufficiently supported queries, the model improves positive recommendation rate and sharply reduces negatives versus random selection, especially for brand-only, brand + 1 and brand + 2 recommendation contexts.<\/p>\n<!-- \/wp:paragraph --><\/blockquote>\n<!-- \/wp:quote -->\n\n<!-- wp:heading -->\n<h2 class=\"wp-block-heading\">What's next?<\/h2>\n<!-- \/wp:heading -->\n\n<!-- wp:paragraph -->\n<p>There are several directions we\u2019re excited to explore from here. One is understanding when the model is confident enough in a recommendation to make it useful, combining signals such as retrieval distance, predicted CTR and the diversity of the candidate pool.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>We\u2019re also interested in making recommendations easier to interpret: which modalities contribute most to predicted CTR, and what makes a particular campaign completion rank highly?<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Another direction is improving calibration across high-CTR segments and objectives such as <code>LINK_CLICKS<\/code>, while continuing to validate open-vocabulary recommendations with expert review.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>Ultimately, the most interesting question is how these recommendations perform beyond offline evaluation. A controlled online experiment comparing model-recommended campaign completions with existing planning approaches would be a natural next step.<\/p>\n<!-- \/wp:paragraph -->\n\n<!-- wp:paragraph -->\n<p>There\u2019s more to explore here \u2014 stay tuned.<\/p>\n<!-- \/wp:paragraph -->","content_quarter":"Q3 2026","related_pods":["1956"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/1760","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/35"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/1956"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1760"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1760"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=1760"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=1760"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}