{"id":1789,"date":"2026-07-01T06:07:39","date_gmt":"2026-07-01T06:07:39","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=1789"},"modified":"2026-07-02T09:20:02","modified_gmt":"2026-07-02T09:20:02","slug":"federated-learning-meets-llm-training","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=federated-learning-meets-llm-training","title":{"rendered":"Federated LLM Fine-Tuning for Marketing: Collective Intelligence from Sovereign Data"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Every click, scroll, and purchase leaves a trace. The more of these traces a Machine Learning (ML) model can study, the sharper its predictions and the stronger its commercial returns. The challenge is that no single organisation holds the full picture. The data is split across agencies, brands, and technology partners, each sitting on their own piece of the puzzle.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In our earlier <a href=\"https:\/\/research.wpp.com\/blog\/training-together-sharing-nothing-the-promise-of-federated-learning\">work<\/a>, we laid the groundwork for a federated way for organisations to collaboratively train a shared ML model for media performance, without having to centralize all their data puzzle pieces in one place. Each participant trains locally within their own infrastructure, and exchanges only model updates (weights) rather than raw data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We then tested this idea with a set of experiments on simple ML models, comparing a centralised baseline (pooling everyone\u2019s data in the same place and training a model on the full puzzle) with a federated counterpart that keeps the puzzle distributed across several data nodes. Alongside this, we examined how fragmentation to an increasing number of participating nodes affects overall performance.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Those initial experiments rested on two assumptions:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>All the data puzzle pieces look alike. Conceptually, this assumes that every data partner observes and records similar patterns and behaviors. This is a simplification of the real-world, where brands, agencies, and platforms address a diverse spectrum of audiences, products, and experiences.<\/li>\n\n\n\n<li>The federated ML model follows a traditional \u201ctabular\u201d architecture. Every row in its input is an observation (e.g. a marketing campaign), every column is a an aspect of that observation (e.g. the targeted audience, the advertised product). The last column is the outcome we are trying to model (e.g. the campaign\u2019s performance). The model then looks for statistical associations among the columns from scratch, without having its own prior knowledge of the world. Again, this is a simplification of today\u2019s AI\/ML model landscape, where LLMs with vast prior world knowledge dominate.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">In this next phase of our work, we relax both of these assumptions by exploring:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Federated Learning in the presence of nodes with heterogeneous datasets<\/li>\n\n\n\n<li>Federated Fine-tuning of LLM-based predictive models<\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Contextualising the data: Marketing campaigns<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Before turning to the experiments, it helps to revisit what the marketing dataset represents and where it comes from. We use synthetic datasets generated by a purpose-built <a href=\"https:\/\/research.wpp.com\/blog\/using-synthetic-data-to-train-and-stress-test-marketing-ml-models\">pipeline<\/a> designed to mirror the dynamics of real marketing campaigns. This approach gives us fine-grained control over each dataset\u2019s composition, allowing us to construct targeted scenarios that stress-test the model under controlled conditions while retaining full visibility into the factors that drive campaign outcomes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At the heart of this pipeline is a signed marketing knowledge graph, which encodes how marketing attributes interact. Each node is an attribute value\u2014the <strong>brand<\/strong> being promoted, the <strong>audience<\/strong> it targets, the content\u2019s <strong>creative<\/strong> tone and style, the <strong>platform<\/strong> it runs on, or the <strong>geography<\/strong> it reaches\u2014and each edge captures whether two attributes are compatible (positive signal), incompatible (negative signal), or unrelated (missing signal).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This graph is the ground truth from which we sample subgraphs to build various synthetic datasets. Each dataset contains campaigns defined as combinations of these attributes: one brand, one audience, one creative (image or video), one platform, and one geographical location, drawn from the graph. Labels follow from the signals linking them: campaigns built from compatible pairings tend to over-perform (<strong>Positive<\/strong>), those from incompatible ones tend to under-perform (<strong>Negative<\/strong>), and mixed cases fall in between (<strong>Average<\/strong>). For full details on how the graph is built and sampled, see the <a href=\"https:\/\/research.wpp.com\/pods\/synthetic-dataset-generation-pod\">synthetic data generator pod<\/a>.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"708\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/graph-4-1024x708.png\" alt=\"\" class=\"wp-image-1791\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/graph-4-1024x708.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/graph-4-300x207.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/graph-4-768x531.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/graph-4.png 1433w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 1:<\/strong> Visualisation of two synthetic marketing subgraphs for the Base and 10% Flip datasets. \u201cFlipping\u201d means changing the edge type of every edge incident to a randomly selected 10% of nodes (highlighted by the dashed rectangle).<\/figcaption><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">Generating heterogeneous data<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For our experiments, we utilised the ground-truth graph to extract four subgraphs for four respective datasets. To meet the requirement of using heterogeneous datasets, we deliberately add structural noise by randomly \u201cflipping\u201d the connected edges of a subset of subgraph nodes from positive to negative and vice versa. Figure 1 illustrates this on a sample subgraph: the left-hand panel (\u201cBase Subgraph\u201d) shows it in its original state. The right-hand panel (\u201c10% Flipped Subgraph\u201d) then shows a flip in action: we randomly select a fraction of nodes (here <em>Brand A<\/em>, boxed by the dashed rectangle) and reverse the signal on every edge connected to them, turning positive (green) edges negative (red) and vice versa. The result is a graph that tells a partly contradictory story, simulating the noisy, conflicting evidence of real marketing data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The dataset names follow directly the fraction of nodes flipped, represented as percentage in each name. More specifically:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Base<\/strong> (0%) leaves the graph untouched as a clean reference<\/li>\n\n\n\n<li><strong>10% Flip<\/strong> reverses a random 10% of nodes (the case in Figure 1)<\/li>\n\n\n\n<li><strong>20% Flip<\/strong> and <strong>50% Flip<\/strong> apply the same step to larger fractions, adding steadily more noise<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Since every dataset shares one origin and differs only in flip rate, this provides a controlled way to observe how the model copes as conflicting evidence accumulates\u2014each dataset large enough to reach stable performance even without training to full equilibrium.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>From tables to text: Federated LLM fine\u2011tuning with LoRA<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The synthetic datasets start out as tables, where each row is a single campaign described by its features (brand, geography, audience, platform, and creative), and the last column records how that campaign performed\u2014Positive, Negative, or Average. To prepare the data, we first replace these three labels with simple number codes (0, 1, or 2). We then treat the problem as a text classification task, where the model&#8217;s job is to read a campaign and produce a single answer at the end indicating its predicted performance.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Before training, every dataset goes through the same preparation steps, so all the inputs look consistent. Each campaign row is rewritten as a short, plain-language description that the LLM model can read, with each feature clearly labelled. For example, one campaign might look like this:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Platform: Platform A\nBrand: Premium outdoor apparel brand\nAudience: Young professionals interested in fitness and travel\nCreative: Humorous\nGeo: &#91;1023, 2045, 9876] # Zip codes expression of the place\n<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This description is then placed inside a fixed set of instructions that presents the task in exactly the same way for every experiment. During training, we include the correct performance label so the model can learn the patterns. During testing, we leave the label out, and the model has to work it out on its own\u2014deciding whether the campaign is Positive, Negative, or Average.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1011\" height=\"612\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/new_graph.png\" alt=\"\" class=\"wp-image-1821\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/new_graph.png 1011w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/new_graph-300x182.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/new_graph-768x465.png 768w\" sizes=\"auto, (max-width: 1011px) 100vw, 1011px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 2:<\/strong> Visualization of the Federated Learning setup using FlowerTune with LoRA. Black numbered dots depict the steps of the FL training process (described below).<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">We recast the marketing problem as a task an LLM can solve, enabling us to fine-tune an LLM in a federated setup using FlowerTune with LoRA (Low-Rank Adaptation). We chose this approach because marketing data is typically sensitive and siloed across data nodes, and a&nbsp;<em>federated<\/em>&nbsp;setup keeps that data private: the model is trained locally at each node, so only the resulting updates\u2014never the raw data\u2014are ever shared. <a href=\"https:\/\/arxiv.org\/pdf\/2506.02961\">FlowerTune<\/a>, the LLM fine-tuning tool in the Flower federated-learning framework, coordinates this process, while LoRA makes it practical: rather than exchanging billions of parameters each round, LoRA freezes the base model and trains only a small set of adapter weights, reducing the trainable parameters by over 99% and keeping communication cheap. As illustrated in Figure 2, a single communication round then proceeds in four steps:&nbsp;<strong>(1) Local training:<\/strong>&nbsp;each data node freezes the massive pre-trained base model and updates only the small set of LoRA adapters on its private, siloed dataset, so the raw data never leaves the node;&nbsp;<strong>(2) Upload updates:<\/strong>&nbsp;each node sends only these lightweight adapters\u2014not the raw data or the full model\u2014to the central server;&nbsp;<strong>(3) Aggregation:<\/strong>&nbsp;the server combines the adapters from all participating nodes into a single set of aggregated LoRA adapters, producing one shared, improved model; and&nbsp;<strong>(4) Broadcast:<\/strong>&nbsp;the server distributes these aggregated adapters back to every node to serve as the starting point for the next round. This cycle repeats over successive communication rounds until the model converges.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Experiments &amp; lessons learned<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For our experimental setup, we evaluate our text classification task using two LLMs of increasing size:&nbsp;<strong><code>Gemma-3-1b-it<\/code><\/strong>&nbsp;and&nbsp;<strong><code>Llama-3-8B-Instruct<\/code><\/strong>&nbsp;(1 billion and 8 billion parameters, respectively). Working with two models of different scale lets us assess not only the centralised-versus-federated gap, but also whether model capacity influences how gracefully each setup absorbs heterogeneous data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For each model, we compare a <strong>centrally-trained<\/strong>&nbsp;LLM against its&nbsp;<strong>federated-trained<\/strong>&nbsp;counterpart across three progressively more heterogeneous settings:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Setting 1 (2 Data nodes):<\/strong>&nbsp;The centralised model is trained on the combined samples of the&nbsp;<em>Base<\/em>&nbsp;and&nbsp;<em>10% Flip<\/em>&nbsp;datasets. The federated equivalent uses two data nodes, one holding&nbsp;<em>Base<\/em>&nbsp;and the other&nbsp;<em>10% Flip<\/em>.<\/li>\n\n\n\n<li><strong>Setting 2 (3 Data nodes):<\/strong>&nbsp;The&nbsp;<em>20% Flip<\/em>&nbsp;dataset is added, to the centralised training pool, and as an additional federated data node.<\/li>\n\n\n\n<li><strong>Setting 3 (4 Data nodes):<\/strong>&nbsp;The&nbsp;<em>50% Flip<\/em>&nbsp;dataset is added in the same manner, giving the most heterogeneous configuration.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">All models are evaluated on a single, universal test set containing equal numbers of samples from every distribution (<em>Base<\/em>, <em>10%<\/em>, <em>20%<\/em>, and <em>50% Flip<\/em>). This common benchmark provides a consistent view of how well each trained model generalises across the full spread of perspectives present in the federation. Each centralised scenario was trained for two passes over its own data; similarly each federated experiment used two communication cycles: in each cycle, every data node trained the model locally for one pass over its own data, then sent only the resulting model update to the server, which combined the updates before starting the next and final cycle, keeping total training exposure comparable between the two regimes. All other training parameters remain the same across the centralized version and its federated counterpart.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">All models are evaluated on a single, universal test set containing equal numbers of samples from every distribution (<em>Base<\/em>, <em>10%<\/em>, <em>20%<\/em>, and <em>50% Flip<\/em>). This common benchmark provides a consistent view of how well each trained model generalises across the full spread of perspectives present in the federation. Each centralised scenario was trained for two passes over its own data; similarly each federated experiment used two communication cycles: in each cycle, every data node trained the model locally for one pass over its own data, then sent only the resulting model update to the server, which combined the updates before starting the next and final cycle, keeping total training exposure comparable between the two regimes. All other training parameters remain the same across the centralized version and its federated counterpart.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Centralised vs. federated: The cost of heterogeneity<\/h3>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"586\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/overall_f1_comparison-1024x586.png\" alt=\"\" class=\"wp-image-1816\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/overall_f1_comparison-1024x586.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/overall_f1_comparison-300x172.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/overall_f1_comparison-768x439.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/overall_f1_comparison-1536x878.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/overall_f1_comparison-2048x1171.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 3:<\/strong> F1 score comparison among the centralized vs FL versions under the 3 different experimental settings between the <code>Gemma-3-1b-it<\/code> and <code>Llama-3-8B-Instruct<\/code> models.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 3 reports the resulting F1 scores for <code>Gemma-3-1b-it<\/code> and <code>Llama-3-8B-Instruct<\/code>, respectively spanning all three settings. As anticipated, the centralised model outperforms its federated counterpart in every setting. This is an expected outcome, since centralised training enjoys unrestricted access to all data at once. One clearer pattern, however, emerges as the experiments scale.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The two regimes (centralized and federated) move in opposite directions**.**&nbsp;As more data is folded into the centralised pool, its overall score holds steady or even edges upward (for Gemma, from 70.21 \u2192 70.98 \u2192 71.87; for Llama, from 71.32 \u2192 72.34 \u2192 72.92). The federated model does the reverse: its score declines progressively as each new, more divergent data node joins (for Gemma, 64.56 \u2192 62.09 \u2192 60.16; for Llama, 68.64 \u2192 66.67 \u2192 61.08). In other words:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><em>Additional data seems to be an asset when pooled centrally, but a growing liability when it arrives as conflicting federated contributions.<\/em><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">To understand why, it helps to look more closely at how the federated server combines what its data nodes send back. The aggregation method we use here is FedAvg (Federated Averaging), the standard baseline in federated learning: after each round, the server takes the model updates returned by every data node and averages them into a new shared model. This works well when data nodes hold similar data, because their updates broadly agree and reinforce one another. In our setup, however, each data node learns from a deliberately different distribution, so their updates pull in different\u2014sometimes opposing\u2014directions. Averaging then cancels out these distribution-specific signals, leaving a global model that is a bland compromise, serving no single distribution well. This is why federated performance erodes on a universal test set drawn from all distributions, and why that erosion deepens as we add data nodes that diverge ever further from the rest.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Lesson Learned:<\/strong> Standard federated learning is not a drop-in replacement for centralised training under heterogeneous data: as data nodes diverge, na\u00efve FedAvg averaging cancels out their conflicting updates, so more participants can lower performance rather than raise it.<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Our ongoing work in this space is now exploring alternative federated techniques that can overcome this limitation.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The effect of model scale under federation<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Looking beyond the headline F1 comparison, we now turn to a more specific question: how does\u00a0LLM\u2019s size\u00a0affect performance on this text-classification task under federation? The scores reported so far have been\u00a0macro F1\u00a0scores, that is, the unweighted average of the F1 achieved on each of the three performance classes (Positive,\u00a0Negative, and\u00a0Average), treating every class as equally important regardless of how many samples it contains.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A single macro F1 figure is convenient for ranking configurations, but it hides&nbsp;where&nbsp;a model succeeds or struggles. Two models can reach the same macro score while being strong on different classes. To look underneath that average, Figure 2 decomposes performance for the three federated settings, plotting the&nbsp;macro (Total) F1&nbsp;alongside the three&nbsp;per-class&nbsp;scores (Positive,&nbsp;Negative, and&nbsp;Average&nbsp;F1) for both models on a single radar chart. The three panels are ordered left to right by increasing heterogeneity, letting us see not just&nbsp;whether&nbsp;each model degrades, but&nbsp;which classes&nbsp;drive that decline.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"381\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/new_all_metrics_radar-1024x381.jpeg\" alt=\"\" class=\"wp-image-1817\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/new_all_metrics_radar-1024x381.jpeg 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/new_all_metrics_radar-300x112.jpeg 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/new_all_metrics_radar-768x286.jpeg 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/new_all_metrics_radar-1536x572.jpeg 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/new_all_metrics_radar-2048x762.jpeg 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><strong>Figure 4:<\/strong>&nbsp;Per-class F1 comparison across the federated settings (2, 3, and 4 data nodes) for the two LLMs of differing parameter size, <code>Gemma-3-1b-it<\/code>&nbsp;and&nbsp;<code>Llama-3-8B-Instruct<\/code>.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Llama&#8217;s greater capacity appears to absorb mild and moderate heterogeneity more gracefully, reconciling divergent data node updates that a smaller model cannot: as Table 1 shows, at two data nodes its centralised\u2013federated gap is less than half of Gemma&#8217;s (2.68 versus 5.65), and it remains markedly smaller at three nodes (5.67 versus 8.89). This advantage, however, does not persist indefinitely. Once the adversarial&nbsp;<em>50% Flip<\/em>&nbsp;dataset enters the federation at four data nodes, both models degrade sharply and their gaps converge almost exactly (11.71 versus 11.84). The implication is clear:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">Greater model capacity provides meaningful resilience to heterogeneity, but even a state-of-the-art model reaches a breaking point when a sufficiently divergent data node is introduced.<\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Experimental setting<\/th><th>Gemma 3 (1B) Gap<\/th><th>Llama 3 (8B) Gap<\/th><\/tr><\/thead><tbody><tr><td>2 Data Nodes<\/td><td>5.65<\/td><td>2.68<\/td><\/tr><tr><td>3 Data Nodes<\/td><td>8.89<\/td><td>5.67<\/td><\/tr><tr><td>4 Data Nodes<\/td><td>11.71<\/td><td>11.84<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Table 1:<\/strong>&nbsp;Centralised\u2013federated F1 gap for each LLM across the three experimental setups; a smaller gap indicates the federated model remains closer to its centralised performance.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 4 makes this dynamic visible. In the two-node panel, the Llama-3 (8B) contour (orange) sits clearly outside the Gemma-3 (1B) contour (green) on every axis, evidence of the larger model\u2019s stronger retained performance under mild heterogeneity. That margin narrows in the three-node panel as a more divergent data node joins, and nearly disappears in the four-node panel, where the 50% Flip data node pulls both contours toward the centre until they overlap, mirroring the convergence shown in the table. Because the radar chart separates the classes, it also reveals how the collapse occurs: both models maintain the same asymmetric profile throughout, extended toward Average F1 yet compressed at Positive F1. This indicates that, regardless of model size, FedAvg protects the&nbsp;Average&nbsp;class, while sacrificing the&nbsp;Positive&nbsp;class. Because the strong Average score inflates the macro average, the single headline F1 number hides this lopsided trade-off entirely.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Lesson Learned:<\/strong> Greater model capacity buffers heterogeneity but does not cure it: a larger LLM narrows the centralised\u2013federated gap under mild and moderate divergence. Yet once a sufficiently adversarial data node joins, both models collapse similarly, because scale changes <em>how much<\/em> a model degrades, not <em>how<\/em>\u2014leaving FedAvg\u2019s aggregation as the true bottleneck.<\/p>\n<\/blockquote>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>The impact and looking ahead<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Our experiments suggest a clear pattern: as more <em>heterogeneous (non\u2011IID)<\/em> data nodes join, federated performance declines.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What this implies:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Heterogeneity is the driver: na\u00efve federated aggregation averages conflicting updates, diluting distribution\u2011specific signals.<\/li>\n\n\n\n<li>A single \u201cglobal\u201d model may be the wrong target: when data nodes hold different \u201cflavours\u201d of data, a one-size-fits-all approach under-serves everyone.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Where we go next:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Explore server-side strategies beyond na\u00efve aggregation (FedAvg), e.g. personalisation methods that group data nodes by update similarity.<\/li>\n\n\n\n<li>Treat heterogeneity as a signal to exploit (discover who can beneficially collaborate), not just an obstacle.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Heterogeneity, it turns out, is not merely an obstacle to tolerate but a signal to exploit. Learning to harness it intelligently may ultimately determine how far federated learning can scale across the real, fragmented landscape of the marketing industry.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Every click, scroll, and purchase leaves a trace. The more of these traces a Machine Learning (ML) model can study, the sharper its predictions and the stronger its commercial returns. The challenge is that no single organisation holds the full picture. The data is split across agencies, brands, and technology partners, each sitting on their [&hellip;]<\/p>\n","protected":false},"author":9,"featured_media":1790,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":""},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":9,"display_name":"Emmanouil Kritharakis","first_name":"Emmanouil","last_name":"Kritharakis","nickname":"emmanouil.kritharakis","user_nicename":"emmanouil-kritharakis","user_email":"emmanouil.kritharakis@satalia.com","biographical_info":"Emmanouil (Manos) Kritharakis is a Data Scientist at Satalia, working in the Research Lab on Graph Machine Learning and Federated Learning Systems. His research focuses on scalable graph-based methods and privacy-preserving distributed learning, with publications at leading venues including VLDB and ECML-PKDD. He brings hands-on experience building production ML systems end-to-end, from data preprocessing to deployment, bridging the gap between research and real-world applications.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/03\/Manos.jpg","job_title":"Researcher","is_lead":false,"display_as_researcher":true,"order_priority":null}],"class_list":["post-1789","research_feed","type-research_feed","status-publish","has-post-thumbnail","hentry","content_type-article"],"acf":{"content":"","content_quarter":"Q2 2026","related_pods":[192]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"Q2 2026","related_pods":["192"],"featured":"","legacy_perspective_source_id":""},"featured_image_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/021d5b91-ccfc-4b57-8943-5ba2b328ac24-1024x270.png","featured_image_sizes":{"thumbnail":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/021d5b91-ccfc-4b57-8943-5ba2b328ac24-150x150.png","medium":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/021d5b91-ccfc-4b57-8943-5ba2b328ac24-300x79.png","large":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/021d5b91-ccfc-4b57-8943-5ba2b328ac24-1024x270.png","full":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/021d5b91-ccfc-4b57-8943-5ba2b328ac24.png"},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/1789","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/9"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/192"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/media\/1790"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1789"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1789"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=1789"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=1789"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}