Author: Emmanouil Kritharakis

  • Learning from the best: Finding Compatible Partners in Federated Learning

    Marketing organisations hold records of every campaign they have run, which can train models to predict how new plans will perform. But coverage is limited: each organisation has tried only a narrow range of campaigns, and none is willing to share records with competitors. Federated learning offers a way to train a shared model without moving anyone’s data, but it assumes all participants are solving the same problem. Organisations often aren’t: the same combination of audience, creative and platform can work well for one organisation and poorly for another, for reasons that are not apparent from either organisation’s data alone. Averaging their contributions can therefore produce a model that fits none of them; in our experiments, it reduced every participant’s accuracy. We show how to identify compatible partners in advance using only what participants already share during training, so no campaign records change hands. Averaging within clusters rather than across everyone improves the model of every organisation with a compatible partner.


    Setting the problem

    One of the most important questions in campaign planning is also the hardest: Will this work? Organisations already have the raw material to answer it. Every campaign they have run is documented by its brand, audience, format, platform and market, along with its performance. Those records can train a machine learning model to predict the outcome of a plan that has not yet run.

    But predictive performance depends less on the model than on the records behind it. Most organisations work with a limited set of brands, a handful of audience segments, and a small number of platforms: a narrow slice of all possible campaigns. Asked about a combination beyond that experience, the model guesses. The experience of other organisations could close that gap.

    The obvious solution is to pool data across organisations. The obvious obstacle is that nobody is willing to share raw data. Campaign performance data reveals what you spend, whom you target, and what actually works: three things you do not hand to a competitor. Data protection regulation also restricts organisations from sharing records containing personal data.

    Federated Learning (FL) was designed for this challenge: the model moves to the data rather than the data moving to the model. A coordinating node sends the current model to each participant, which trains it on its own records and returns only the resulting parameter updates. The central node combines those updates into a model that, in theory, outperforms what any participant could build alone. We described this collaborative scenario for marketing records in an earlier post on fine-tuning a language model across organisations that keep their own data.

    That earlier post left open the question that determines whether collaboration is worth joining: Are participants solving the same prediction problem? A model trained on campaign records maps campaign attributes to a verdict on performance. But that verdict is not an intrinsic property of a campaign; it reflects each organisation’s own benchmarks. As a result, two organisations can reach opposite conclusions about identical results. Their records then offer conflicting answers to the same question rather than complementary evidence about it.

    Averaging updates in this situation pulls the shared model in opposing directions, producing a result that fits neither participant. In our experiments, every participant ended with a worse model than it would have had by training on its own records alone. This post describes two clustering methods for grouping compatible organisations before averaging. We show that the methods recover the intended grouping from the model updates participants already upload, and report their performance on a testbed built for that purpose.


    Why naive averaging updates fails

    The server combines orgnisations updates, weighting each by the number of records used for training. This is appropriate only if every organisation approximates the same target function: their updates are then independent estimates of the same quantity, and averaging reduces error.

    That assumption fails in two distinct ways, each requiring a different remedy.

    • Different outcomes for the same choices. A combination of campaign attributes that performs well for one organisation can perform poorly for another. The mapping from inputs to labels therefore differs between them, even when both describe campaigns identically and judge performance by the same standards.
    • Different vocabularies. Each organisation describes campaigns in its own terms. Even when their tables have matching columns, the values may differ. Organisations may use different audience taxonomies, target different markets, or describe creative content in free text that a table cannot capture. As a result, the model’s inputs are not comparable across participants.

    This post focuses on the first issue. Figure 1 illustrates a typical example.

    Figure 1. Concept shift. Both organisations run the campaign described above, using the same five attributes and the same performance criterion: click-through rate (CTR). Pairing a Gen Z audience with a summer-themed creative lifts performance in Organisation A’s records, yielding a CTR of 6.0% and an over-performing label. In Organisation B’s records, the same pairing depresses performance, yielding a CTR of 0.1% and an under-performing label. The identical combination therefore receives opposite labels.

    Two organisations run the same campaign, a summer-themed video advertisement for a fashion brand aimed at a Gen Z audience on one platform in one market. They describe the campaign with the same feature attributes, drawn from the same value space, and they judge the result by the same metric, click-through rate (CTR), and agree on which CTR values count as over-performing and which as under-performing. Organisation A measures a CTR of 6.0%, which both would call high. Organisation B measures 0.1%, which both would call low. Nothing about the description or the criteria differs, and the labels are opposite. This is known as “concept shift”: the mapping from inputs to labels differs between organisations [1], [2]. When the same combination of attributes leads to strong performance for one organisation and weak performance for another, averaging their updates does not reduce noise. Instead, it averages conflicting labelling functions, producing a result that fits neither.

    The second issue, different vocabularies, can be addressed at the representation level. This is known as “covariate shift”: the distribution of the inputs differs between participants, while the mapping from inputs to labels stays the same [2], [4]. We replace the fixed table with free-text campaign descriptions and fine-tune a pre-trained large language model (LLM) with a prediction head. This offers two benefits. Differences in schema no longer disrupt training, because there is no schema to reconcile. The model’s existing knowledge helps it connect descriptions that a table would treat as unrelated: it can recognise that “Low-Income Millennials who enjoy hiking” and “Budget-conscious Gen Y trail walkers” describe much the same people, where a table sees two unrelated strings.

    One important limitation: our experiments use a fixed input representation, with the same schema across all organisations. This article therefore addresses concept shift, not covariate shift. Testing covariate shift would require organisations whose attribute values genuinely differ, which is a direction for future work. Nevertheless, the free-text representation explains our choice of a language model as the baseline architecture.

    Which participants to average, not how

    Most published work on participant heterogeneity in federated learning changes how updates are combined: by keeping each participant close to the shared model [3], correcting for drift between participants [5], or adding a regularisation term to the local objective [6]. However, these methods cannot help when no single model fits every participant’s labels, as Figure 1 illustrates.

    The real question is which participants should be averaged together. Participants whose criteria agree hold complementary evidence about one function and can usefully be averaged. Participants whose criteria conflict cannot, whatever the averaging rule. The server has to answer that without access to any participant’s records. All it receives is the parameter updates each participant uploads, so a usable method has to find compatibility in those.

    What the uploaded model parameters contain

    Full fine-tuning a pre-trained LLM and transmitting it every round is impractical, so we focus on a parameter-efficient fine-tuning (PEFT) method called Low-Rank Adaptation (LoRA), which freezes the pre-trained model and trains a pair of small adapter matrices, A and B, alongside it [7]. Participants exchange only those LoRA adapters, which are a small fraction of the model’s parameters.

    The two matrices do different work. A holds general linguistic information, shared across participants because every participant describes campaigns in the same terms. B is task-specific: it holds the criteria by which a campaign record receives a performance label, and those criteria are what differ between participants. The same asymmetry is reported for low-rank adapters generally [8] and for federated fine-tuning in particular, where A carries knowledge transferable across clients and B carries what is specific to each one [9].

    That asymmetry tells us which matrix to read, not which to share. Participants upload both, one of each for every layer of the model, and our methods use the B matrices to decide who is compatible with whom, then average the A and B matrices separately within each group that decision produces. The server already holds everything it needs as a normal part of training, so deciding which participants are compatible requires no campaign records and nothing beyond what federated training already sends.

    How the server turns the LoRA adapters into groups

    The server compares each organisation’s B matrices against every other’s. It measures how closely two organisations’ matrices align: closely aligned means the two are applying similar criteria, poorly aligned means they are not. Those comparisons are the only evidence the server has for forming the groups. Figure 2 shows the full sequence, from those comparisons to the final groups, and the rest of this section takes it step by step.

    Figure 2. How the server forms the groups, shown for four organisations. (1) B matrices are compared against every other organisation’s. (2) The closest pair merges repeatedly into a dendrogram, whose height records how dissimilar each merge was. (3) Each candidate cut of the dendrogram is scored by the Silhouette and Davies-Bouldin measures. (4) The best-scoring cut fixes the groups, and averaging stays inside each one.

    Building a hierarchy and choosing where to cut

    From these comparisons, the server builds a hierarchy using agglomerative clustering. It starts with each participant in its own group and repeatedly merges the two closest groups until only one remains. The resulting tree does not prescribe a number of groups: cutting it high leaves everyone together, while cutting it lower divides them into two, three or more.

    The key decision is where to cut, and the server makes it without being told how many groups to expect. It tests each candidate number of groups and scores the division in two ways. The Silhouette score [10] compares how close each participant is to its own group with how close it is to the nearest other group, rewarding clear membership. The Davies-Bouldin index [10] compares the spread within groups with the distance between them, penalising groups that are loose relative to the gaps separating them. A high Silhouette score and a low Davies-Bouldin index both indicate groups that are internally similar and clearly separated. The best-scoring division fixes the groups, and the decision is made once. In every round that follows, the server sends each group’s averaged adapter only to the organisations in that group, those organisations train locally and upload their updates, and the server averages them within the same groups again.

    Two ways to apply the grouping

    We developed two versions of this approach, Model-Wise Clustering and Layer-Wise Clustering, both adapting existing work on clustering clients in federated LoRA. They differ in one respect: whether the hierarchy is cut once for the whole model, or cut again at every layer, so that groups divide as the model deepens. Figure 3 shows how each version assigns organisations to groups.

    Figure 3. The two grouping strategies, shown for four organisations. Numbers identify organisations, and each coloured box is a group whose contributions the server averages together at that layer. Model-Wise settles on a single grouping and applies it at every layer. Layer-Wise starts with every organisation in one group and lets the groups divide as the model deepens, never merging back, so organisations share the shallow layers and separate in the deeper ones.

    Model-Wise Clustering, influenced by [11], averages the comparisons across every layer into one distance per pair, builds a single hierarchy and cuts it once. Each organisation belongs to exactly one group, and that group decides which peers it is averaged with from the first layer to the last. The assumption is that an organisation’s criteria affect every layer to a similar degree.

    Layer-Wise Clustering, influenced by [12], drops that assumption. The layers of the model do different work: the shallow ones handle general language that applies to any campaign description, while the deeper ones hold the mapping from a description to a performance label. It is that mapping which varies between organisations, so a single grouping applied throughout must either over-share in the deep layers or under-share in the shallow ones.

    Layer-Wise therefore builds the same hierarchy, then walks the model from the first layer to the last, choosing at each depth how far down to cut. The evidence for each cut is the layer’s own distances, because two organisations can agree across the model as a whole and still diverge sharply at one depth. Because every cut comes from the same hierarchy, the groupings are nested rather than unrelated. Two rules govern the walk: the groups never merge back, so a divergence once detected holds deeper in, and a layer with no convincing split keeps the grouping of the one above it. The shallow layers stay unified and the groups divide as the model deepens, which is the pattern Figure 3 shows.


    How we evaluated the two grouping methods

    To judge a grouping, you have to know which grouping was right. A real federation never tells you: each participant’s criteria are private, so if accuracy changes, there is no way to attribute it to the grouping rather than to anything else. Therefore, we built a testbed where the answer is known in advance.

    Generating campaign records from a rule graph

    The testbed uses synthetic campaign records from a generator we built for this purpose, described in an earlier post on using synthetic data to stress-test marketing models. At its centre is a signed marketing knowledge graph, shown in the left panel of Figure 4. Each node is an attribute value, a brand, an audience, a creative treatment, a platform or a geography, and each edge records whether two of them work well together, shown in green, clash, shown in red, or have no bearing on each other.

    A campaign record is one value of each kind, and its label follows from the edges between them: mostly green gives over-performing, mostly red gives under-performing, and mixed cases give average. We generated 27,000 records per organisation, split into a training set of 18,000 and a held-out test set of 9,000.

    How the four organisations were made to disagree

    We created four organisations and gave each a copy of the rule graph, flipping a controlled share of the nodes in three of the four. Flipping a node means changing the type of every edge incident to it, so that green becomes red and red becomes green. A campaign involving one of those choices then carries a different balance of green and red edges, and its label changes where that shift reverses the majority. The campaign records themselves are unchanged, so what differs between the four organisations is the labels alone. Figure 4 shows the base graph beside one of the three derived from it, with 10% of its nodes flipped.

    Figure 4. The base rule graph, left, and a graph derived from it by flipping a randomly selected 10% of its nodes, right, here the single node Brand A, marked by the dashed rectangle. These are the graphs held by Organisation A and Organisation B in Table 1.

    Because we create the four graphs, we know the label each organisation’s criteria assign to every record it holds, and therefore how far any two organisations’ criteria diverge. Each organisation received a different share of flipped nodes. Table 1 gives the four organisations and how each one’s criteria relate to the original rules.

    OrganisationShare of nodes flippedRelation to the original rules
    A0%, the base graph.It labels campaigns by the original rules and is the reference for the other three.
    B10%It agrees with Organisation A on most campaigns and differs on a few.
    C20%It agrees with Organisation A on the broad picture and differs on more cases than Organisation B does.
    D50%Its labels run close to the opposite of Organisation A’s.
    Table 1. The four organisations in the federation, with the share of nodes flipped in each one’s rule graph and how the resulting labelling criteria relate to Organisation A’ s.

    The four configurations span the range of disagreement we expect between advertising organisations. Two brands in adjacent categories agree about what a strong campaign looks like and differ at the margins. A brand with a different commercial model, a luxury house against a discount retailer, can reach the opposite conclusion from the same measured outcome.

    Table 1 also fixes the grouping a method should produce. Organisations A, B and C belong in one group, because their criteria differ in degree rather than in kind and each holds records bearing on the others’ cases. Organisation D belongs on its own, because averaging its updates with the other three produces a model that fits none of the four. The sections that follow report which grouping each method produced.

    Results

    After federated training under each grouping strategy, we asked each organisation’s model two questions:

    Does it work on your own campaigns? We trained each model on that organisation’s training set and tested it on its own held-out test set. This measures whether the model has learned that organisation’s own labelling criteria, the question an advertiser cares about day to day.

    Does it generalise beyond your campaigns? We pooled the held-out test sets from all four
    organisations into one large merged set (36,000 records) and tested each model against that. This measures whether a model has learned something about campaign performance in general, or has merely memorised one labelling criteria. It is the harder test, and it is the one that reveals whether collaboration delivered anything at all.

    Both questions are scored with the same measure, the macro-averaged F1 score, reported as a percentage. It rewards a model for getting all three outcome campaign performance labels right (under-performing, average and over-performing), weighting the three equally, so failure on one cannot hide behind success on the others. Differences between approaches are quoted in percentage points.

    Figure 5. Results for the four organisations, labelled by the share of nodes whose incident edges were flipped (0%, 10%, 20%, 50%). (A) The groupings the two methods produced from the uploaded adapter matrices alone. (B) Macro-averaged F1, as a percentage, on each organisation’s own held-out test set. (C) The same score on the four held-out test sets merged into one. Bars in panels B and C are the unweighted mean of the four organisations’ scores, with the difference against training alone at the right of each bar.

    Both methods recovered the intended grouping

    As panel A of Figure 5 shows, both placed the 50% organisation on its own and grouped the 0%, 10% and 20% organisations together, without access to any participant’s records. The server held nothing except the adapter matrices participants upload as part of training, and from the pair-wise similarities between those matrices it recovered the relationship between the four organisations’ labelling criteria, which it was not given.

    The two methods then differ. Model-Wise applies that single grouping at every layer. Layer-Wise keeps it in the early layers and draws one further distinction in the deeper layers, separating the 0% organisation from the 10% and 20% pair, which matches the construction in Table 1: the three agree closely enough to share the layers that read a campaign and differ enough to separate in the layers that label it. That the distinction falls in the deeper layers is consistent with concept shift affecting the label rather than the reading.

    On its own test set, each organisation scored at or above training alone

    Panel B of Figure 5 reports each organisation’s model tested on its own held-out test set. Table 2 gives the individual scores behind those bars.

    Training ArrangementOrg. A – 0%Org. B – 10%Org. C – 20%Org. D – 50%Average
    Train alone78.472.670.059.070.0
    Train with everyone59.955.255.845.454.1
    Model-Wise Clustering78.975.472.260.571.8
    Layer-Wise Clustering80.475.772.760.472.3
    Table 2. Macro-averaged F1, as a percentage, on each organisation’s own held-out test set, under the four training arrangements. Columns are the four organisations by share of nodes flipped, with their unweighted mean at the right.

    The first row of Table 2 shows that an organisation training alone scores between 59.0 and 78.4 depending on how self-consistent its own rulebook is, and those four figures average to 70.0. Two findings follow.

    Clustering tracks each organisation’s own performance, and lifts it. Each clustering row follows the same shape as training alone, highest for the 0% organisation, lowest for the 50%, but sits above it at every point, gaining 1.8 points on average for Model-Wise and 2.3 for Layer-Wise. The margins are narrow, as expected of an organisation tested on its own campaigns with its full history available. What matters is the consistency: no organisation scored lower under either clustering method than it did training alone on its own test set.

    Naive averaging loses to training alone for every organisation, by 15.9 points on average. This is the result that should give any prospective data collaboration pause. Every participant would have been better off never joining.

    On the merged test set, clustering’s advantage over training alone grows

    Panel C reports the same models on a single merged test set containing campaigns from all four organisations. This is a harder question, because three quarters of those campaigns are labelled by a standard the model was never trained on. Table 3 gives the individual scores behind panel C of Figure 5.

    Training ArrangementOrg. A – 10%Org. B – 10%Org. C – 20%Org. D – 50%Average
    Train alone66.365.465.357.663.7
    Train with everyone54.254.254.254.254.2
    Model-Wise Clustering70.170.170.156.766.7
    Layer-Wise Clustering69.569.969.956.766.5
    Table 3. Macro-averaged F1, as a percentage, on the four held-out test sets pooled into one, under the same four arrangements. Columns identify the model tested rather than the test records, which are the same throughout, with the unweighted mean at the right.

    Every method scores lower here than on home ground, which is as it should be. Two further lessons follow.

    The clustering advantage roughly doubles on the merged test. Clustering gains 3.1 points on average for Model-Wise and 2.9 for Layer-Wise, where naive averaging is 9.5 points behind. Compared with Table 2, the advantage is larger once models are asked to judge campaigns labelled by criteria other than their own. That is where collaboration is supposed to help, and where it does.

    For the organisations that actually had partners, the gains are larger still. The averages in Table 3 include the 50% organisation, which by design had no partners. The three grouped organisations gained between 3.8 and 4.8 points over training alone: 66.3 rising to 70.1, 65.4 to 70.1, and 65.3 to 70.1. A model built on one dataset standard does not improve; one built on three related dataset standards does, because it has seen the same underlying dynamics judged in three slightly different ways and learned what they have in common.


    What comes next

    Compatibility determines which organisations can usefully train together, but not whether they should. In a real federation, access to a partner’s records comes at a price. Each organisation is both a seller, charging for access to its own records, and a buyer, using a fixed budget to acquire access to others’ data. Partner selection is therefore an economic as well as a statistical decision.

    The methods described here address the statistical question: which participants are likely to benefit from training together? The next step is to give the server pricing information alongside the similarities it already computes, so it can form groups that deliver the greatest expected benefit to each participant within its budget. A partner who offers only a small improvement may not be worth the price, and the grouping that looks best by similarity may exceed every participant’s budget. We are now focused on extending these methods to partner selection under cost and budget constraints.


    Conclusion

    Federated learning creates value not by bringing more participants into the same model, but by ensuring that the right participants learn together. In our experiments, averaging updates from organisations with conflicting definitions of campaign success made every participant worse off than training alone, while clustering compatible organisations preserved or improved performance on their own campaigns and delivered larger gains beyond their home data. Crucially, the server identified those partnerships using only the LoRA matrices participants already upload, without seeing their campaign records, labels or business criteria. The practical implication is clear: privacy-preserving collaboration needs a compatibility step; the question is not whether more data helps, but whose data helps whom.


    References

    [1] Moreno-Torres et al. A Unifying View on Dataset Shift in Classification. In Pattern Recognition, 2012. https://doi.org/10.1016/j.patcog.2011.06.019

    [2] Kairouz et al. Advances and Open Problems in Federated Learning. In Foundations and Trends in Machine Learning, 2021. https://arxiv.org/abs/1912.04977

    [3] Li et al. Federated Optimization in Heterogeneous Networks. In MLSys 2020. https://arxiv.org/abs/1812.06127

    [4] Zhu et al. Federated Learning on Non-IID Data: A Survey. In Neurocomputing, 2021. https://doi.org/10.1016/j.neucom.2021.07.098

    [5] Karimireddy et al. SCAFFOLD: Stochastic Controlled Averaging for Federated Learning. In ICML 2020. https://arxiv.org/abs/1910.06378

    [6] Acar et al. Federated Learning Based on Dynamic Regularization. In ICLR 2021. https://openreview.net/forum?id=B7v4QMR6Z9w

    [7] Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR 2022. https://arxiv.org/abs/2106.09685

    [8] Tian et al. HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning. In NeurIPS 2024. https://arxiv.org/abs/2404.19245

    [9] Guo et al. Selective Aggregation for Low-Rank Adaptation in Federated Learning. In ICLR 2025. https://arxiv.org/abs/2410.01463

    [10] Davies and Bouldin. A Cluster Separation Measure. In IEEE Transactions on Pattern Analysis and Machine Intelligence, 1979. https://doi.org/10.1109/TPAMI.1979.4766909

    [11] Wang et al. Adaptive LoRA Experts Allocation and Selection for Federated Fine-Tuning. In NeurIPS 2025. https://arxiv.org/abs/2509.15087

    [12] Bian et al. FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-Tuning. In ICML 2026. https://arxiv.org/abs/2603.13282


    Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.

  • Federated LLM Fine-Tuning for Marketing: Collective Intelligence from Sovereign Data

    Federated LLM Fine-Tuning for Marketing: Collective Intelligence from Sovereign Data

    Every click, scroll, and purchase leaves a trace. The more of these traces a Machine Learning (ML) model can study, the sharper its predictions and the stronger its commercial returns. The challenge is that no single organisation holds the full picture. The data is split across agencies, brands, and technology partners, each sitting on their own piece of the puzzle.

    In our earlier work, we laid the groundwork for a federated way for organisations to collaboratively train a shared ML model for media performance, without having to centralize all their data puzzle pieces in one place. Each participant trains locally within their own infrastructure, and exchanges only model updates (weights) rather than raw data.

    We then tested this idea with a set of experiments on simple ML models, comparing a centralised baseline (pooling everyone’s data in the same place and training a model on the full puzzle) with a federated counterpart that keeps the puzzle distributed across several data nodes. Alongside this, we examined how fragmentation to an increasing number of participating nodes affects overall performance.

    Those initial experiments rested on two assumptions:

    • All the data puzzle pieces look alike. Conceptually, this assumes that every data partner observes and records similar patterns and behaviors. This is a simplification of the real-world, where brands, agencies, and platforms address a diverse spectrum of audiences, products, and experiences.
    • The federated ML model follows a traditional “tabular” architecture. Every row in its input is an observation (e.g. a marketing campaign), every column is a an aspect of that observation (e.g. the targeted audience, the advertised product). The last column is the outcome we are trying to model (e.g. the campaign’s performance). The model then looks for statistical associations among the columns from scratch, without having its own prior knowledge of the world. Again, this is a simplification of today’s AI/ML model landscape, where LLMs with vast prior world knowledge dominate.

    In this next phase of our work, we relax both of these assumptions by exploring:

    • Federated Learning in the presence of nodes with heterogeneous datasets
    • Federated Fine-tuning of LLM-based predictive models

    Contextualising the data: Marketing campaigns

    Before turning to the experiments, it helps to revisit what the marketing dataset represents and where it comes from. We use synthetic datasets generated by a purpose-built pipeline designed to mirror the dynamics of real marketing campaigns. This approach gives us fine-grained control over each dataset’s composition, allowing us to construct targeted scenarios that stress-test the model under controlled conditions while retaining full visibility into the factors that drive campaign outcomes.

    At the heart of this pipeline is a signed marketing knowledge graph, which encodes how marketing attributes interact. Each node is an attribute value—the brand being promoted, the audience it targets, the content’s creative tone and style, the platform it runs on, or the geography it reaches—and each edge captures whether two attributes are compatible (positive signal), incompatible (negative signal), or unrelated (missing signal).

    This graph is the ground truth from which we sample subgraphs to build various synthetic datasets. Each dataset contains campaigns defined as combinations of these attributes: one brand, one audience, one creative (image or video), one platform, and one geographical location, drawn from the graph. Labels follow from the signals linking them: campaigns built from compatible pairings tend to over-perform (Positive), those from incompatible ones tend to under-perform (Negative), and mixed cases fall in between (Average). For full details on how the graph is built and sampled, see the synthetic data generator pod.


    Figure 1: Visualisation of two synthetic marketing subgraphs for the Base and 10% Flip datasets. “Flipping” means changing the edge type of every edge incident to a randomly selected 10% of nodes (highlighted by the dashed rectangle).

    Generating heterogeneous data

    For our experiments, we utilised the ground-truth graph to extract four subgraphs for four respective datasets. To meet the requirement of using heterogeneous datasets, we deliberately add structural noise by randomly “flipping” the connected edges of a subset of subgraph nodes from positive to negative and vice versa. Figure 1 illustrates this on a sample subgraph: the left-hand panel (“Base Subgraph”) shows it in its original state. The right-hand panel (“10% Flipped Subgraph”) then shows a flip in action: we randomly select a fraction of nodes (here Brand A, boxed by the dashed rectangle) and reverse the signal on every edge connected to them, turning positive (green) edges negative (red) and vice versa. The result is a graph that tells a partly contradictory story, simulating the noisy, conflicting evidence of real marketing data.

    The dataset names follow directly the fraction of nodes flipped, represented as percentage in each name. More specifically:

    • Base (0%) leaves the graph untouched as a clean reference
    • 10% Flip reverses a random 10% of nodes (the case in Figure 1)
    • 20% Flip and 50% Flip apply the same step to larger fractions, adding steadily more noise

    Since every dataset shares one origin and differs only in flip rate, this provides a controlled way to observe how the model copes as conflicting evidence accumulates—each dataset large enough to reach stable performance even without training to full equilibrium.


    From tables to text: Federated LLM fine‑tuning with LoRA

    The synthetic datasets start out as tables, where each row is a single campaign described by its features (brand, geography, audience, platform, and creative), and the last column records how that campaign performed—Positive, Negative, or Average. To prepare the data, we first replace these three labels with simple number codes (0, 1, or 2). We then treat the problem as a text classification task, where the model’s job is to read a campaign and produce a single answer at the end indicating its predicted performance.

    Before training, every dataset goes through the same preparation steps, so all the inputs look consistent. Each campaign row is rewritten as a short, plain-language description that the LLM model can read, with each feature clearly labelled. For example, one campaign might look like this:

    Platform: Platform A
    Brand: Premium outdoor apparel brand
    Audience: Young professionals interested in fitness and travel
    Creative: Humorous
    Geo: [1023, 2045, 9876] # Zip codes expression of the place
    

    This description is then placed inside a fixed set of instructions that presents the task in exactly the same way for every experiment. During training, we include the correct performance label so the model can learn the patterns. During testing, we leave the label out, and the model has to work it out on its own—deciding whether the campaign is Positive, Negative, or Average.

    Figure 2: Visualization of the Federated Learning setup using FlowerTune with LoRA. Black numbered dots depict the steps of the FL training process (described below).

    We recast the marketing problem as a task an LLM can solve, enabling us to fine-tune an LLM in a federated setup using FlowerTune with LoRA (Low-Rank Adaptation). We chose this approach because marketing data is typically sensitive and siloed across data nodes, and a federated setup keeps that data private: the model is trained locally at each node, so only the resulting updates—never the raw data—are ever shared. FlowerTune, the LLM fine-tuning tool in the Flower federated-learning framework, coordinates this process, while LoRA makes it practical: rather than exchanging billions of parameters each round, LoRA freezes the base model and trains only a small set of adapter weights, reducing the trainable parameters by over 99% and keeping communication cheap. As illustrated in Figure 2, a single communication round then proceeds in four steps: (1) Local training: each data node freezes the massive pre-trained base model and updates only the small set of LoRA adapters on its private, siloed dataset, so the raw data never leaves the node; (2) Upload updates: each node sends only these lightweight adapters—not the raw data or the full model—to the central server; (3) Aggregation: the server combines the adapters from all participating nodes into a single set of aggregated LoRA adapters, producing one shared, improved model; and (4) Broadcast: the server distributes these aggregated adapters back to every node to serve as the starting point for the next round. This cycle repeats over successive communication rounds until the model converges.


    Experiments & lessons learned

    For our experimental setup, we evaluate our text classification task using two LLMs of increasing size: Gemma-3-1b-it and Llama-3-8B-Instruct (1 billion and 8 billion parameters, respectively). Working with two models of different scale lets us assess not only the centralised-versus-federated gap, but also whether model capacity influences how gracefully each setup absorbs heterogeneous data.

    For each model, we compare a centrally-trained LLM against its federated-trained counterpart across three progressively more heterogeneous settings:

    • Setting 1 (2 Data nodes): The centralised model is trained on the combined samples of the Base and 10% Flip datasets. The federated equivalent uses two data nodes, one holding Base and the other 10% Flip.
    • Setting 2 (3 Data nodes): The 20% Flip dataset is added, to the centralised training pool, and as an additional federated data node.
    • Setting 3 (4 Data nodes): The 50% Flip dataset is added in the same manner, giving the most heterogeneous configuration.

    All models are evaluated on a single, universal test set containing equal numbers of samples from every distribution (Base, 10%, 20%, and 50% Flip). This common benchmark provides a consistent view of how well each trained model generalises across the full spread of perspectives present in the federation. Each centralised scenario was trained for two passes over its own data; similarly each federated experiment used two communication cycles: in each cycle, every data node trained the model locally for one pass over its own data, then sent only the resulting model update to the server, which combined the updates before starting the next and final cycle, keeping total training exposure comparable between the two regimes. All other training parameters remain the same across the centralized version and its federated counterpart.

    All models are evaluated on a single, universal test set containing equal numbers of samples from every distribution (Base, 10%, 20%, and 50% Flip). This common benchmark provides a consistent view of how well each trained model generalises across the full spread of perspectives present in the federation. Each centralised scenario was trained for two passes over its own data; similarly each federated experiment used two communication cycles: in each cycle, every data node trained the model locally for one pass over its own data, then sent only the resulting model update to the server, which combined the updates before starting the next and final cycle, keeping total training exposure comparable between the two regimes. All other training parameters remain the same across the centralized version and its federated counterpart.

    Centralised vs. federated: The cost of heterogeneity

    Figure 3: F1 score comparison among the centralized vs FL versions under the 3 different experimental settings between the Gemma-3-1b-it and Llama-3-8B-Instruct models.

    Figure 3 reports the resulting F1 scores for Gemma-3-1b-it and Llama-3-8B-Instruct, respectively spanning all three settings. As anticipated, the centralised model outperforms its federated counterpart in every setting. This is an expected outcome, since centralised training enjoys unrestricted access to all data at once. One clearer pattern, however, emerges as the experiments scale.

    The two regimes (centralized and federated) move in opposite directions**.** As more data is folded into the centralised pool, its overall score holds steady or even edges upward (for Gemma, from 70.21 → 70.98 → 71.87; for Llama, from 71.32 → 72.34 → 72.92). The federated model does the reverse: its score declines progressively as each new, more divergent data node joins (for Gemma, 64.56 → 62.09 → 60.16; for Llama, 68.64 → 66.67 → 61.08). In other words:

    Additional data seems to be an asset when pooled centrally, but a growing liability when it arrives as conflicting federated contributions.

    To understand why, it helps to look more closely at how the federated server combines what its data nodes send back. The aggregation method we use here is FedAvg (Federated Averaging), the standard baseline in federated learning: after each round, the server takes the model updates returned by every data node and averages them into a new shared model. This works well when data nodes hold similar data, because their updates broadly agree and reinforce one another. In our setup, however, each data node learns from a deliberately different distribution, so their updates pull in different—sometimes opposing—directions. Averaging then cancels out these distribution-specific signals, leaving a global model that is a bland compromise, serving no single distribution well. This is why federated performance erodes on a universal test set drawn from all distributions, and why that erosion deepens as we add data nodes that diverge ever further from the rest.

    Lesson Learned: Standard federated learning is not a drop-in replacement for centralised training under heterogeneous data: as data nodes diverge, naïve FedAvg averaging cancels out their conflicting updates, so more participants can lower performance rather than raise it.

    Our ongoing work in this space is now exploring alternative federated techniques that can overcome this limitation.

    The effect of model scale under federation

    Looking beyond the headline F1 comparison, we now turn to a more specific question: how does LLM’s size affect performance on this text-classification task under federation? The scores reported so far have been macro F1 scores, that is, the unweighted average of the F1 achieved on each of the three performance classes (Positive, Negative, and Average), treating every class as equally important regardless of how many samples it contains.

    A single macro F1 figure is convenient for ranking configurations, but it hides where a model succeeds or struggles. Two models can reach the same macro score while being strong on different classes. To look underneath that average, Figure 2 decomposes performance for the three federated settings, plotting the macro (Total) F1 alongside the three per-class scores (Positive, Negative, and Average F1) for both models on a single radar chart. The three panels are ordered left to right by increasing heterogeneity, letting us see not just whether each model degrades, but which classes drive that decline.

    Figure 4: Per-class F1 comparison across the federated settings (2, 3, and 4 data nodes) for the two LLMs of differing parameter size, Gemma-3-1b-it and Llama-3-8B-Instruct.

    Llama’s greater capacity appears to absorb mild and moderate heterogeneity more gracefully, reconciling divergent data node updates that a smaller model cannot: as Table 1 shows, at two data nodes its centralised–federated gap is less than half of Gemma’s (2.68 versus 5.65), and it remains markedly smaller at three nodes (5.67 versus 8.89). This advantage, however, does not persist indefinitely. Once the adversarial 50% Flip dataset enters the federation at four data nodes, both models degrade sharply and their gaps converge almost exactly (11.71 versus 11.84). The implication is clear:

    Greater model capacity provides meaningful resilience to heterogeneity, but even a state-of-the-art model reaches a breaking point when a sufficiently divergent data node is introduced.

    Experimental settingGemma 3 (1B) GapLlama 3 (8B) Gap
    2 Data Nodes5.652.68
    3 Data Nodes8.895.67
    4 Data Nodes11.7111.84

    Table 1: Centralised–federated F1 gap for each LLM across the three experimental setups; a smaller gap indicates the federated model remains closer to its centralised performance.

    Figure 4 makes this dynamic visible. In the two-node panel, the Llama-3 (8B) contour (orange) sits clearly outside the Gemma-3 (1B) contour (green) on every axis, evidence of the larger model’s stronger retained performance under mild heterogeneity. That margin narrows in the three-node panel as a more divergent data node joins, and nearly disappears in the four-node panel, where the 50% Flip data node pulls both contours toward the centre until they overlap, mirroring the convergence shown in the table. Because the radar chart separates the classes, it also reveals how the collapse occurs: both models maintain the same asymmetric profile throughout, extended toward Average F1 yet compressed at Positive F1. This indicates that, regardless of model size, FedAvg protects the Average class, while sacrificing the Positive class. Because the strong Average score inflates the macro average, the single headline F1 number hides this lopsided trade-off entirely.

    Lesson Learned: Greater model capacity buffers heterogeneity but does not cure it: a larger LLM narrows the centralised–federated gap under mild and moderate divergence. Yet once a sufficiently adversarial data node joins, both models collapse similarly, because scale changes how much a model degrades, not how—leaving FedAvg’s aggregation as the true bottleneck.


    The impact and looking ahead

    Our experiments suggest a clear pattern: as more heterogeneous (non‑IID) data nodes join, federated performance declines.

    What this implies:

    • Heterogeneity is the driver: naïve federated aggregation averages conflicting updates, diluting distribution‑specific signals.
    • A single “global” model may be the wrong target: when data nodes hold different “flavours” of data, a one-size-fits-all approach under-serves everyone.

    Where we go next:

    • Explore server-side strategies beyond naïve aggregation (FedAvg), e.g. personalisation methods that group data nodes by update similarity.
    • Treat heterogeneity as a signal to exploit (discover who can beneficially collaborate), not just an obstacle.

    Heterogeneity, it turns out, is not merely an obstacle to tolerate but a signal to exploit. Learning to harness it intelligently may ultimately determine how far federated learning can scale across the real, fragmented landscape of the marketing industry.

    Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.

  • Training together, sharing nothing: The promise of Federated Learning

    Why Federated Learning now?

    In marketing, data is a competitive edge. The more audience signals, campaign performance data, and consumer behaviour a Machine Learning (ML) model can learn from, the sharper its predictions and the greater its business impact. Across the marketing ecosystem, spanning agencies, brands, and technology partners, organisations individually hold rich and valuable datasets. The potential to learn shared patterns across these assets, without ever exposing proprietary information, could unlock new capabilities: better audience targeting, smarter media spend, and faster creative optimisation at a global scale.

    But here’s the challenge. Across the marketing industry, the most impactful data is inherently distributed. Each organisation, whether agency, brand, or technology partner, holds a unique piece of the data puzzle. Client contracts, privacy regulations like GDPR, and the sheer sensitivity of consumer-level data mean this data has to stay within each organisation’s walls. This is a structural reality of the industry, not a limitation of any single company. The question is whether there is a way to learn collectively from this distributed knowledge without compromising the privacy boundaries that exist for good reason.

    The traditional solution, centralised ML, pools raw data from multiple sources into a single cloud to train a global model. But uploading terabytes of sensitive data to a central server creates severe network latency and exposes collaborators to data breaches and potential violations of privacy regulations.

    Distributed ML methods attempted to address this by splitting training across local worker nodes. Whilst this reduces latency and avoids centralising raw data, these architectures were designed for internal computing clusters, not secure collaboration between independent companies. Without cross-organisation coordination, each organisation’s models are limited to what their own data can teach them, with no mechanism to benefit from shared learning.

    Problem: The collaboration vs. privacy bottleneck

    Organisations face a fundamental tension: gaining the benefits of shared learning typically requires centralising sensitive data, which privacy and contractual obligations rightly prevent. Yet without a way to learn collectively, each organisation’s models are limited to their own data alone. Neither traditional architecture offers a path to collaborative model improvement whilst keeping private data strictly where it belongs.

    Federated Learning (FL) offers a way out of this dilemma by bringing the model to the data, rather than the other way around. To understand why this shift matters, let’s look at how FL actually works under the bonnet.


    How Federated Learning works

    Figure 1: Overview of the Federated Learning communication cycle between a central node and distributed client nodes.

    Federated Learning (FL) enables multiple organisations to collaboratively train a shared model without ever centralising raw data. Instead of moving data to the model, FL brings the model to the data. Training proceeds through iterative rounds:

    1. A central server sends the current global model to all participating nodes (Blue arrow).
    2. Each node trains the model locally on its own private data (Green arrow).
    3. Nodes send back only their model updates, never the underlying data (Pink arrow).
    4. The server aggregates these updates into an improved global model and starts the next round.

    Throughout this process, raw data never leaves its source. Only learned model representations are exchanged across the network.

    The multimodal challenge

    Whilst the above privacy-preserving framework is valuable in its own right, modern marketing data adds another layer of complexity. Organisations do not just work with spreadsheets and numbers. They work with images, video, text, audio, and structured data, often all at once. A single campaign might involve visual brand assets, ad copy, audience segments, and performance metrics across channels. Training models that can reason across these different data types, known as multimodal learning, is already one of the most demanding challenges in ML.

    Now combine that with the constraints of federated learning. Each client may hold different combinations of modalities, in different formats and volumes. One partner might contribute rich visual data, another mostly text and tabular records. Coordinating a single global model that learns effectively from this fragmented, heterogeneous landscape, without ever seeing the raw data, pushes the problem to a new level of complexity.

    This is precisely what makes the intersection of FL and multimodal learning so important, and so hard. If it can be made to work, it unlocks collaborative intelligence across organisations at a scale that neither approach could achieve alone.


    Our objective: can Federated Learning deliver?

    The promise of FL is compelling, but before investing in real-world deployment, we need to answer a fundamental question:

    Does federated learning actually work well enough on multimodal marketing data to justify the tradeoff?

    Centralised training will always have an inherent advantage, it sees all the data at once. The question is not whether FL can beat centralised performance, but whether it can get close enough to make the privacy and collaboration benefits worthwhile. And beyond raw performance, we need to understand how FL behaves under realistic stress conditions: more partners joining, noisy data, and complex cross-modal relationships.

    To answer this, we designed a series of experiments around four key questions:

    Experiment 1 — Centralised vs. federated performance

    • How close can FL get to centralised performance? In a centralised setup, the model sees all the data at once, the ideal scenario for learning. FL, by design, fragments this data across nodes. The first question is whether this tradeoff costs us meaningful accuracy, or whether FL can match centralised results despite never accessing the full dataset.
    • What happens as more clients join? In practice, a federated network might involve a handful of partners or dozens. As the number of participants grows, each node holds a smaller, potentially less representative slice of the overall data. We tested how model performance scales as we increase the number of nodes.

    Experiment 2 — Resilience to noisy data

    • How robust are centralised and federated models to noisy data? Real-world datasets are messy, labels can be wrongly defined, and data quality varies across partners. We deliberately introduced noise into the multimodal dataset to simulate these imperfections and measure how much degradation the model can tolerate before performance breaks down.

    Experiment 3 — Cross-modal relationships

    • How sensitive centralised and federated models to underlying cross-modal patterns? Multimodal models learn by connections between different types of data. For example, a luxury brand might target a high-income audience through a premium creative tone on a specific platform. Some of these connections appear frequently in the data, whilst others are rare. We tested whether emphasising the most frequent cross-modal patterns in our synthetic data improves performance compared to emphasising the least frequent ones, helping us understand how much the model benefits from common, naturally occurring relationships versus rare, atypical ones.

    The data

    For our experiments, we used a multimodal synthetic dataset generated by our own well-tested synthetic data generator, designed to mirror real-world marketing dynamics. The generator allows us to customise various elements of the data and design targeted datasets that stress-test our model architecture under controlled conditions, giving us full visibility into the factors that drive campaign performance.

    Each campaign in the dataset is described using five key modalities:

    • Audience – the consumer segment being targeted
    • Brand – the positioning and perception of the brand
    • Creative – the tone and message of the campaign
    • Platform – where the campaign runs
    • Geography – the markets being targeted

    Each dataset’s sample is assigned a target label, Positive (Over performing), Negative (Under performing), or Average (Average performance), indicating whether that particular combination of modalities would lead to a successful, underperforming, or average campaign outcome.

    Experimental results

    All federated experiments are implemented using Flower, a widely adopted open-source framework for federated learning research and deployment. Flower allows us to simulate multi-client federated setups in a controlled environment, making it possible to rigorously test different configurations before moving to a fully distributed architecture.

    To ensure a fair comparison between centralised and federated setups, we kept the playing field level. Both setups use the exact same model architecture, so any performance differences come from how the model is trained, not what is being trained. In the federated setup, data is split equally across nodes, so that each partner sees a representative sample. This way, when we increase the number of nodes, any change in performance can be attributed to the scaling itself, not to differences in what each node’s data looks like.

    Experiment 1: How much accuracy do we trade for privacy?

    Figure 2: Impact of increasing node fragmentation on Federated Learning performance. Performance clearly degrades as the number of nodes increases from 5 to 15, compared to the baseline centralised model version.

    The centralised model sets the performance ceiling at 79.67%. This is expected, when a single model has direct access to all the data at once, it has the best possible conditions to learn. No information is lost to partitioning, and no coordination overhead is introduced. It’s the ideal scenario, and the benchmark everything else is measured against.

    The federated results tell a clear story: as we add more nodes, performance gradually declines. With 5 nodes, the model reaches 76.23%, a modest drop from the centralised baseline. But as we scale to 10 and then 15 nodes, scores fall to 70.29% and 67.65% respectively. The same pattern holds across all metrics, with the sharpest drops in the model’s ability to correctly identify both positive and negative cases.

    Why does this happen? As more nodes join, the total dataset gets divided into smaller slices. Each node sees less data, which means each node’s local training produces a less reliable picture of the overall patterns. When the server combines these local updates, the differences between them make it harder to converge on a strong global model, an effect we call the “aggregation penalty.”

    Lesson learned: FL with 5 nodes comes remarkably close to centralised performance, showing that federated collaboration is viable with minimal accuracy loss. However, as the number of nodes grows, makes it progressively harder for the global model to match centralised results.

    Experiment 2: How robust are centralised and federated models to noisy data? 

    In practice, marketing data is never perfectly clean. Campaign outcomes don’t fall neatly into “this worked” or “this didn’t.” Was a campaign that slightly exceeded expectations truly a success, or just average? Was a modest underperformance a failure, or noise in the measurement? Different teams may label the same outcome differently, tracking systems introduce inconsistencies, and the line between a “positive” and “average” campaign is often blurry.

    To simulate this reality, we deliberately introduced noise into our synthetic dataset by blurring the boundaries between performance classes. With no noise, the labels are clean — positive, negative, and neutral outcomes are clearly separated. As we increase the noise level from low, to medium, and then to high, the boundaries between these classes increasingly overlap, making it harder for the model to tell them apart. Think of it like gradually turning up the fog: the underlying patterns are still there, but they become harder to see. The federated learning simulation for this experiment was configured with 5 participating clients, consistent with the best-performing federated setup identified in Experiment 1.

    Figure 3: Performance comparison of centralised (left) and federated learning (right) configurations across increasing noise levels. Both paradigms degrade gradually, with Positive F1 and Negative F1 most affected, whilst the performance gap between the two remains approximately constant across all conditions.

    As expected, both models perform best on clean data and gradually decline as noise increases. At high noise:

    • The centralised model’s score drops from 82.05% to 78.11%
    • The FL model’s score drops from 80.74% to 76.09%

    The good news: neither model collapses. Even at the highest noise level both models still perform reasonably well. The overall accuracy dips, and the models struggle most with distinguishing clearly positive or clearly negative campaigns, which makes sense, since those are exactly the boundaries we blurred. However, their ability to capture general patterns across the dataset remains stable throughout.

    As in Experiment 1, the centralised model maintains a consistent edge over the federated setup at every noise level, but the gap between them stays roughly the same. This means that FL doesn’t become more fragile in noisy conditions; it handles data messiness about as well as its centralised counterpart.

    Lesson learned: Real-world data is inherently noisy, and any viable model must be able to handle that. Both centralised and FL models show strong resilience — performance declines gradually rather than breaking down, even when the data is heavily corrupted. Importantly, FL’s relative performance holds steady across noise levels, suggesting it is no more vulnerable to messy data than centralised training.

    Experiment 3: How sensitive are centralised & federated models to underlying cross-modal patterns?

    Our synthetic data generator creates campaign data based on a graph of relationships between five key factors: Audience, Brand, Creative, Platform, and Geography. Each relationship captures whether a particular combination of these factors tends to drive strong or weak campaign performance. Some of these relationships are common and obvious — they show up frequently and reflect well-known marketing dynamics. Others are rare and subtle — unusual combinations that don’t appear often but may carry uniquely valuable signal about what makes a campaign succeed or fail.

    Understanding how these different types of patterns affect learning is important for both training paradigms. If the nature of the underlying data patterns matters, we need to know whether centralised and federated models respond to them in the same way — or whether one setup handles certain patterns better than the other. To investigate this, we generated three versions of our dataset, keeping everything else the same:

    • Common-first: The generator focuses on the most frequently occurring combinations and downplays the rarest ones. This gives us a dataset dominated by typical, familiar marketing patterns.
    • Rare-first: The opposite — the generator prioritises the rarest combinations and downplays the most common. This fills the dataset with unusual, less obvious patterns.
    • Middle-ground: The generator focuses on combinations that fall in the middle of the frequency spectrum, neither the most common nor the rarest.

    As in Experiment 2, the federated learning simulation was run with 5 participating nodes, and performance was compared against the centralised baseline across all three dataset versions.

    Figure 4: Impact of cross-modal relationships on model performance. Prioritising rare feature combinations (Rare-first) substantially improves accuracy compared with focusing on common patterns, showing that atypical relationships provide a stronger learning signal for both centralised and federated learning paradigms.

    The results were striking. The Rare-first configuration dramatically outperformed the other two, achieving peak scores of 94.41% (Centralised) and 93.43% (FL), compared to scores in the 86–88% range for the Common-first and Middle-ground setups.

    This tells us something counterintuitive: the model learns far more from unusual feature combinations than from common ones. The typical, frequently seen patterns are in some sense “easy”, they don’t give the model much new information. But rare combinations force the model to learn more nuanced and distinctive boundaries between what makes a campaign succeed or fail.

    As in previous experiments, the centralised model maintains a small edge over FL, but the ranking between dataset strategies stays the same in both setups. Whether training centrally or federally, prioritising rare patterns is the winning strategy.

    Lesson learned: Not all data is equally valuable. Prioritising on rare, atypical feature combinations produces significantly better models than focusing mostly on common patterns. This has direct implications for how we design synthetic datasets: rather than mimicking the most typical marketing dynamics, we should deliberately include uncommon combinations to give the model a richer and more discriminative learning signal.

    The impact and looking ahead

    This work is just the initial spark for our federated learning efforts. Verifying that the centralised ML model performance our company provides is slightly degraded under a reasonable number of users opens the discussion about delivering ML solutions that address shared challenges among clients who are reluctant to share data to tackle a common industry problem. The FL approach allows companies to securely train a shared global model on their own datasets without the risk of data leakage throughout the training process.

    Although Federated Learning has been an established collaborative learning method since 2017, it remains a highly active research domain in academia and a strategic priority for industrial implementation.

    Our next objective is to stress-test and further expand our federated learning (FL) infrastructure by enabling learning across nodes that hold highly heterogeneous data, with substantially different shapes, feature spaces, and underlying distributions. This introduces a number of technical challenges, including how to align representations, aggregate knowledge effectively, and maintain stable performance when local data varies significantly from node to node. Overcoming these challenges will unlock deeper insights into the robustness and scalability of our FL framework, and will allow our models to learn more effectively in realistic, decentralised settings where data heterogeneity is the norm rather than the exception.

    Ready to explore the specifics? Read our full technical deep dive into Multimodal Federated Learning for a closer look at our methodology.

    Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP AI Lab team.

  • Multimodal Federated Learning Pod: Technical walkthrough

    Building strong marketing Machine Learning (ML) models requires diverse data spanning multiple facets of campaign performance, but the most valuable data is distributed across organisations and, for good reason, stays firmly within their walls. This structural reality limits the industry’s ability to build truly intelligent campaign optimisation at scale. Federated Learning (FL) offers a way forward: each participant trains on their own data locally, sharing only model updates, never a single row of raw data. In our experiments on synthetic multimodal marketing data, FL with a small number of participants came remarkably close to centralised performance, proved resilient to noisy data, and benefited equally from rare feature combinations, a surprisingly powerful driver of model quality. The findings lay the groundwork for deploying privacy-preserving ML solutions that enable collaborative learning whilst ensuring every organisation’s data remains strictly under its own control.

    If you don’t care about the technical details, read our blog post instead. Alternatively you can check out our code on Github.

    Multimodal Federated Learning for marketing outcome prediction: A deep dive Flower-based simulation analysis

    Training high-capacity models for marketing prediction is often constrained not by algorithmic complexity but by data locality. Across the marketing industry, informative signals naturally reside within individual organisations, and raw data is typically non-transferable due to privacy and governance restrictions. Conventional centralised training, pooling all data to train a single model, is therefore rarely feasible, whilst standard distributed training within a single trust domain does not address the fundamental challenge of improving models when relevant knowledge is held independently by separate entities.

    The multimodal challenge

    The problem is further complicated by the inherently multimodal nature of modern marketing data. A single campaign may span visual brand assets, textual ad copy, structured audience segments, and cross-channel performance metrics. Multimodal learning, training models to reason jointly across such heterogeneous inputs, is already among the most demanding areas in machine learning. Under collaborative constraints, the challenge intensifies: each participating organisation may hold different modality combinations in varying formats and volumes, and a shared model must learn effectively from this fragmented landscape without any raw data ever leaving its source.

    Why use Federated Learning?

    Federated Learning (FL) offers a principled alternative by keeping data at the edge and exchanging only model updates. In the standard cross-silo formulation, each participant trains locally on its own private data, a central coordinator aggregates the resulting model parameters, and the process repeats over communication rounds until convergence. This report studies that pipeline in a multimodal classification setting, where each sample is a structured composition of contextual modalities and the task is ternary outcome prediction.

    The objective

    A fundamental question must be answered before committing to real-world deployment:

    Does federated learning perform well enough on multimodal marketing data to justify its tradeoffs?

    Centralised training retains an intrinsic advantage, full visibility into the data distribution at every optimisation step. The goal of this work is not to show that FL surpasses centralised performance, but to determine whether it approximates it closely enough for the privacy and collaboration benefits to be practically worthwhile. We further examine FL robustness under realistic stress conditions: scaling the number of participating nodes, introducing data noise, and investigating the complexity of relationships among different modalities. If this intersection of federated and multimodal learning can be made viable, it enables collaborative intelligence across organisations at a scale neither approach could achieve alone.

    To ensure a controlled, reproducible analysis, we generate multimodal datasets with a configurable synthetic generator and assign them to virtual nodes using identically independently distributed (IID)-like, homogeneous partitioning to simulate collaborative training. We implement the full FL loop in simulation mode using Flower, a state-of-the-art federated learning framework. The experiments that follow quantify how federation affects predictive quality relative to a centralised baseline, and how performance varies with (i) the number of nodes, (ii) dataset noise levels, and (iii) the prioritisation of cross-modal relationships in the synthetic data generation based on their frequencies.


    Handling the data

    Synthetic data generator

    For our experiments, we used proprietary multimodal synthetic datasets generated with an internally developed framework. The synthetic data was designed to replicate real-world marketing dynamics, where campaign outcomes are shaped by the interplay of multiple contextual factors (modalities). The framework provides precise control over the composition of each data instance, enabling rigorous evaluation of the proposed model under known conditions and full transparency into the factors driving campaign performance. The specific hyper parameters governing each dataset’s generation are detailed in the experimental sections below, where they are adjusted according to each experiment’s objectives.

    Each campaign instance within the dataset is characterised by five distinct modalities:

    • Audience – the target consumer segment,
    • Brand – the strategic positioning and perceptual attributes of the brand,
    • Creative – the tonal and messaging characteristics of the campaign,
    • Platform – the distribution channel on which the campaign is deployed, and
    • Geography – the market region(s) in which the campaign is activated.

    Each sample is assigned a categorical performance label — Positive (Overperforming), Negative (Underperforming), or Average (Performing within expected bounds) — indicating the projected campaign outcome for a given multimodal configuration. This ternary classification scheme enables the model to learn discriminative representations across the full spectrum of campaign effectiveness.s across the full spectrum of campaign effectiveness.

    Data partitioning

    After generating the training and test datasets with the synthetic data generator, the training data must be partitioned across federated clients to simulate a real-world scenario in which multiple companies collaboratively train the model, each holding its own proprietary data. In practice, this partitioning step would not be necessary, as each company would naturally possess its own local dataset. However, in our simulated environment, the codebase provides the partition-dataset script, which splits the training dataset across a user-defined number of nodes using either a homogeneous or heterogeneous partitioning strategy.

    Under homogeneous partitioning, the script groups training samples by label and divides each label’s indices into equal-sized segments across the specified number of clients. This ensures each node receives a statistically representative subset of the data with approximately uniform class proportions. Under heterogeneous partitioning, the script samples a Dirichlet distribution with concentration parameter alpha to determine the proportion of each class assigned to each node, producing non-IID label skew of configurable severity. Lower alpha values yield more extreme label imbalance across clients, whilst higher values produce distributions that progressively approach the homogeneous case. For our proof-of-concept research project, we focus our experiments on a homogeneous data split.

    Upon completion, the script outputs individual training files per node, a shared server test set, and a heat map visualising the label distribution across all partitions. An example of a homogeneous partition heat map across 5 nodes is shown below:

    Figure 1: Heat map showing homogeneous data partitioning across 5 clients in the federated learning scenario

    With the dataset now generated and partitioned into per-node training files, we can move from data preparation to the federated training setup. In the next section, we introduce Flower, the framework we use to orchestrate node selection, parameter exchange, and aggregation over these partitions.


    Flower: The Federated Learning framework

    What is Flower?

    Flower is an open-source, framework-agnostic federated learning framework that enables collaborative model training across decentralised data holders without requiring raw data to leave its source. It supports both real-world distributed deployment over gRPC and local simulation through a Virtual Client Engine (VCE) backed by the Ray distributed runtime. In simulation mode, nodes are virtualised as ephemeral objects instantiated on demand, allowing researchers to simulate federations of arbitrary size on a single machine with fine-grained resource control.

    Simulation cycle

    The simulation follows a centralised client-server architecture executed over iterative communication rounds. Each round proceeds through five sequential stages:

    Figure 2: Breakdown of a federated learning communication round into steps under the Flower framework.

    Each virtual node is instantiated at the beginning of its task, executes training or evaluation, returns results, and is immediately destroyed. This allows all five clients in this configuration to run concurrently without persistent memory allocation.

    The above diagram illustrates a complete FL communication round orchestrated by the Flower framework. The process begins with Client Selection, where the Flower Server selects a subset of participants from a broader pool. The server then distributes the current global model parameters to all selected clients simultaneously. Each client performs Local Training in parallel using Ray workers, training on its own private dataset to produce a locally updated model. These Local Updates are sent back to the server, which performs Federated Averaging (FedAvg). This is a weighted average, so clients with more data have proportionally more influence on the updated global model. A readable way to write the FedAvg update is:

    θglobal=∑k∈𝒦nkNθktheta_{text{global}} = sum_{k in mathcal{K}} frac{n_k}{N} , theta_k

    Where:

    • θglobal theta_{text{global}} is the updated global model parameters after aggregation.
    • 𝒦mathcal{K} is the set of clients selected to participate in this round.
    • θktheta_k is the model parameters after client kk finishes local training.
    • nkn_k is the number of training samples held by client kk.
    • N=∑k∈𝒦nkN = sum_{k in mathcal{K}} n_k is the total number of samples across the participating clients.

    Finally, a Federated Evaluation step assesses the updated global model across all clients and reports per-client accuracy metrics. This pipeline repeats across successive FL rounds until convergence.

    A key architectural detail of Flower’s VirtualClientEngine underpins this workflow: each virtual client is instantiated at the start of its task, runs training or evaluation, returns results, and is immediately destroyed. This enables all clients in the configuration to run concurrently without persistent memory allocation, making the system highly scalable even on resource-constrained hardware. By keeping raw data decentralised at the edge and exchanging only model parameters, this design preserves data privacy by default whilst still enabling collaborative model improvement across heterogeneous environments.

    Configuration reference

    The simulation is governed by a YAML configuration organized into five sections that correspond to Flower’s core components. Here, we present the YAML file with its default values.

    server:
      strategy: Mean
      fraction_fit: 1.0
      fraction_eval: 1.0
      num_rounds: 4
      server_dataset_included: false
    client:
      num_clients: 5
    model:
      name: GeneralisedJointModel_3cls_feat_drop
      hidden_embed_dim: 256
      embed_dim: 256
      hidden_dim: 256
      dropout_prob: 0.3
      modality_dropout_prob: 0
      local_epochs: 5
      batch_size: 64
      lr: 1e-3
      weight_decay: 1e-4
      optuna_n_trials: 3
    general:
      use_wandb: true
      random_seed: 42
      text_embedding_task: classification
      data_version: V28_noise_0_percent
      data_split: homogeneous
      alpha: 0.1
      one_hot_modalities:
    backend:
      client_resources:
        num_cpus: 2.0
        num_gpus: 0.0
    

    To run an experiment, the user only needs to adjust the desired parameters in this YAML file and execute the simulation with:

    poetry run simulation
    

    The framework reads the configuration, initialises all components accordingly, and executes the full federated learning cycle without any additional setup. Below, we explain what each section’s default values represent.

    Server configuration

    Defines the central orchestrator responsible for client coordination, parameter distribution, and aggregation.

    ParameterDefault ValueDescription
    strategyMeanAggregation strategy implementing Federated Averaging (FedAvg), where client model parameters are averaged weighted by each client’s local dataset size.
    fraction_fit1.0Fraction of clients selected for training each round. At 1.0, all clients participate in every round.
    fraction_eval1.0Fraction of clients selected for evaluation each round. At 1.0, all clients evaluate after each aggregation.
    num_rounds4Total number of federated communication rounds. Combined with 5 local epochs per round (default value), each data sample is exposed to 20 effective training epochs.
    server_dataset_includedfalseThe server holds no data partition and acts purely as a parameter aggregator.

    Client configuration

    Defines the federated client pool. Each client represents an independent data silo with its own private partition.

    ParameterDefault ValueDescription
    num_clients5Total number of virtual clients in the federation. With full participation, this constitutes a cross-silo setting with a small number of reliable, always-available participants.

    Model configuration

    Defines the neural network architecture and the local training hyper parameters applied on each client.

    Architecture parameters

    ParameterDefault ValueDescription
    nameGeneralisedJointModel_3cls_feat_dropA custom multimodal fusion model with a modality-agnostic joint embedding space, three classification heads for multi-task prediction, and feature-level dropout regularisation.
    hidden_embed_dim256Dimensionality of hidden layers within each modality-specific encoder, applied before the fusion stage.
    embed_dim256Dimensionality of the joint fused embedding space shared across all classification heads.
    hidden_dim256Dimensionality of hidden layers within each of the three classification heads.
    dropout_prob0.3Dropout probability applied in the classification heads and hidden layers.
    modality_dropout_prob0Probability of dropping entire modality branches during training. At 0, all modalities are always present (disabled).

    Training parameters

    ParameterDefault ValueDescription
    local_epochs5Number of complete passes over a client’s local dataset per communication round.
    batch_size64Mini-batch size for local stochastic gradient descent.
    lr1e-3Learning rate for the local optimiser.
    weight_decay1e-4L2 regularization coefficient penalising large weight magnitudes.
    optuna_n_trials3Number of Optuna Bayesian hyper parameter optimisation trials.

    Data configuration

    Controls dataset versioning, partitioning strategy, and modality-specific preprocessing.

    ParameterDefault ValueDescription
    data_versionV28_noise_0_percentDataset version identifier pointing to a specific preprocessed dataset.
    data_splithomogeneousPartitioning strategy across clients. Points to a homogeneous data distribution folder where data is split in an IID manner and each client receives a statistically representative partition with similar class distributions. The alternative heterogeneous points to a non-IID distribution folder where splits are governed by alpha.
    alpha0.1Dirichlet concentration parameter controlling non-IID severity when data_split points to a heterogeneous distribution folder. Currently inactive as the configuration uses the homogeneous split. When active, lower values produce more extreme label skew across clients.
    text_embedding_taskclassificationConfigures the text modality encoder for a classification objective, affecting pooling strategy and embedding optimisation.
    one_hot_modalitiesnullSpecifies modalities requiring one-hot encoding. Currently none.

    The data_split and alpha parameters in the simulation YAML configuration must match the partitioning strategy used when running the partition-dataset script, because the simulation reads client data from the output directory, whose folder name encodes both the split type and the number of clients.

    Simulation backend & experiment management

    Controls reproducibility, logging, and resource allocation for parallel execution of virtual clients through the Ray runtime.

    ParameterDefault ValueDescription
    use_wandbtrueEnables Weights & Biases experiment tracking for real-time metric logging and cross-experiment comparison.
    random_seed42Global seed ensuring reproducibility in our experiments.
    num_cpus2.0CPU cores reserved per virtual client task. Ray schedules a client only when the required cores are available.
    num_gpus0.0GPU allocation per client. At 0.0, all computation runs on CPU and concurrency is bounded solely by CPU availability.

    Experiments

    The following section outlines three experiments that compare federated learning with centralised training under different conditions.

    • Experiment 1 establishes a baseline comparison and tests how FL performance changes as the number of clients increases.
    • Experiment 2 examines how robust both approaches are when controlled noise is added to the training data.
    • Experiment 3 evaluates whether emphasising common vs. rare cross-modal relationships during synthetic data generation affects model performance.

    Overall, these experiments measure both the absolute performance of each approach and whether the performance gap between centralised and federated training remains consistent as the setting becomes more challenging.

    To ensure a fair comparative evaluation, all experimental variables were held constant across centralised and federated configurations. Both setups employ an identical model architecture, ensuring that any observed performance differences are attributable solely to the training paradigm rather than structural variations in the model itself.

    The total computational training budget was also standardised: the centralised model undergoes 20 sequential training epochs, whilst the federated configuration distributes this across 4 global communication rounds with each client performing 5 local epochs per round, yielding an equivalent total of 20 epochs. Furthermore, the training data is partitioned uniformly across all participating clients under a homogeneous assumption, ensuring that each client’s local dataset is a representative subset of the global distribution. This controlled partitioning ensures that any performance degradation observed with increasing client counts can be attributed to the effects of federation and aggregation at scale, rather than to statistical heterogeneity across client data.

    Experiment 1: Centralised vs. federated performance

    The data

    We employed the V16 synthetic dataset, generated using the synthetic data generator, comprising 84K training samples and 36K test samples. For both centralised and federated training configurations, raw input features, specifically Audience, Brand, Creative, Platform, and Geography, were encoded into dense representations using Vertex AI embeddings as a preprocessing step.

    The results

    Training ConfigurationScoreNegative F1Positive F1Average F1
    Centralised (Baseline)0.79670.70740.73050.9523
    FL with 5 clients0.76230.68410.65710.9458
    FL with 10 clients0.70290.57810.59020.9405
    FL with 15 clients0.67650.53450.55810.9368

    The centralised training configuration establishes the upper performance bound at a score of 0.7967. This outcome is theoretically expected, as the model benefits from unrestricted access to the complete dataset, without information loss due to partitioning or coordination overhead inherent in distributed paradigms. It therefore serves as the reference benchmark for all federated configurations.

    The federated learning results reveal a consistent and monotonic degradation in performance as the number of participating clients increases. With 5 clients, the model achieves a score of 0.7623 — a modest decline of approximately 3.5 points from the centralised baseline. However, scaling to 10 and 15 clients yields more substantial reductions to 0.7029 and 0.6765, respectively. This pattern is uniformly reflected across all evaluation metrics; however, the decline is most pronounced in the Negative F1 and Positive F1 scores, which degrade at a markedly steeper rate than Average F1. This suggests that class-specific discriminative performance is more sensitive to data partitioning than overall classification ability.

    This observed degradation is primarily attributable to the aggregation penalty. As the number of clients grows, the training corpus is divided into progressively smaller subsets, resulting in local model updates that are less representative of the global data distribution. The increased variance among these updates introduces noise during server-side aggregation, impeding convergence toward a robust global model.

    Lesson learned: FL with 5 clients comes remarkably close to centralised performance, showing that federated collaboration is viable with minimal accuracy loss. However, as the number of clients grows, makes it progressively harder for the global model to match centralised results.

    Experiment 2: Resilience to noisy data

    The data

    To conduct this experiment, it was necessary to generate synthetic data with controlled levels of noise. To understand what noise means in this context, it is important to first describe how the synthetic data is generated.

    The data generation process is grounded in a predefined graph structure. In this graph, nodes represent distinct values for each modality, namely Audience, Brand, Creative, Platform, and Geography, whilst edges encode the pairwise relationships between these values. Each edge carries a label of either Positive (indicating an over performing campaign) or Negative (indicating an underperforming campaign).

    The generator samples from this graph to produce a user-defined number of data points, subject to a set of hard constraints governed by configurable hyper parameters. Specifically, the user defines the desired number of samples for each target label: Positive, Negative, and Average. The generation of a single data sample proceeds as follows:

    1. Value Selection: One or more unique values are selected for each modality.
    2. Pairwise Evaluation: All pairwise combinations among the selected values are evaluated against the graph. Each combination is classified as positive, negative, or missing — the latter indicating that no edge exists between the two values in the graph.
    3. Proportion Calculation: The proportions of positive, negative, and missing combinations are computed relative to the total number of pairwise combinations.
    4. Label Assignment: These proportions are then compared against predefined acceptable ranges specified in the hyper parameters for each target label. If the proportions fall within the range defined for Positive, Negative, or Average, the sample is assigned the corresponding label. If the proportions do not satisfy any of the defined ranges, the sample is discarded and the generation process is repeated.

    A key question that arises from this process is: How are the predefined acceptable ranges for each target label determined?

    To address this, we conducted the following preliminary experiment. We randomly sampled 10,000 subgraphs, each comprising 1,000 edges, from the initial graph. For each subgraph, we computed the proportions of positive, negative, and missing pairwise combinations. From these 10,000 samples, we derived the mean and standard deviation among Positive, Negative, and Missing values. These statistics were then used to define the acceptable range for the Average target label, representing the typical composition of a randomly sampled subgraph.

    The acceptable ranges for the Positive and Negative target labels were subsequently defined by shifting the boundaries of the Average range along the respective axes. Specifically, the Positive range requires the proportion of positive combinations to exceed the Average upper bound by at least 5 standard deviations, and similarly, the Negative range requires the proportion of negative combinations to exceed the Average upper bound by the same margin. This ensures a clear statistical separation between the three label categories, such that samples assigned to the Positive or Negative class exhibit meaningfully distinct distributional characteristics from those labelled as Average.

    Based on the above methodology, we established the appropriate acceptable ranges for each target label. This, however, raises a subsequent question: What constitutes noise in this context?

    Figure 3: Impact of additive noise on the acceptable ranges for each target label in the synthetic data generator. As noise increases from zero to high, the opposing acceptable range for each sample’s target label progressively widens. This increases the acceptable proportion of negative combinations for the Positive label, positive combinations for the Negative label, and both equally for the Average label, thereby reducing the distributional separation between label categories.

    In our framework, noise is defined as the relaxation of the opposing acceptable range for a given target label. Specifically, introducing noise to the Positive target label corresponds to increasing its acceptable proportion of negative combinations — effectively reducing the degree of “positiveness” required for a sample to be classified as Positive. Conversely, adding noise to the Negative target label increases its acceptable proportion of positive combinations. For the Average target label, the additive noise is distributed equally across both the Positive and Negative acceptable ranges.

    This noise mechanism is applied at three levels of intervention — low, medium, and high — each progressively widening the acceptable range of the opposing value for a given target label. The figure above illustrates how the acceptable ranges for each target label are impacted under each level of intervention.

    To support this experiment, four synthetic datasets were generated, each comprising 84K training samples and 36K test samples:

    • Clean: No noise intervention applied.
    • Low Noise: Low-level relaxation of the opposing acceptable ranges.
    • Medium Noise: Medium-level relaxation of the opposing acceptable ranges.
    • High Noise: High-level relaxation of the opposing acceptable ranges.

    The federated learning simulation was configured with five participating clients, and performance was evaluated against the centralised baseline across all four dataset conditions.

    The Results

    Training ConfigurationScoreNegative F1Average F1Positive F1
    FL with no noise0.80740.80290.86330.7559
    Centralised with no noise0.82050.81650.86920.7760
    FL with low noise0.79230.79020.85550.7313
    Centralised with low noise0.81590.80390.87500.7686
    FL with medium noise0.78260.77180.85700.7189
    Centralised with medium noise0.80170.78760.86430.7533
    FL with high noise0.76090.73440.86250.6859
    Centralised with high noise0.78110.75510.86590.7221

    As anticipated, both the centralised and federated models achieve their highest performance on clean data and exhibit a gradual decline as noise levels increase. At the highest noise intervention, the centralised model’s score decreases from 0.8205 to 0.7811, whilst the federated model’s score declines from 0.8074 to 0.7609 — representing drops of approximately 3.9 and 4.7 percentage points, respectively.

    Notably, neither model exhibits catastrophic degradation under any noise condition. Even at the highest level of intervention, both configurations maintain reasonable performance. The most pronounced declines are observed in the Positive F1 and Negative F1 scores, which is consistent with the noise injection methodology described above: since noise is introduced by relaxing the opposing acceptable range for each target label, the boundaries between Positive and Negative classes become increasingly blurred, making these the most challenging distinctions for the model. In contrast, the Average F1 remains remarkably stable across all noise levels for both configurations, indicating that the models’ capacity to capture general distributional patterns is largely unaffected by the introduced noise.

    Consistent with the findings from Experiment 1, the centralised model maintains a performance advantage over the federated configuration at every noise level. However, the magnitude of this gap remains approximately constant across all noise conditions. This observation is significant: it indicates that the federated setup does not exhibit increased sensitivity to noisy data relative to its centralised counterpart. The performance differential between the two paradigms is attributable to the aggregation penalty discussed in Experiment 1, rather than to any compounding effect of noise on the federated training process.

    Lesson learned: Real-world data is inherently noisy, and any viable model must be able to handle that. Both centralised and FL models show strong resilience, performance declines gradually rather than breaking down, even when the data is heavily corrupted. Importantly, FL’s relative performance holds steady across noise levels, suggesting it is no more vulnerable to messy data than centralised training.

    Experiment 3: Impact of cross-modal relationships under synthetic data generation

    The data

    Leveraging the synthetic data generator enables the investigation of additional structural characteristics of the initial marketing graph — specifically, the edge types representing pairwise relationships between modalities. Understanding which modality pairs (e.g., Audience–Brand or Creative–Geography) are most influential on model performance is of particular interest.

    To this end, we conducted a preliminary analysis: 10,000 subgraphs, each comprising 1,000 edges, were sampled from the initial multimodal graph, and the mean and standard deviation of the observed edge-type frequencies were computed. The table below presents the modality relationships ranked by frequency, from the most common to the most rare.

    Modality RelationshipMean Frequency
    Brand to Content8.48
    Audience to Content7.29
    Content to Geography6.57
    Audience to Brand5.90
    Audience to Geography5.84
    Brand to Geography5.55
    Content to Content3.24
    Brand to Brand3.18
    Content to Platform1.88
    Brand to Platform1.52
    Audience to Audience1.18
    Audience to Platform1.16

    This frequency distribution informed the design of a subsequent performance-based experiment. The synthetic data generator exposes two relevant hyper parameters: a high pair preference, which increases the likelihood of sampling edges from the specified modality relationships, and a low pair preference, which suppresses them. Using these controls, three synthetic datasets were generated under distinct configurations:

    • Common First: The two highest-frequency modality relationships are assigned as the high pair, and the two lowest-frequency relationships as the low pair.
    • Rare First: The inverse configuration, where the two lowest-frequency relationships are assigned as the high pair and the two highest as the low pair.
    • Middle Ground: The four middle-ranked relationships from the frequency table are assigned to the high and low pairs accordingly.

    All remaining hyper parameters were held constant across the three configurations: noise levels were set to zero, and the federated learning simulation was conducted with five participating clients, consistent with the setup described in prior experiments. The training size remains 84k, as the test size is equal to 36K.

    The Results

    Training ConfigurationScoreNegative F1Average F1Positive F1
    Centralised Common First0.88030.88560.92450.8308
    FL Common First0.87480.87750.91500.8320
    Centralised Rare First0.94410.95030.95790.9242
    FL Rare First0.93430.94700.95050.9054
    Centralised Middle Ground0.88130.87440.91430.8551
    FL Middle Ground0.86250.86770.90090.8188

    The results reveal a notable disparity in performance across the three dataset configurations. The Rare First configuration substantially outperforms the other two, achieving scores of 0.9441 (Centralised) and 0.9343 (FL) — a margin of approximately 6–8 percentage points over the Common First and Middle Ground configurations, which yield scores in the 0.86–0.88 range. This performance advantage is consistently reflected across all evaluation metrics, with particularly pronounced gains in Positive F1, where the Rare First configuration achieves 0.9242 (Centralised) and 0.9054 (FL), compared to values in the 0.81–0.85 range for the alternative configurations.

    This finding is counterintuitive yet theoretically interpretable. Frequently occurring modality combinations, by virtue of their prevalence, contribute comparatively less discriminative information to the learning process — the decision boundaries they define are, in effect, already well-represented and easily separable. In contrast, rare combinations compel the model to learn more nuanced and distinctive feature interactions, resulting in richer decision boundaries between Positive and Negative campaign outcomes. The learning signal provided by atypical patterns is therefore disproportionately more informative per sample.

    Consistent with findings from prior experiments, the centralised model maintains a modest performance advantage over the federated configuration across all three dataset strategies. Crucially, however, the relative ranking of dataset configurations remains identical under both training paradigms: Rare First consistently outperforms Common First and Middle Ground, regardless of whether training is conducted centrally or in a federated manner.

    Lesson learned: Not all data is equally valuable. Prioritising rare, atypical feature combinations produces significantly better models than focusing mostly on common patterns. This has direct implications for how we design synthetic datasets: rather than mimicking the most typical marketing dynamics, we should deliberately include uncommon combinations to give the model a richer and more discriminative learning signal.


    Impact and future directions

    This work represents an initial investigation into the viability of federated learning within our operational context. The finding that centralised model performance degrades only marginally under a reasonable number of participating clients opens a promising avenue for delivering machine learning solutions that address shared industry challenges among organisations reluctant to pool their data. The federated learning paradigm enables multiple entities to collaboratively train a shared global model on their respective proprietary datasets, without exposing raw data at any stage of the training process, thereby mitigating the risk of data leakage.


    Although Federated Learning has been an established collaborative learning paradigm since its introduction in 2017, it remains a highly active area of research in academia and a strategic priority for industrial adoption. Our initial findings establish the foundation for continued exploration, with future work organised around the following directions:

    1. Privacy guarantees in adversarial federated environments

    Whilst FL enables organisations to collaborate on a shared model without moving or centralising raw data, it’s important to be clear about the remaining risk surface: the exchanged model updates must be handled securely. Extensive literature shows that, in adversarial settings, updates can be targeted by malicious participants or exposed by a compromised coordinator. In practice, this is addressed with a defence baseline, e.g. secure aggregation, privacy protections, and integrity monitoring, so partners can benefit from FL’s collaboration gains whilst maintaining strong privacy and trust throughout training.

    2. Evaluation under advanced and realistic federated scenarios

    Whilst simulating collaborative training with uniformly distributed data provides a valuable baseline for foundational FL research, it does not fully capture the complexities inherent in real-world deployments. Future work will extend our preliminary investigations into data heterogeneity, building upon the noise-injection experiments conducted on synthetic datasets in this study. Additionally, we intend to evaluate the efficacy of maintaining a shared synthetic dataset on the central server as a reference benchmark for assessing the integrity of incoming model updates and detecting potentially malicious contributions. Finally, we plan to transition from the simulated FL environment currently facilitated by the Flower framework to a fully distributed architecture. By deploying distinct computational nodes to represent separate organisational entities, we aim to empirically investigate and address the communication bottlenecks inherent in practical federated deployments.

    Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP AI Lab team.