{"id":1879,"date":"2026-07-16T11:06:47","date_gmt":"2026-07-16T11:06:47","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=1879"},"modified":"2026-07-16T11:28:53","modified_gmt":"2026-07-16T11:28:53","slug":"does-structural-similarity-improve-transfer-learning-across-business-domains","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=does-structural-similarity-improve-transfer-learning-across-business-domains","title":{"rendered":"Does structural similarity improve transfer learning across business domains?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Transfer learning is widely used to adapt large language models (LLMs) to new tasks by leveraging knowledge acquired during previous training. Rather than fine-tuning a model from scratch for every new dataset, a common strategy is to first adapt the model to a related dataset before continuing training on the target task. If the source and target domains share useful structure, the resulting model may provide a better starting point for optimisation than the original foundation model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For organisations working across multiple clients, brands, or business units, this is an attractive proposition. <a href=\"https:\/\/research.wpp.com\/pods\/campaign-performance-modelling-pod\">In marketing applications, models are often trained<\/a> on campaign data collected across different organisations, products, and customer segments in order to estimate the success potential of a future marketing campaign. Although these campaigns span different products, audiences, and markets, they typically share the same prediction objective: estimating how a campaign will perform. This raises an opportunity: if the underlying relationships that drive performance are partly shared across business domains, then a model adapted to one domain should carry useful structure into another, making it a stronger starting point than the foundation model alone.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The objective of this experiment is therefore to answer the following question:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><em>&#8220;Does sequential fine-tuning on one business domain improve performance on another, and how does the relationship between the two domains influence the results?&#8221;<\/em><\/p>\n<\/blockquote>\n\n\n\n<h2 class=\"wp-block-heading\">Our datasets<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To study how well models transfer across business environments, we construct two synthetic business domains &#8211; <strong>Domain A and Domain B<\/strong> &#8211; using the graph-based <a href=\"https:\/\/research.wpp.com\/reports\/synthetic-dataset-generation-pod-technical-walkthrough\"><strong>Synthetic Data Generator<\/strong> <\/a>developed in <em>WPP Research<\/em>. The generator first constructs a signed knowledge graph by estimating the compatibility between pairs of marketing attributes using LLMs, historical campaign data, domain expertise, or a combination of these. This graph then serves as the foundation for generating synthetic campaign datasets. Its nodes represent possible values of five key campaign factors: <strong>Audience<\/strong>, <strong>Brand<\/strong>, <strong>Creative<\/strong>, <strong>Platform<\/strong>, and <strong>Geography<\/strong>. For example:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Audience<\/strong><\/th><th><strong>Gen Z fitness enthusiasts (18\u201324)<\/strong><\/th><\/tr><\/thead><tbody><tr><td><strong>Brand<\/strong><\/td><td>AcmeSports<\/td><\/tr><tr><td><strong>Creative<\/strong><\/td><td>15-second vertical video: a fast-cut montage of a pre-dawn city run, set to an upbeat trap beat and ending on a bold &#8220;Always running beside you&#8221; text overlay with a product close-up. Energetic and aspirational<\/td><\/tr><tr><td><strong>Platform<\/strong><\/td><td>TikTok<\/td><\/tr><tr><td><strong>Geography<\/strong><\/td><td>Los Angeles metro area (Southern California)<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\"><em>Example campaign instance represented as structured marketing attributes used by the synthetic data generator.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Each edge in the graph encodes the compatibility between two values. For instance, how well a given audience matches a given brand. A campaign is then represented as a subgraph: a subset of nodes drawn from the broader graph, together with the distribution of compatibility values on the edges connecting them. The campaign&#8217;s performance label (its outcome) is then derived from those edge-level compatibility values.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In our experiments, we consider three discrete performance outcomes for a campaign:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Underperforming<\/li>\n\n\n\n<li>Average Performing<\/li>\n\n\n\n<li>Overperforming<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The underlying graph can be configured to represent different business domains. For instance, <strong>&#8220;Gen Z fitness enthusiasts&#8221;<\/strong> may be a strong target audience for <strong>AcmeSports<\/strong> in one domain and a poor match in another. This is achieved by modifying the compatibility values assigned to the graph&#8217;s edges. Because the graph contains thousands of nodes and tens of thousands of edges, it can represent a wide range of business domains. The compatibility values themselves may be inferred by LLMs, learned from historical campaign data, defined by domain experts, or derived from a combination of these sources.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Figure 2<\/em> below shows samples from the two distinct knowledge graphs that were used to generate the two domain datasets that we use in our experiments: Domain A and Domain B.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"406\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.22.00-1024x406.png\" alt=\"\" class=\"wp-image-1887\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.22.00-1024x406.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.22.00-300x119.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.22.00-768x304.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.22.00-1536x609.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.22.00-2048x812.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><em>Independent business-domain graphs used to generate Domain A and Domain B. Each graph defines modality-specific campaign attributes and signed pairwise relationships that determine whether sampled campaign configurations are labelled as underperforming, average performing, or overperforming.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">To test how sensitive our models are to small, structured changes within a single domain, we also build a modified version of Domain A, called&nbsp;Domain A-10% flip. We start from the original Domain-A graph and make one controlled change: we randomly pick 10% of the nodes and flip the sign of every relationship connected to them, so what used to be a good match becomes a bad one, and vice versa.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The nodes of the knowledge graph thus stay identical, while modifying some of the edges. The result is a domain that looks identical on the surface but has had a handful of its underlying &#8220;rules&#8221; quietly reversed.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"400\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.25.08-1024x400.png\" alt=\"\" class=\"wp-image-1888\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.25.08-1024x400.png 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.25.08-300x117.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.25.08-768x300.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.25.08-1536x601.png 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-16-at-12.25.08-2048x801.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><em>Original Domain A graph and its 10% node-flipped perturbation. The perturbed variant preserves the same schema and node space, but inverts all edge signs incident to a randomly selected 10% of nodes, creating a controlled structural shift within the same domain.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Importantly, <strong>Domain A-10% flip is not an independent business domain<\/strong>, but a structured perturbation of Domain A that preserves its underlying generation mechanism while altering relational semantics.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Together, these three datasets (A, B , A10%) define two complementary types of distribution shift:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Cross-domain shift:<\/strong> Domain A \u2194 Domain B, where the underlying graphs are independently generated.<\/li>\n\n\n\n<li><strong>Structural shift:<\/strong> Domain A \u2194 Domain A-10% flip, where the underlying graph is shared except for a small percentage of perturbed edges.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This setup enables controlled comparison of transfer behaviour under cross-domain shift and structured perturbation, while holding everything else constant.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Distributional validation of the synthetic domains<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Before evaluating transfer learning, we first verify that the three datasets exhibit the intended structural relationships. Although Domain A, Domain B, and Domain A-10% flip are generated from different graph configurations, we expect Domain A to remain much closer to Domain A-10% flip than to the independently generated Domain B.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To quantify this, we compare the distributions of the datasets&#8217; internal representations produced by the <strong>base (un-tuned) LLM model <\/strong>before any fine-tuning. Specifically, each campaign is converted into the natural-language prompt described above, passed through the pre-trained base model, and its final hidden-layer representation is extracted. We then compute the <strong>spherical sliced Wasserstein distance <\/strong>between the resulting sets of embeddings. Comparing datasets in this semantic feature space allows us to measure how similarly the foundation model represents the underlying campaign distributions, rather than comparing the raw campaign descriptions directly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The Wasserstein distance compares two distributions by measuring the minimum &#8220;cost&#8221; required to transform one into the other. A common analogy is to imagine each distribution as a pile of earth: the distance corresponds to the least amount of work needed to move the earth so that one pile exactly matches the other, where the work depends on both how much material is moved and how far it must travel. Smaller values therefore indicate that two datasets have more similar overall distributions, while larger values indicate greater structural differences.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Computing the exact Wasserstein distance in high-dimensional spaces is computationally expensive. The spherical sliced Wasserstein distance provides an efficient approximation by repeatedly projecting the high-dimensional embeddings onto many random one-dimensional directions, computing the Wasserstein distance for each projection, and averaging the results. Increasing the number of projections produces a more stable estimate of the true distance while remaining computationally tractable.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"641\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/ai-edited-image-1024x641.jpg\" alt=\"\" class=\"wp-image-1883\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/ai-edited-image-1024x641.jpg 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/ai-edited-image-300x188.jpg 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/ai-edited-image-768x481.jpg 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/ai-edited-image.jpg 1304w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\"><em>Spherical sliced Wasserstein distance from Domain A to Domain A-10% flip and Domain B across projection counts. Lower distances to Domain A-10% flip indicate that the source domain remains geometrically closer to its perturbed dataset than to the independently generated Domain B, supporting the observed transfer asymmetry.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The results confirm the intended construction of the datasets. Across all projection counts, <strong>Domain A-10% flip consistently exhibits a substantially smaller Wasserstein distance from Domain A than Domain B does<\/strong>. This is expected because Domain A-10% flip is produced by introducing controlled perturbations into the original Domain A graph, whereas Domain B is generated independently from a different underlying graph.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This validation establishes an objective measure of structural similarity before any model training takes place. The experiments that follow therefore test whether datasets that are closer in the foundation model&#8217;s representation space also exhibit stronger transfer learning during sequential fine-tuning.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Data pre-processing<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Before training, all datasets are processed using a unified preprocessing pipeline to ensure consistency across experiments. Each campaign instance is converted into a structured natural-language prompt describing its attributes across all modalities. The input format is fixed as:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>Platform: Twitter\nBrand: Premium outdoor apparel brand\nAudience: Young professionals interested in fitness and travel\nCreative: Humorous product-focused storytelling\nGeo: &#91;Mexico, USA, Germany]<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Each dataset is then split into training and test partitions using a <strong>70\/30 stratified split<\/strong> to preserve class balance. Where required, controlled oversampling is applied to address class imbalance in the training set. Importantly, test sets remain untouched to ensure unbiased evaluation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">During training, each example is formatted as an instruction-following task where the model receives the structured campaign description and is trained to <strong>predict the corresponding performance label<\/strong>. At evaluation time, the label is omitted and the model must infer it directly from the input prompt. This ensures a consistent task formulation across all datasets and training stages.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Our LLM of choice<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">All experiments are performed using <strong>Qwen3-4B-Base<\/strong> (<a href=\"https:\/\/huggingface.co\/unsloth\/Qwen3-4B-Base\">unsloth\/Qwen3-4B-Base<\/a>), adapted with <strong>Rank-Stabilised Low-Rank Adaptation<\/strong> (<a href=\"https:\/\/arxiv.org\/abs\/2607.09757\">rsLoRA<\/a>). The base model weights remain frozen, while trainable low-rank matrices are inserted into selected attention and MLP projection layers.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Model Component<\/th><th>Configuration<\/th><\/tr><\/thead><tbody><tr><td>Base model<\/td><td>unsloth\/Qwen3-4B-Base<\/td><\/tr><tr><td>Method<\/td><td>rsLoRA<\/td><\/tr><tr><td>Precision<\/td><td>4-bit quantisation<\/td><\/tr><tr><td>LoRA rank<\/td><td>16<\/td><\/tr><tr><td>LoRA alpha<\/td><td>16<\/td><\/tr><tr><td>Dropout<\/td><td>0.0<\/td><\/tr><tr><td>Target modules<\/td><td>Attention + MLP projections<\/td><\/tr><tr><td>Optimiser<\/td><td>AdamW (8-bit)<\/td><\/tr><tr><td>Learning rate<\/td><td>1e-4<\/td><\/tr><tr><td>Scheduler<\/td><td>Cosine decay<\/td><\/tr><tr><td>Warmup<\/td><td>10 steps<\/td><\/tr><tr><td>Weight decay<\/td><td>0.01<\/td><\/tr><tr><td>Batch size<\/td><td>32<\/td><\/tr><tr><td>Gradient accumulation<\/td><td>1<\/td><\/tr><tr><td>Training<\/td><td>1 epoch per dataset<\/td><\/tr><tr><td>Seed<\/td><td>3407<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\"><em>Fine-tuning configuration used across all experiments. The same Qwen3-4B-Base model, rsLoRA setup, optimiser, learning schedule, and training hyperparameters are held fixed so that observed differences can be attributed to dataset ordering and distribution shift rather than model or optimisation changes.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The model is further configured as a <strong>three-class classification system<\/strong>, where the output space is restricted to a single token representing one of:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>0 \u2192 Underperforming<\/li>\n\n\n\n<li>1 \u2192 Average Performing<\/li>\n\n\n\n<li>2 \u2192 Overperforming<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This is implemented by restricting the prediction head to a three-token classification vocabulary, ensuring the model performs direct categorical prediction rather than free-form text generation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">All experiments begin from the same pre-trained checkpoint with a newly initialised LoRA adapter, ensuring that dataset ordering is the only experimental variable.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Experimental design<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Every experiment ends by <strong>evaluating&nbsp;Macro F1 on the held-out Domain A test set<\/strong>, so that all training histories are compared on the same target task and the same untouched evaluation data. The only thing that varies across runs is the&nbsp;<em>sequence of datasets<\/em>&nbsp;the model is fine-tuned on before that evaluation. This isolates the effect of prior adaptation from any change in model, optimiser, or evaluation protocol.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We compare five training histories, chosen to probe the two forms of distribution shift defined earlier:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Condition<\/strong><\/th><th><strong>Training sequence<\/strong><\/th><th><strong>Shift probed<\/strong><\/th><\/tr><\/thead><tbody><tr><td>Baseline<\/td><td>Domain A<\/td><td>none (reference)<\/td><\/tr><tr><td>Cross-domain transfer<\/td><td>Domain B \u2192 Domain A<\/td><td>cross-domain, source seen first<\/td><\/tr><tr><td>Cross-domain forgetting<\/td><td>Domain A \u2192 Domain B<\/td><td>cross-domain, target seen first<\/td><\/tr><tr><td>Structural forgetting<\/td><td>Domain A \u2192 Domain A-10% flip<\/td><td>structural, target seen first<\/td><\/tr><tr><td>Structural transfer<\/td><td>Domain A-10% flip \u2192 Domain A<\/td><td>structural, source seen first<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\"><em>Sequential fine-tuning histories used to compare baseline training, cross-domain transfer and forgetting, and structurally related transfer and forgetting.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The&nbsp;<strong>baseline<\/strong>&nbsp;(train on A only) is the reference every sequential history is measured against: sequential fine-tuning is only worthwhile if it beats training directly on the target.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The three <strong>&#8220;source \u2192 A&#8221;<\/strong> conditions (baseline, B \u2192 A, A-10% flip \u2192 A) all&nbsp;<em>end<\/em>&nbsp;on Domain A and therefore test the core question directly: <em>does pre-adapting to another domain give a better starting point for Domain A than the foundation model alone, and does it matter whether that source domain is independently generated (B) or a structural perturbation of A (A-10% flip)?<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The two &#8220;A \u2192 other&#8221; conditions (A \u2192 B, A \u2192 A-10% flip) instead measure the flip side: how much Domain A knowledge survives once the model is subsequently adapted to a different domain. This quantifies catastrophic forgetting and lets us separate the&nbsp;<em>recency<\/em>&nbsp;of a domain in the training sequence from the&nbsp;<em>similarity<\/em>&nbsp;between domains.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Each history is run at&nbsp;<strong>1 epoch and 5 epochs per stage<\/strong>, holding all other hyperparameters fixed (Figure 5). Comparing the two lets us check whether any transfer benefit is robust to longer adaptation, or whether additional training amplifies overfitting and forgetting.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The results<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Figure 7 reports Macro F1 on the Domain A test set for each training history at 1 and 5 epochs. Three findings stand out.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"806\" height=\"501\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Training-sequentially-for-1-5-epochs-1-2.png\" alt=\"\" class=\"wp-image-1884\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Training-sequentially-for-1-5-epochs-1-2.png 806w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Training-sequentially-for-1-5-epochs-1-2-300x186.png 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Training-sequentially-for-1-5-epochs-1-2-768x477.png 768w\" sizes=\"auto, (max-width: 806px) 100vw, 806px\" \/><figcaption class=\"wp-element-caption\"><em>Comparison of the F1 score obtained on the held-out Domain A test set after different sequential training histories.<\/em><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>First, the domain seen&nbsp;<em>last<\/em>&nbsp;dominates performance.<\/strong>&nbsp;The strongest predictor of Domain A performance is simply whether Domain A was the final training stage. Every history ending on A scores in the 0.69\u20130.72 range (baseline 0.7079, B \u2192 A 0.7120, A-10% flip \u2192 A 0.7188 at 1 epoch), whereas every history ending on a&nbsp;<em>different<\/em>&nbsp;domain collapses: A \u2192 B falls to 0.5794 and A \u2192 A-10% flip to 0.6264. In other words, adapting the model to any other domain after A substantially erases what it learned about A, classic catastrophic forgetting. This is the single largest effect in the experiment and frames how the transfer results should be read: sequential fine-tuning only helps the target task when the target is the final stage.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Second, structural similarity transfers better than cross-domain similarity.<\/strong>&nbsp;Among the histories that end on A, the ordering at 1 epoch is A-10% flip \u2192 A (0.7188) &gt; B \u2192 A (0.7120) &gt; baseline (0.7079). Both forms of pre-adaptation give a small lift over training on A alone, but the structurally similar source (A-10% flip, which shares Domain A&#8217;s graph and node space and differs only in a fraction of edge signs) gives the largest gain. This is the direct answer to the research question:&nbsp;<em>the relationship between source and target matters, and structural similarity yields the most useful starting point.<\/em>&nbsp;The effect is consistent with the forgetting story above\u2014because A-10% flip shares most of Domain A&#8217;s structure, adapting to it first leaves the model closer to a good Domain A solution than adapting to an independently generated domain does.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Third, more epochs do not help transfer, and can hurt it.&nbsp;<\/strong>Training the source domain longer does not strengthen transfer. Direct training on A improves marginally with more epochs (0.7079 \u2192 0.7197), and structural transfer is essentially flat (0.7188 \u2192 0.7186), but cross-domain transfer&nbsp;<em>degrades<\/em>: B \u2192 A drops from 0.7120 at 1 epoch to 0.6904 at 5 epochs, below the baseline. Longer adaptation to an independent domain appears to push the model further from a good Domain A initialisation, so the extra training spent on the source has to be undone during the target stage. The forgetting conditions show the same trend (A \u2192 B: 0.5794 \u2192 0.5570; A \u2192 A-10% flip: 0.6264 \u2192 0.6178): more epochs on the &#8220;wrong&#8221; final domain make the forgetting worse.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Taken together,<\/strong> sequential fine-tuning across&nbsp;<em>independent<\/em>&nbsp;business domains offers little practical benefit for the target task: the gain over direct training is at best marginal at 1 epoch (+0.4 F1 points for B \u2192 A) and turns negative with more training. The benefit only becomes clear, and only comes with no downside at 5 epochs, when the source domain is&nbsp;<em>structurally<\/em>&nbsp;related to the target (A-10% flip \u2192 A matches the direct baseline while starting from a related domain). The results therefore support a nuanced answer to the opening question: structural similarity does improve transfer, but the dominant factor governing target performance is which domain the model is adapted to last.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A few caveats worth stating explicitly:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>These are single-seed runs and the gaps among the &#8220;ends-on-A&#8221; conditions are small (within ~1 F1 point), so the transfer&nbsp;<em>benefit<\/em>&nbsp;should be read as &#8220;at best modest and never harmful when the source is structurally similar,&#8221; rather than a large improvement.<\/li>\n\n\n\n<li>The&nbsp;<em>forgetting<\/em>&nbsp;effects, by contrast, are large and unambiguous. Repeating the key conditions across multiple seeds, and reporting the same histories evaluated on the Domain B test set, would let us confirm the transfer ranking and check whether the recency effect is symmetric across domains.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Final takeaway<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Our experiments point to a single dominant factor:&nbsp;<strong>under shared parameterisation, sequential fine-tuning behaves as domain adaptation, not knowledge accumulation.<\/strong>&nbsp;Whichever domain the model sees last governs its performance; adapting to any new domain after Domain A erased much of what it had learned (Macro F1 fell from ~0.71 to 0.58\u20130.63), regardless of how the earlier stages were arranged.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Within that constraint,&nbsp;<strong>structural similarity does help.<\/strong>&nbsp;Pre-adapting to a structurally related source (Domain A-10% flip, which shares Domain A&#8217;s graph and differs only in a fraction of edge signs) gave the best starting point for Domain A and matched direct training even at five epochs, whereas pre-adapting to an independently generated domain (Domain B) offered at best a marginal gain at one epoch and actively hurt with more training. <strong>This pattern mirrors the distributional validation:<\/strong> Domain A-10% flip was also substantially closer to Domain A than Domain B in the foundation model&#8217;s representation space. Transfer therefore tracks the alignment between the source and target relational structures: greater similarity yields a better initialisation, but it never overrides the interference introduced by later adaptation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The practical implication is that sequential fine-tuning across independent business domains is unlikely to build a cumulatively stronger model, each stage reshapes the network toward its own domain rather than adding to previous knowledge. It is worthwhile mainly when the source domain is structurally close to the target, and even then the target must be trained last. Accumulating knowledge across genuinely distinct domains would require mechanisms that explicitly protect earlier representations (e.g., separate adapters per domain, replay, or regularisation against forgetting) rather than a single shared adapter adapted in sequence.<\/p>\n\n\n\n<p class=\"is-style-default has-small-font-size wp-block-paragraph\"><sub>Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.<\/sub><\/p>\n","protected":false},"excerpt":{"rendered":"<p>TL;DR: Sequential fine-tuning does not automatically build a stronger model across business domains. Performance is driven mainly by whichever domain the model is adapted to most recently, meaning later training can overwrite earlier learning. However, transfer works better when the source and target datasets share the same underlying structure, even if some relationships have been perturbed. In practice, fine-tuning is most useful when the earlier dataset is structurally close to the target, and the target domain is trained last.<\/p>\n","protected":false},"author":10,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":""},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":10,"display_name":"Elektra Papazoglou","first_name":"Elektra","last_name":"Papazoglou","nickname":"elektra.papazoglou","user_nicename":"elektra-papazoglou","user_email":"elektra.papazoglou@satalia.com","biographical_info":"Elektra Papazoglou is a machine learning engineer at Satalia, working in the Research Lab on NLP and LLM systems. Her previous work has focused on large-scale information extraction from unstructured data, with an emphasis on combining prompt engineering and efficient fine-tuning to build methods that are both performant and practical.\r\nShe brings prior experience building production ML systems across recommendation, experimentation, and content understanding, bridging the gap between research and deployment.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/04\/image-e1776864019813.png","job_title":"Senior Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":4}],"class_list":["post-1879","research_feed","type-research_feed","status-publish","hentry","content_type-article"],"acf":{"content":"","content_quarter":"","related_pods":[136]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"","related_pods":["136"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/1879","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/10"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/136"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1879"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1879"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=1879"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=1879"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}