Author: Simon Philp

  • Embedding Alignment for Profile Matching 

    Can two companies collaborate on audience sizing and enrichment without sharing raw customer data, and without having to use the same AI embedding model? In this initial phase, we tested whether independent customer embedding spaces can be aligned using a geometric transformation guided by a small sample of shared customer profiles. We found that models built on similar foundations align almost seamlessly, and even entirely different models reliably place customers into their correct behavioural neighbourhoods for effective audience discovery. Furthermore, this alignment remains robust even with very few shared customer records. These findings prove that cross-model audience matching is viable in practice, establishing a strong foundation for our next goal: aligning spaces without needing any shared customers at all.

    (more…)

  • Are open-source embedding models good enough? A comparative study

    Throughout the years, a very large number of embedding models have emerged, each one having different strengths and weaknesses. We set out to answer a practical question many data professionals face often:

    Do premium embedding models significantly outperform free, open-source alternatives when predicting social media success?

    Figure 1

    In our latest research, we examined the performance of various embedding models across two distinct tasks. Our findings suggest that whilst premium models from industry leaders like OpenAI and Google technically produce the optimal results for most cases, the margin of victory is surprisingly narrow. Our main takeaway? For many predictive marketing tasks and Retrieval-Augmented Generation (RAG) tasks, free open alternatives are often good enough.

    The methodology and model selection

    To test the capabilities of different embeddings, we designed two evaluation tasks:

    1. Predictive downstream modelling: We generated embeddings from our data sources and fed them into a downstream LightGBM model to predict the final target (such as post popularity).
    2. Information retrieval: We compared the embeddings inside a RAG framework using a vector index to measure direct context retrieval performance.

    We tested seven different embedding models to get a comprehensive view of the landscape:

    ModelProviderBase ArchitectureCategory
    text-embedding-005GoogleProprietaryPremium (Paid API)
    text-embedding-3-largeOpenAIProprietaryPremium (Paid API)
    thenlper/gte-baseAlibabaBERTOpen-Source
    BAAI/bge-base-en-v1.5BAAIBERTOpen-Source
    all-mpnet-base-v2SBERTMPNet (Microsoft)Open-Source
    all-roberta-large-v1SBERTRoBERTa (Meta AI)Open-Source
    all-miniLM-L6-v2SBERTMiniLM (Microsoft)Open-Source
    Embedding models

    These open-source alternatives were primarily chosen to correlate with the embedding dimensions of the premium baselines, while specifically including a smaller (all-miniLM-L6-v2) and a larger (all-roberta-large-v1) alternative to see how model size impacted performance.


    Experiment 1: Predicting Instagram post popularity

    For our first experiment, we used the public Instagram Influencer Dataset, which consists of various posts from different online influencers, to predict post popularity.

    Since the raw data is highly visual, we first used Gemini to generate detailed textual descriptions of each post. We then converted these descriptions into embeddings using our seven chosen models, which were then used to train our downstream LightGBM model.

    The results: OpenAI’s text-embedding-3-large produced the best overall results with an average R² score of 0.475, closely followed by Google’s text-embedding-005 at 0.470. However, the smallest model we tested (all-miniLM-L6-v2) still achieved a score of 0.440. This competitive showing from the open-source models is particularly impressive when you consider the potential “family alignment” advantage in the workflow, where descriptions generated by Gemini might naturally favour Google’s own embedding model. Despite premium models taking the lead, the R² performance gap between the best and worst models was a mere 0.035.


    Experiment 2: The SMP challenge image dataset

    To validate our initial findings, we applied the exact same framework to the image dataset from the Social Media Prediction (SMP) Challenge, which was instead used to predict popularity of Flickr posts.

    The results: Once again, the premium models topped the charts, with both OpenAI and Google producing an identical R² score of 0.234. Just like our Instagram experiment, the weakest free alternative trailed by that same narrow margin, coming in at 0.196.


    Experiment 3: Retrieval-augmented generation (RAG) performance

    To see how these findings hold up beyond downstream regression tasks, we introduced a third experiment: a traditional Retrieval-Augmented Generation (RAG) evaluation.

    At its core, a RAG framework acts as an open-book exam for an LLM. Instead of relying solely on its pre-trained internal knowledge, the system first searches an external database to retrieve the most relevant documents matching a user’s prompt. It then passes these documents alongside the question to the LLM, ensuring the final generated response is accurate, contextually grounded, and factual.

    Using the industry-standard BEIR (SciFact) dataset, we indexed 5,180+ scientific documents and evaluated how effectively each embedding model could retrieve the exact context needed to answer 200 distinct queries.

    We measured two key retrieval metrics:

    • Hit rate: The percentage of queries where the correct document was successfully retrieved in the top 5 results.
    • Mean reciprocal rank (MRR): A measure of where the correct document ranked (closer to 1 is better).

    The results: This experiment provided an interesting twist. Whilst OpenAI’s premium text-embedding-3-large achieved the highest MRR (0.741), the open-source gte-base model proved remarkably competitive, securing a strong second-place MRR of 0.729 and comfortably beating Google’s premium offering (0.692).

    When it came to the overall Hit Rate, the open-source alternative stood completely shoulder-to-shoulder with the premium giants. Instead of a clear winner, we saw a three-way tie at 86% between gte-base, OpenAI, and Google.


    Results overview: Scores across multiple experiments

    To ensure a robust evaluation across various datasets, we measured performance using the R² score averaged over multiple random training splits, allowing us to establish a reliable variance margin (±) for each model.

    In practice, this means instead of training our LightGBM model just once on a single slice of data, we shuffled and split the dataset into different training and testing sets multiple times. Doing this ensures that a model’s high score wasn’t just a fluke resulting from a “lucky” data split. The resulting variance margin (±) acts like an error bar: a tighter margin indicates the model is highly stable, telling us exactly how consistent and reliable its predictions will be when exposed to entirely new data.

    ModelSize CategoryExp 1: Instagram (R²)Exp 2: SMP Challenge (R²)Exp 3: RAG (MRR)Exp 3: RAG (Hit Rate)
    text-embedding-3-large (OpenAI)Premium / Baseline0.475 ± 0.0220.234 ± 0.0240.7410.86
    text-embedding-005 (Google)Premium / Baseline0.470 ± 0.0250.234 ± 0.0230.6920.86
    thenlper/gte-baseBaseline Match0.458 ± 0.0180.212 ± 0.0270.7290.86
    BAAI/bge-base-en-v1.5Baseline Match0.457 ± 0.0200.217 ± 0.0220.6880.84
    all-mpnet-base-v2Baseline Match0.446 ± 0.0210.216 ± 0.0270.6140.74
    all-roberta-large-v1Larger Alternative0.440 ± 0.0210.196 ± 0.0200.5940.71
    all-miniLM-L6-v2Smaller Alternative0.439 ± 0.0100.197 ± 0.0150.5960.75
    Experiment results
    Experiment results

    The true cost of performance: Latency and pricing

    Whilst the minimal gap in predictive accuracy alone makes a compelling case for utilising open-source alternatives, factoring in latency and execution costs makes the decision even clearer.

    To illustrate this, we tracked the time and money spent to generate embeddings for each experiment. For reference, the dataset was of size 11,500 rows for Experiment 1, size 10,000 rows for Experiment 2 and 5,180 documents with 200 queries for Experiment 3.

    When analysing costs, it is important to separate the costs into two main categories: the cost of calling each embedding model’s API and the cost of the underlying compute resources (such as a virtual machine or a local laptop) required to process the embeddings. For our experiments, we ran everything on a standard Virtual Machine (VM). We have excluded that infrastructure cost from this breakdown, as it fluctuates wildly depending on a developer’s specific deployment preferences and scaling needs.

    ModelExp 1: Token FeesExp 1: Latency (Seconds)Exp 2: Token FeesExp 2: Latency (Seconds)Exp 3: Token FeesExp 3: Latency (Seconds)
    text-embedding-3-large (OpenAI)$1.7775$0.5141$0.22298
    text-embedding-005 (Google)$1.2978$0.4150$0.17241
    all-roberta-large-v1Free512Free438Free3,777
    BAAI/bge-base-en-v1.5Free344Free247Free2,229
    all-mpnet-base-v2Free282Free240Free1,990
    thenlper/gte-baseFree84Free60Free11,420
    all-miniLM-L6-v2Free43Free30Free251
    Latency and costs

    The latency vs. cost trade-off: Who wins on speed?

    Whilst open-source models eliminate recurring API costs entirely, a common concern is whether self-hosting sacrifices processing speed. To find out, we timed how long each model took to generate the embeddings across our experiments.

    To keep the comparison fair, we ran all the open-source models on a standard Google Cloud Platform (GCP) virtual machine: an n1-standard-8 instance (8 vCPUs, 30 GB memory) equipped with a single NVIDIA T4 GPU.

    The results revealed a highly competitive landscape with massive implications for production pipelines:

    1. In small-to-medium tasks, light open source dominates

    For standard datasets like Experiment 1 (Instagram) and Experiment 2 (SMP Challenge), the lightweight open-source models proved that you don’t need to pay for a premium API to get blazing-fast speeds:

    • The absolute champion: The tiny all-miniLM-L6-v2 absolutely shredded the competition, completing Experiment 2 in a mere 30 seconds (and Experiment 1 in 43 seconds), beating Google’s premium API by up to 40% whilst costing $0 in token fees.
    • Premium APIs are fast, but costly: Google’s text-embedding-005 and OpenAI’s text-embedding-3-large blazed through Experiment 2 in 50 and 41 seconds respectively, but carrying token bills of $0.41 and $0.51. Scaled across millions of rows, those micro-transactions add up quickly.

    2. At scale (Experiment 3), the infrastructure tax emerges

    When we scaled up to the larger RAG retrieval dataset in Experiment 3, the dynamics shifted heavily, highlighting the core trade-off of hosting your own models:

    • Premium APIs pull ahead on massive batches: Google’s model completed the entire RAG dataset in just 241 seconds (costing $0.17), whilst OpenAI finished in 298 seconds ($0.22). Because their infrastructure is globally distributed and massively parallelised, they handle large batches effortlessly.
    • The self-hosting bottleneck: Whilst all-miniLM-L6-v2 stayed nimble at 251 seconds, heavier open-source models struggled on our single-GPU setup. For instance, thenlper/gte-base, our RAG accuracy champion, took over 11,400 seconds (more than 3 hours) to complete the run on our T4 GPU.

    The Takeaway: When budgeting for your pipeline, remember to separate API token costs from your underlying infrastructure costs (the VM compute time). Model latency serves as an excellent proxy for your infrastructure bill as the longer a self-hosted model runs, the longer your VM has to operate.

    If you are running light-to-medium real-time tasks, optimised open-source models like all-miniLM-L6-v2 give you double savings: zero API fees and lower VM uptime. But if you’re processing massive, enterprise-scale batches and don’t want to invest in scaling a heavy local cluster of GPUs, paying a premium API fee might be the more cost-effective route.


    Key findings and conclusion

    Our experiments across predictive marketing tasks and search retrieval pipelines point to a nuanced conclusion:

    • Premium edges out open-source in predictive modelling, but only just: Both OpenAI’s text-embedding-3-large and Google’s text-embedding-005 consistently delivered the optimal results across our data splits. However, the difference in predictive performance between the premium giants and top open-source models remains minimal, highlighting that free alternatives are highly capable.
    • Open-source matches premium performance in RAG hit rates: For semantic search and context retrieval, the free open-source gte-base model achieved an identical 86% Hit Rate to both OpenAI and Google, demonstrating that free options can deliver matching retrieval coverage. However, OpenAI’s premium offering did secure the crown for tighter search precision via the highest Mean Reciprocal Rank (MRR) of 0.741, with gte-base following closely behind at 0.729.
    • The speed vs. cost trade-off has shifted: Highly optimised premium APIs have closed the speed gap, actually outpacing mid-sized open-source models like gte-base. Whilst open-source still completely dominates on cost (being 100% free), the absolute speed crown belongs specifically to hyper-lightweight open-source models like all-miniLM-L6-v2, which outran everything at 30 seconds.
    • Size isn’t everything: Interestingly, large open models like all-roberta-large-v1 underperformed across the board compared to optimised, smaller ones like gte-base.

    Our study focused on specific, mainstream tasks, providing strong evidence that free and open-source models can often be enough to get the job done effectively for standard applications.

    However, it is important to acknowledge that frontier embedding models, as well as massive open-weight models, still hold distinct advantages. For highly demanding use cases requiring massive context windows, multilingual support, or nuanced reasoning across obscure domains, state-of-the-art models consistently prove their worth. A quick look at popular industry benchmarks, such as Hugging Face’s Massive Text Embedding Benchmark (MTEB) leaderboard, clearly demonstrates this. The frontrunners at the very top of those charts are continually pushing the boundaries of what’s possible across dozens of highly specialised datasets.

    Ultimately, depending on whether your priority is:

    • Squeezing out the absolute highest KPI value and conquering highly complex edge cases
    • Minimising cost and latency whilst maintaining competitive results

    You now have a highly capable spectrum of both premium and open-source models to choose from.

    Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.