{"id":1823,"date":"2026-07-03T12:29:29","date_gmt":"2026-07-03T12:29:29","guid":{"rendered":"https:\/\/cms.research.wpp.com\/?post_type=research_feed&#038;p=1823"},"modified":"2026-07-03T12:34:21","modified_gmt":"2026-07-03T12:34:21","slug":"are-open-source-embedding-models-good-enough-a-comparative-study","status":"publish","type":"research_feed","link":"https:\/\/cms.research.wpp.com\/?research_feed=are-open-source-embedding-models-good-enough-a-comparative-study","title":{"rendered":"Are open-source embedding models good enough? A comparative study"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Throughout the years, a very large number of embedding models have emerged, each one having different strengths and weaknesses. We set out to answer a practical question many data professionals face often:<\/strong><\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Do premium embedding models significantly outperform free, open-source alternatives when predicting social media success?<\/strong><\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"572\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Gemini_Generated_Image_yix5z6yix5z6yix5-1024x572.jpg\" alt=\"\" class=\"wp-image-1832\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Gemini_Generated_Image_yix5z6yix5z6yix5-1024x572.jpg 1024w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Gemini_Generated_Image_yix5z6yix5z6yix5-300x167.jpg 300w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Gemini_Generated_Image_yix5z6yix5z6yix5-768x429.jpg 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Gemini_Generated_Image_yix5z6yix5z6yix5-1536x857.jpg 1536w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/Gemini_Generated_Image_yix5z6yix5z6yix5.jpg 1926w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption class=\"wp-element-caption\">Figure 1<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">In our latest research, we examined the performance of various embedding models across two distinct tasks. Our findings suggest that whilst premium models from industry leaders like OpenAI and Google technically produce the optimal results for most cases, the margin of victory is surprisingly narrow. Our main takeaway? For many predictive marketing tasks and Retrieval-Augmented Generation (RAG) tasks, free open alternatives are often good enough.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The methodology and model selection<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To test the capabilities of different embeddings, we designed two evaluation tasks:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Predictive downstream modelling<\/strong>: We generated embeddings from our data sources and fed them into a downstream <strong>LightGBM<\/strong> model to predict the final target (such as post popularity).<\/li>\n\n\n\n<li><strong>Information retrieval:<\/strong> We compared the embeddings inside a RAG framework using a vector index to measure direct context retrieval performance.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">We tested seven different embedding models to get a comprehensive view of the landscape:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th><strong>Model<\/strong><\/th><th><strong>Provider<\/strong><\/th><th><strong>Base Architecture<\/strong><\/th><th><strong>Category<\/strong><\/th><\/tr><\/thead><tbody><tr><td><code>text-embedding-005<\/code><\/td><td>Google<\/td><td>Proprietary<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#595bff\" class=\"has-inline-color\">Premium (Paid API)<\/mark><\/strong><\/td><\/tr><tr><td><code>text-embedding-3-large<\/code><\/td><td>OpenAI<\/td><td>Proprietary<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#595bff\" class=\"has-inline-color\">Premium (Paid API)<\/mark><\/strong><\/td><\/tr><tr><td><code>thenlper\/gte-base<\/code><\/td><td>Alibaba<\/td><td>BERT<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#54a473\" class=\"has-inline-color\">Open-Source<\/mark><\/strong><\/td><\/tr><tr><td><code>BAAI\/bge-base-en-v1.5<\/code><\/td><td>BAAI<\/td><td>BERT<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#54a473\" class=\"has-inline-color\">Open-Source<\/mark><\/strong><\/td><\/tr><tr><td><code>all-mpnet-base-v2<\/code><\/td><td>SBERT<\/td><td>MPNet (Microsoft)<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#54a473\" class=\"has-inline-color\">Open-Source<\/mark><\/strong><\/td><\/tr><tr><td><code>all-roberta-large-v1<\/code><\/td><td>SBERT<\/td><td>RoBERTa&nbsp;(Meta AI)<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#54a473\" class=\"has-inline-color\">Open-Source<\/mark><\/strong><\/td><\/tr><tr><td><code>all-miniLM-L6-v2<\/code><\/td><td>SBERT<\/td><td>MiniLM&nbsp;(Microsoft)<\/td><td><strong><mark style=\"background-color:rgba(0, 0, 0, 0);color:#54a473\" class=\"has-inline-color\">Open-Source<\/mark><\/strong><\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Embedding models<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">These open-source alternatives were primarily chosen to correlate with the embedding dimensions of the premium baselines, while specifically including a smaller (<code>all-miniLM-L6-v2<\/code>) and a larger (<code>all-roberta-large-v1<\/code>) alternative to see how model size impacted performance.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h3 class=\"wp-block-heading\">Experiment 1: Predicting Instagram post popularity<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For our first experiment, we used the public <a href=\"https:\/\/sites.google.com\/site\/sbkimcv\/dataset\/instagram-influencer-dataset\">Instagram Influencer Dataset<\/a>, which consists of various posts from different online influencers, to predict post popularity.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Since the raw data is highly visual, we first used Gemini to generate detailed textual descriptions of each post. We then converted these descriptions into embeddings using our seven chosen models, which were then used to train our downstream LightGBM model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The results:<\/strong> OpenAI\u2019s <code>text-embedding-3-large<\/code> produced the best overall results with an average R\u00b2 score of <strong>0.475<\/strong>, closely followed by Google\u2019s <code>text-embedding-005<\/code> at 0.470. However, the smallest model we tested (<code>all-miniLM-L6-v2<\/code>) still achieved a score of <strong>0.440<\/strong>. This competitive showing from the open-source models is particularly impressive when you consider the potential &#8220;family alignment&#8221; advantage in the workflow, where descriptions generated by Gemini might naturally favour Google&#8217;s own embedding model. Despite premium models taking the lead, the R\u00b2 performance gap between the best and worst models was a mere 0.035.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h3 class=\"wp-block-heading\">Experiment 2: The SMP challenge image dataset<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">To validate our initial findings, we applied the exact same framework to the image dataset from the <a href=\"https:\/\/smp-challenge.com\/\">Social Media Prediction (SMP) Challenge<\/a>, which was instead used to predict popularity of Flickr posts.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The results:<\/strong> Once again, the premium models topped the charts, with both OpenAI and Google producing an identical R\u00b2 score of <strong>0.234<\/strong>. Just like our Instagram experiment, the weakest free alternative trailed by that same narrow margin, coming in at <strong>0.196<\/strong>.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h3 class=\"wp-block-heading\">Experiment 3: Retrieval-augmented generation (RAG) performance<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">To see how these findings hold up beyond downstream regression tasks, we introduced a third experiment: a traditional Retrieval-Augmented Generation (RAG) evaluation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At its core, a RAG framework acts as an open-book exam for an LLM. Instead of relying solely on its pre-trained internal knowledge, the system first searches an external database to retrieve the most relevant documents matching a user&#8217;s prompt. It then passes these documents alongside the question to the LLM, ensuring the final generated response is accurate, contextually grounded, and factual.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Using the industry-standard <a href=\"https:\/\/huggingface.co\/datasets\/BeIR\/scifact\">BEIR (SciFact)<\/a> dataset, we indexed 5,180+ scientific documents and evaluated how effectively each embedding model could retrieve the exact context needed to answer 200 distinct queries.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We measured two key retrieval metrics:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Hit rate:<\/strong> The percentage of queries where the correct document was successfully retrieved in the top 5 results.<\/li>\n\n\n\n<li><strong>Mean reciprocal rank (MRR):<\/strong> A measure of <em>where<\/em> the correct document ranked (closer to 1 is better).<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The results:<\/strong> This experiment provided an interesting twist. Whilst OpenAI&#8217;s premium&nbsp;<code>text-embedding-3-large<\/code>&nbsp;achieved the highest MRR (0.741), the open-source&nbsp;<code>gte-base<\/code>&nbsp;model proved remarkably competitive, securing a strong second-place MRR of 0.729 and comfortably beating Google&#8217;s premium offering (0.692).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When it came to the overall Hit Rate, the open-source alternative stood completely shoulder-to-shoulder with the premium giants. Instead of a clear winner, we saw a three-way tie at 86% between&nbsp;<code>gte-base<\/code>, OpenAI, and Google.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Results overview: Scores across multiple experiments<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To ensure a robust evaluation across various datasets, we measured performance using the R\u00b2 score averaged over multiple random training splits, allowing us to establish a reliable variance margin (\u00b1) for each model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In practice, this means instead of training our LightGBM model just once on a single slice of data, we shuffled and split the dataset into different training and testing sets multiple times. Doing this ensures that a model&#8217;s high score wasn&#8217;t just a fluke resulting from a &#8220;lucky&#8221; data split. The resulting variance margin (\u00b1) acts like an error bar: a tighter margin indicates the model is highly stable, telling us exactly how consistent and reliable its predictions will be when exposed to entirely new data.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Model<\/th><th>Size Category<\/th><th>Exp 1: Instagram (R\u00b2)<\/th><th>Exp 2: SMP Challenge (R\u00b2)<\/th><th>Exp 3: RAG (MRR)<\/th><th>Exp 3: RAG (Hit Rate)<\/th><\/tr><\/thead><tbody><tr><td><code>text-embedding-3-large (OpenAI)<\/code><\/td><td>Premium \/ Baseline<\/td><td><strong>0.475 \u00b1 0.022<\/strong><\/td><td><strong>0.234 \u00b1 0.024<\/strong><\/td><td><strong>0.741<\/strong><\/td><td><strong>0.86<\/strong><\/td><\/tr><tr><td><code>text-embedding-005 (Google)<\/code><\/td><td>Premium \/ Baseline<\/td><td>0.470 \u00b1 0.025<\/td><td><strong>0.234 \u00b1 0.023<\/strong><\/td><td>0.692<\/td><td><strong>0.86<\/strong><\/td><\/tr><tr><td><code>thenlper\/gte-base<\/code><\/td><td>Baseline Match<\/td><td>0.458 \u00b1 0.018<\/td><td>0.212 \u00b1 0.027<\/td><td>0.729<\/td><td><strong>0.86<\/strong><\/td><\/tr><tr><td><code>BAAI\/bge-base-en-v1.5<\/code><\/td><td>Baseline Match<\/td><td>0.457 \u00b1 0.020<\/td><td>0.217 \u00b1 0.022<\/td><td>0.688<\/td><td>0.84<\/td><\/tr><tr><td><code>all-mpnet-base-v2<\/code><\/td><td>Baseline Match<\/td><td>0.446 \u00b1 0.021<\/td><td>0.216 \u00b1 0.027<\/td><td>0.614<\/td><td>0.74<\/td><\/tr><tr><td><code>all-roberta-large-v1<\/code><\/td><td>Larger Alternative<\/td><td>0.440 \u00b1 0.021<\/td><td>0.196 \u00b1 0.020<\/td><td>0.594<\/td><td>0.71<\/td><\/tr><tr><td><code>all-miniLM-L6-v2<\/code><\/td><td>Smaller Alternative<\/td><td>0.439 \u00b1 0.010<\/td><td>0.197 \u00b1 0.015<\/td><td>0.596<\/td><td>0.75<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Experiment results<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"805\" height=\"1024\" src=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-805x1024.png\" alt=\"\" class=\"wp-image-1846\" srcset=\"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-805x1024.png 805w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-236x300.png 236w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-768x977.png 768w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-1207x1536.png 1207w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-1609x2048.png 1609w, https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/07\/embedding_model_performance_universal-2-scaled.png 2011w\" sizes=\"auto, (max-width: 805px) 100vw, 805px\" \/><figcaption class=\"wp-element-caption\">Experiment results<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">The true cost of performance: Latency and pricing<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Whilst the minimal gap in predictive accuracy alone makes a compelling case for utilising open-source alternatives, factoring in latency and execution costs makes the decision even clearer.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To illustrate this, we tracked the time and money spent to generate embeddings for each experiment. For reference, the dataset was of size 11,500 rows for Experiment 1, size 10,000 rows for Experiment 2 and 5,180 documents with 200 queries for Experiment 3.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When analysing costs, it is important to separate the costs into two main categories: the cost of calling each embedding model&#8217;s API and the cost of the underlying compute resources (such as a virtual machine or a local laptop) required to process the embeddings. For our experiments, we ran everything on a standard Virtual Machine (VM). We have excluded that infrastructure cost from this breakdown, as it fluctuates wildly depending on a developer&#8217;s specific deployment preferences and scaling needs.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Model<\/th><th>Exp 1: Token Fees<\/th><th>Exp 1: Latency (Seconds)<\/th><th>Exp 2: Token Fees<\/th><th>Exp 2: Latency (Seconds)<\/th><th>Exp 3: Token Fees<\/th><th>Exp 3: Latency (Seconds)<\/th><\/tr><\/thead><tbody><tr><td><code>text-embedding-3-large (OpenAI)<\/code><\/td><td>$1.77<\/td><td>75<\/td><td>$0.51<\/td><td>41<\/td><td>$0.22<\/td><td>298<\/td><\/tr><tr><td><code>text-embedding-005 (Google)<\/code><\/td><td>$1.29<\/td><td>78<\/td><td>$0.41<\/td><td>50<\/td><td>$0.17<\/td><td>241<\/td><\/tr><tr><td><code>all-roberta-large-v1<\/code><\/td><td>Free<\/td><td>512<\/td><td>Free<\/td><td>438<\/td><td>Free<\/td><td>3,777<\/td><\/tr><tr><td><code>BAAI\/bge-base-en-v1.5<\/code><\/td><td>Free<\/td><td>344<\/td><td>Free<\/td><td>247<\/td><td>Free<\/td><td>2,229<\/td><\/tr><tr><td><code>all-mpnet-base-v2<\/code><\/td><td>Free<\/td><td>282<\/td><td>Free<\/td><td>240<\/td><td>Free<\/td><td>1,990<\/td><\/tr><tr><td><code>thenlper\/gte-base<\/code><\/td><td>Free<\/td><td>84<\/td><td>Free<\/td><td>60<\/td><td>Free<\/td><td>11,420<\/td><\/tr><tr><td><code>all-miniLM-L6-v2<\/code><\/td><td>Free<\/td><td>43<\/td><td>Free<\/td><td>30<\/td><td>Free<\/td><td>251<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">Latency and costs<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">The latency vs. cost trade-off: Who wins on speed?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Whilst open-source models eliminate recurring API costs entirely, a common concern is whether self-hosting sacrifices processing speed. To find out, we timed how long each model took to generate the embeddings across our experiments.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To keep the comparison fair, we ran all the open-source models on a standard Google Cloud Platform (GCP) virtual machine: an <strong>n1-standard-8 instance (8 vCPUs, 30 GB memory) equipped with a single NVIDIA T4 GPU<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The results revealed a highly competitive landscape with massive implications for production pipelines:<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">1. In small-to-medium tasks, light open source dominates<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For standard datasets like Experiment 1 (Instagram) and Experiment 2 (SMP Challenge), the lightweight open-source models proved that you don&#8217;t need to pay for a premium API to get blazing-fast speeds:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>The absolute champion:<\/strong> The tiny <code>all-miniLM-L6-v2<\/code> absolutely shredded the competition, completing Experiment 2 in a mere <strong>30 seconds<\/strong> (and Experiment 1 in <strong>43 seconds<\/strong>), beating Google&#8217;s premium API by up to 40% whilst costing $0 in token fees.<\/li>\n\n\n\n<li><strong>Premium APIs are fast, but costly:<\/strong> Google&#8217;s <code>text-embedding-005<\/code> and OpenAI&#8217;s <code>text-embedding-3-large<\/code> blazed through Experiment 2 in <strong>50<\/strong> and <strong>41 seconds<\/strong> respectively, but carrying token bills of $0.41 and $0.51. Scaled across millions of rows, those micro-transactions add up quickly.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">2. At scale (Experiment 3), the infrastructure tax emerges<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">When we scaled up to the larger RAG retrieval dataset in Experiment 3, the dynamics shifted heavily, highlighting the core trade-off of hosting your own models:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Premium APIs pull ahead on massive batches:<\/strong> Google&#8217;s model completed the entire RAG dataset in just <strong>241 seconds<\/strong> (costing $0.17), whilst OpenAI finished in <strong>298 seconds<\/strong> ($0.22). Because their infrastructure is globally distributed and massively parallelised, they handle large batches effortlessly.<\/li>\n\n\n\n<li><strong>The self-hosting bottleneck:<\/strong> Whilst <code>all-miniLM-L6-v2<\/code> stayed nimble at <strong>251 seconds<\/strong>, heavier open-source models struggled on our single-GPU setup. For instance, <code>thenlper\/gte-base<\/code>, our RAG accuracy champion, took over <strong>11,400 seconds<\/strong> (more than 3 hours) to complete the run on our T4 GPU.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The Takeaway:<\/strong> When budgeting for your pipeline, remember to separate API token costs from your underlying infrastructure costs (the VM compute time). Model latency serves as an excellent proxy for your infrastructure bill as the longer a self-hosted model runs, the longer your VM has to operate.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you are running light-to-medium real-time tasks, optimised open-source models like <code>all-miniLM-L6-v2<\/code> give you double savings: zero API fees and lower VM uptime. But if you&#8217;re processing massive, enterprise-scale batches and don&#8217;t want to invest in scaling a heavy local cluster of GPUs, paying a premium API fee might be the more cost-effective route.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">Key findings and conclusion<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Our experiments across predictive marketing tasks and search retrieval pipelines point to a nuanced conclusion:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Premium edges out open-source in predictive modelling, but only just:<\/strong>&nbsp;Both OpenAI\u2019s&nbsp;<code>text-embedding-3-large<\/code>&nbsp;and Google\u2019s&nbsp;<code>text-embedding-005<\/code>&nbsp;consistently delivered the optimal results across our data splits. However, the difference in predictive performance between the premium giants and top open-source models remains minimal, highlighting that free alternatives are highly capable.<\/li>\n\n\n\n<li><strong>Open-source matches premium performance in RAG hit rates:<\/strong>&nbsp;For semantic search and context retrieval, the free open-source&nbsp;<code>gte-base<\/code>&nbsp;model achieved an identical 86% Hit Rate to both OpenAI and Google, demonstrating that free options can deliver matching retrieval coverage. However, OpenAI&#8217;s premium offering did secure the crown for tighter search precision via the highest Mean Reciprocal Rank (MRR) of 0.741, with&nbsp;<code>gte-base<\/code>&nbsp;following closely behind at 0.729.<\/li>\n\n\n\n<li><strong>The speed vs. cost trade-off has shifted:<\/strong>&nbsp;Highly optimised premium APIs have closed the speed gap, actually outpacing mid-sized open-source models like&nbsp;<code>gte-base<\/code>. Whilst open-source still completely dominates on cost (being 100% free), the absolute speed crown belongs specifically to hyper-lightweight open-source models like&nbsp;<code>all-miniLM-L6-v2<\/code>, which outran everything at <strong>30 seconds<\/strong>.<\/li>\n\n\n\n<li><strong>Size isn&#8217;t everything:<\/strong>&nbsp;Interestingly, large open models like&nbsp;<code>all-roberta-large-v1<\/code>&nbsp;underperformed across the board compared to optimised, smaller ones like&nbsp;<code>gte-base<\/code>.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Our study focused on specific, mainstream tasks, providing strong evidence that free and open-source models can often be enough to get the job done effectively for standard applications.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, it is important to acknowledge that frontier embedding models, as well as massive open-weight models, still hold distinct advantages. For highly demanding use cases requiring massive context windows, multilingual support, or nuanced reasoning across obscure domains, state-of-the-art models consistently prove their worth. A quick look at popular industry benchmarks, such as Hugging Face&#8217;s Massive Text Embedding Benchmark (<a href=\"https:\/\/huggingface.co\/spaces\/mteb\/leaderboard\">MTEB<\/a>) leaderboard, clearly demonstrates this. The frontrunners at the very top of those charts are continually pushing the boundaries of what&#8217;s possible across dozens of highly specialised datasets.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ultimately, depending on whether your priority is:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Squeezing out the absolute highest KPI value and conquering highly complex edge cases<\/li>\n\n\n\n<li>Minimising cost and latency whilst maintaining competitive results<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">You now have a highly capable spectrum of both premium and open-source models to choose from.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Are premium embedding models worth the price? We benchmarked seven premium and open-source models on predictive marketing and RAG tasks to analyze the true tradeoffs between cost, speed, and accuracy. Read on to find out when free alternatives are &#8220;good enough&#8221; for your production pipeline.<\/p>\n","protected":false},"author":7,"featured_media":0,"template":"","meta":{"_acf_changed":false,"_ppma_block_editor_authors":""},"tags":[],"content_types":[{"id":50,"name":"Blog Post","slug":"article"}],"ppma_author":[{"id":7,"display_name":"Eirini Kolimatsi","first_name":"Eirini","last_name":"Kolimatsi","nickname":"eirini.kolimatsi","user_nicename":"eirini-kolimatsi","user_email":"eirini.kolimatsi@satalia.com","biographical_info":"Eirini is a Data Scientist at Satalia with a multidisciplinary background in Management Science and Computer Science. She specialises in architecting end-to-end data science solutions, leveraging a deep technical toolkit to solve complex industrial challenges across diverse sectors. Known for bridging the gap between theoretical research and scalable application, she focuses on delivering high-impact models that translate abstract data patterns into actionable strategic intelligence.\r\n\r\nHer current research focuses on sophisticated campaign performance multimodal modelling and the development of data enrichment frameworks to maximise predictive accuracy.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/03\/DSC_0100.jpg","job_title":"Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":2},{"id":8,"display_name":"Jael Freixanet","first_name":"Jael","last_name":"Freixanet","nickname":"jael.freixanet","user_nicename":"jael-freixanet","user_email":"jael.freixanet@satalia.com","biographical_info":"Jael is a Data Scientist at Satalia, leveraging her physics background for a deep analytical foundation in complex systems analysis and modelling. Her experience, spanning foundational research in computational physics and a proven track record in data science consultancy, provides a unique perspective for architecting robust, scalable models in intricate environments. In Satalia's Research Lab, she bridges scientific methodology with industrial innovation to address WPP\u2019s most sophisticated data challenges. Her current research focuses on multimodal fusion models, aiming to improve campaign performance and pioneer state-of-the-art machine learning.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/03\/jael_img.jpeg","job_title":"Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":2},{"id":36,"display_name":"Simon Philp","first_name":"Simon","last_name":"Philp","nickname":"simon.philp","user_nicename":"simon-philp","user_email":"simon.philp@satalia.com","biographical_info":"Simon is a Data Scientist with experience applying Bayesian methods and predictive modeling to design innovative, data-driven solutions for complex operational challenges. His background spans machine learning, deep learning, and advanced statistical analysis, alongside leading impactful AI initiatives that drive organizational efficiency. During his time as a Research Assistant with the Met Office, he utilized Bayesian techniques to model environmental patterns, bringing together rigorous research and practical, real-world application.\r\n\r\nSimon holds a Master's degree with Distinction in Data Science and Analytics from the University of Leeds and a First-Class Honours degree in Mathematics from Durham University. This strong academic and research foundation supports his work in translating complex mathematical concepts into robust, scalable technology that automates processes and delivers measurable value.","avatar_url":"https:\/\/cms.research.wpp.com\/wp-content\/uploads\/2026\/06\/author.jpeg","job_title":"Data Scientist","is_lead":false,"display_as_researcher":true,"order_priority":null}],"class_list":["post-1823","research_feed","type-research_feed","status-publish","hentry","content_type-article"],"acf":{"content":"","content_quarter":"","related_pods":[1362]},"research_categories":[],"raw_acf":{"content":"","content_quarter":"","related_pods":["1362"],"featured":"","legacy_perspective_source_id":""},"_links":{"self":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed\/1823","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_feed"}],"about":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/types\/research_feed"}],"author":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/users\/7"}],"acf:post":[{"embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=\/wp\/v2\/research_pods\/1362"}],"wp:attachment":[{"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1823"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1823"},{"taxonomy":"content_type","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcontent_types&post=1823"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/cms.research.wpp.com\/index.php?rest_route=%2Fwp%2Fv2%2Fppma_author&post=1823"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}