Throughout the years, a very large number of embedding models have emerged, each one having different strengths and weaknesses. We set out to answer a practical question many data professionals face often:
Do premium embedding models significantly outperform free, open-source alternatives when predicting social media success?

In our latest research, we examined the performance of various embedding models across two distinct tasks. Our findings suggest that whilst premium models from industry leaders like OpenAI and Google technically produce the optimal results for most cases, the margin of victory is surprisingly narrow. Our main takeaway? For many predictive marketing tasks and Retrieval-Augmented Generation (RAG) tasks, free open alternatives are often good enough.
The methodology and model selection
To test the capabilities of different embeddings, we designed two evaluation tasks:
- Predictive downstream modelling: We generated embeddings from our data sources and fed them into a downstream LightGBM model to predict the final target (such as post popularity).
- Information retrieval: We compared the embeddings inside a RAG framework using a vector index to measure direct context retrieval performance.
We tested seven different embedding models to get a comprehensive view of the landscape:
| Model | Provider | Base Architecture | Category |
|---|---|---|---|
text-embedding-005 | Proprietary | Premium (Paid API) | |
text-embedding-3-large | OpenAI | Proprietary | Premium (Paid API) |
thenlper/gte-base | Alibaba | BERT | Open-Source |
BAAI/bge-base-en-v1.5 | BAAI | BERT | Open-Source |
all-mpnet-base-v2 | SBERT | MPNet (Microsoft) | Open-Source |
all-roberta-large-v1 | SBERT | RoBERTa (Meta AI) | Open-Source |
all-miniLM-L6-v2 | SBERT | MiniLM (Microsoft) | Open-Source |
These open-source alternatives were primarily chosen to correlate with the embedding dimensions of the premium baselines, while specifically including a smaller (all-miniLM-L6-v2) and a larger (all-roberta-large-v1) alternative to see how model size impacted performance.
Experiment 1: Predicting Instagram post popularity
For our first experiment, we used the public Instagram Influencer Dataset, which consists of various posts from different online influencers, to predict post popularity.
Since the raw data is highly visual, we first used Gemini to generate detailed textual descriptions of each post. We then converted these descriptions into embeddings using our seven chosen models, which were then used to train our downstream LightGBM model.
The results: OpenAI’s text-embedding-3-large produced the best overall results with an average R² score of 0.475, closely followed by Google’s text-embedding-005 at 0.470. However, the smallest model we tested (all-miniLM-L6-v2) still achieved a score of 0.440. This competitive showing from the open-source models is particularly impressive when you consider the potential “family alignment” advantage in the workflow, where descriptions generated by Gemini might naturally favour Google’s own embedding model. Despite premium models taking the lead, the R² performance gap between the best and worst models was a mere 0.035.
Experiment 2: The SMP challenge image dataset
To validate our initial findings, we applied the exact same framework to the image dataset from the Social Media Prediction (SMP) Challenge, which was instead used to predict popularity of Flickr posts.
The results: Once again, the premium models topped the charts, with both OpenAI and Google producing an identical R² score of 0.234. Just like our Instagram experiment, the weakest free alternative trailed by that same narrow margin, coming in at 0.196.
Experiment 3: Retrieval-augmented generation (RAG) performance
To see how these findings hold up beyond downstream regression tasks, we introduced a third experiment: a traditional Retrieval-Augmented Generation (RAG) evaluation.
At its core, a RAG framework acts as an open-book exam for an LLM. Instead of relying solely on its pre-trained internal knowledge, the system first searches an external database to retrieve the most relevant documents matching a user’s prompt. It then passes these documents alongside the question to the LLM, ensuring the final generated response is accurate, contextually grounded, and factual.
Using the industry-standard BEIR (SciFact) dataset, we indexed 5,180+ scientific documents and evaluated how effectively each embedding model could retrieve the exact context needed to answer 200 distinct queries.
We measured two key retrieval metrics:
- Hit rate: The percentage of queries where the correct document was successfully retrieved in the top 5 results.
- Mean reciprocal rank (MRR): A measure of where the correct document ranked (closer to 1 is better).
The results: This experiment provided an interesting twist. Whilst OpenAI’s premium text-embedding-3-large achieved the highest MRR (0.741), the open-source gte-base model proved remarkably competitive, securing a strong second-place MRR of 0.729 and comfortably beating Google’s premium offering (0.692).
When it came to the overall Hit Rate, the open-source alternative stood completely shoulder-to-shoulder with the premium giants. Instead of a clear winner, we saw a three-way tie at 86% between gte-base, OpenAI, and Google.
Results overview: Scores across multiple experiments
To ensure a robust evaluation across various datasets, we measured performance using the R² score averaged over multiple random training splits, allowing us to establish a reliable variance margin (±) for each model.
In practice, this means instead of training our LightGBM model just once on a single slice of data, we shuffled and split the dataset into different training and testing sets multiple times. Doing this ensures that a model’s high score wasn’t just a fluke resulting from a “lucky” data split. The resulting variance margin (±) acts like an error bar: a tighter margin indicates the model is highly stable, telling us exactly how consistent and reliable its predictions will be when exposed to entirely new data.
| Model | Size Category | Exp 1: Instagram (R²) | Exp 2: SMP Challenge (R²) | Exp 3: RAG (MRR) | Exp 3: RAG (Hit Rate) |
|---|---|---|---|---|---|
text-embedding-3-large (OpenAI) | Premium / Baseline | 0.475 ± 0.022 | 0.234 ± 0.024 | 0.741 | 0.86 |
text-embedding-005 (Google) | Premium / Baseline | 0.470 ± 0.025 | 0.234 ± 0.023 | 0.692 | 0.86 |
thenlper/gte-base | Baseline Match | 0.458 ± 0.018 | 0.212 ± 0.027 | 0.729 | 0.86 |
BAAI/bge-base-en-v1.5 | Baseline Match | 0.457 ± 0.020 | 0.217 ± 0.022 | 0.688 | 0.84 |
all-mpnet-base-v2 | Baseline Match | 0.446 ± 0.021 | 0.216 ± 0.027 | 0.614 | 0.74 |
all-roberta-large-v1 | Larger Alternative | 0.440 ± 0.021 | 0.196 ± 0.020 | 0.594 | 0.71 |
all-miniLM-L6-v2 | Smaller Alternative | 0.439 ± 0.010 | 0.197 ± 0.015 | 0.596 | 0.75 |

The true cost of performance: Latency and pricing
Whilst the minimal gap in predictive accuracy alone makes a compelling case for utilising open-source alternatives, factoring in latency and execution costs makes the decision even clearer.
To illustrate this, we tracked the time and money spent to generate embeddings for each experiment. For reference, the dataset was of size 11,500 rows for Experiment 1, size 10,000 rows for Experiment 2 and 5,180 documents with 200 queries for Experiment 3.
When analysing costs, it is important to separate the costs into two main categories: the cost of calling each embedding model’s API and the cost of the underlying compute resources (such as a virtual machine or a local laptop) required to process the embeddings. For our experiments, we ran everything on a standard Virtual Machine (VM). We have excluded that infrastructure cost from this breakdown, as it fluctuates wildly depending on a developer’s specific deployment preferences and scaling needs.
| Model | Exp 1: Token Fees | Exp 1: Latency (Seconds) | Exp 2: Token Fees | Exp 2: Latency (Seconds) | Exp 3: Token Fees | Exp 3: Latency (Seconds) |
|---|---|---|---|---|---|---|
text-embedding-3-large (OpenAI) | $1.77 | 75 | $0.51 | 41 | $0.22 | 298 |
text-embedding-005 (Google) | $1.29 | 78 | $0.41 | 50 | $0.17 | 241 |
all-roberta-large-v1 | Free | 512 | Free | 438 | Free | 3,777 |
BAAI/bge-base-en-v1.5 | Free | 344 | Free | 247 | Free | 2,229 |
all-mpnet-base-v2 | Free | 282 | Free | 240 | Free | 1,990 |
thenlper/gte-base | Free | 84 | Free | 60 | Free | 11,420 |
all-miniLM-L6-v2 | Free | 43 | Free | 30 | Free | 251 |
The latency vs. cost trade-off: Who wins on speed?
Whilst open-source models eliminate recurring API costs entirely, a common concern is whether self-hosting sacrifices processing speed. To find out, we timed how long each model took to generate the embeddings across our experiments.
To keep the comparison fair, we ran all the open-source models on a standard Google Cloud Platform (GCP) virtual machine: an n1-standard-8 instance (8 vCPUs, 30 GB memory) equipped with a single NVIDIA T4 GPU.
The results revealed a highly competitive landscape with massive implications for production pipelines:
1. In small-to-medium tasks, light open source dominates
For standard datasets like Experiment 1 (Instagram) and Experiment 2 (SMP Challenge), the lightweight open-source models proved that you don’t need to pay for a premium API to get blazing-fast speeds:
- The absolute champion: The tiny
all-miniLM-L6-v2absolutely shredded the competition, completing Experiment 2 in a mere 30 seconds (and Experiment 1 in 43 seconds), beating Google’s premium API by up to 40% whilst costing $0 in token fees. - Premium APIs are fast, but costly: Google’s
text-embedding-005and OpenAI’stext-embedding-3-largeblazed through Experiment 2 in 50 and 41 seconds respectively, but carrying token bills of $0.41 and $0.51. Scaled across millions of rows, those micro-transactions add up quickly.
2. At scale (Experiment 3), the infrastructure tax emerges
When we scaled up to the larger RAG retrieval dataset in Experiment 3, the dynamics shifted heavily, highlighting the core trade-off of hosting your own models:
- Premium APIs pull ahead on massive batches: Google’s model completed the entire RAG dataset in just 241 seconds (costing $0.17), whilst OpenAI finished in 298 seconds ($0.22). Because their infrastructure is globally distributed and massively parallelised, they handle large batches effortlessly.
- The self-hosting bottleneck: Whilst
all-miniLM-L6-v2stayed nimble at 251 seconds, heavier open-source models struggled on our single-GPU setup. For instance,thenlper/gte-base, our RAG accuracy champion, took over 11,400 seconds (more than 3 hours) to complete the run on our T4 GPU.
The Takeaway: When budgeting for your pipeline, remember to separate API token costs from your underlying infrastructure costs (the VM compute time). Model latency serves as an excellent proxy for your infrastructure bill as the longer a self-hosted model runs, the longer your VM has to operate.
If you are running light-to-medium real-time tasks, optimised open-source models like all-miniLM-L6-v2 give you double savings: zero API fees and lower VM uptime. But if you’re processing massive, enterprise-scale batches and don’t want to invest in scaling a heavy local cluster of GPUs, paying a premium API fee might be the more cost-effective route.
Key findings and conclusion
Our experiments across predictive marketing tasks and search retrieval pipelines point to a nuanced conclusion:
- Premium edges out open-source in predictive modelling, but only just: Both OpenAI’s
text-embedding-3-largeand Google’stext-embedding-005consistently delivered the optimal results across our data splits. However, the difference in predictive performance between the premium giants and top open-source models remains minimal, highlighting that free alternatives are highly capable. - Open-source matches premium performance in RAG hit rates: For semantic search and context retrieval, the free open-source
gte-basemodel achieved an identical 86% Hit Rate to both OpenAI and Google, demonstrating that free options can deliver matching retrieval coverage. However, OpenAI’s premium offering did secure the crown for tighter search precision via the highest Mean Reciprocal Rank (MRR) of 0.741, withgte-basefollowing closely behind at 0.729. - The speed vs. cost trade-off has shifted: Highly optimised premium APIs have closed the speed gap, actually outpacing mid-sized open-source models like
gte-base. Whilst open-source still completely dominates on cost (being 100% free), the absolute speed crown belongs specifically to hyper-lightweight open-source models likeall-miniLM-L6-v2, which outran everything at 30 seconds. - Size isn’t everything: Interestingly, large open models like
all-roberta-large-v1underperformed across the board compared to optimised, smaller ones likegte-base.
Our study focused on specific, mainstream tasks, providing strong evidence that free and open-source models can often be enough to get the job done effectively for standard applications.
However, it is important to acknowledge that frontier embedding models, as well as massive open-weight models, still hold distinct advantages. For highly demanding use cases requiring massive context windows, multilingual support, or nuanced reasoning across obscure domains, state-of-the-art models consistently prove their worth. A quick look at popular industry benchmarks, such as Hugging Face’s Massive Text Embedding Benchmark (MTEB) leaderboard, clearly demonstrates this. The frontrunners at the very top of those charts are continually pushing the boundaries of what’s possible across dozens of highly specialised datasets.
Ultimately, depending on whether your priority is:
- Squeezing out the absolute highest KPI value and conquering highly complex edge cases
- Minimising cost and latency whilst maintaining competitive results
You now have a highly capable spectrum of both premium and open-source models to choose from.
Disclaimer: This content was created with AI assistance. All research and conclusions are the work of the WPP Research team.