OpenAI text-embedding-3: The New Embedding Models and When to Use Each

OpenAI's text-embedding-3-small and text-embedding-3-large introduce Matryoshka representation learning - you can truncate dimensions without retraining, cutting storage costs while keeping most retrieval quality.

Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

March 16, 2026
8 min read
OpenAI text-embedding-3: The New Embedding Models and When to Use Each

The Two New Models

One AI engineering post, weekly

LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.

OpenAI replaced text-embedding-ada-002 with two new models in January 2024:

  • text-embedding-3-small - 1536 dimensions, $0.02/1M tokens
  • text-embedding-3-large - 3072 dimensions, $0.13/1M tokens

Ada-002 charged $0.10/1M tokens for a model that MTEB scores showed was falling behind newer alternatives. The new models are both cheaper and more capable.

MTEB Leaderboard Performance

The Massive Text Embedding Benchmark (MTEB) covers 56 tasks across retrieval, classification, clustering, and semantic similarity.

ModelMTEB AverageDimensionsCost/1M tokens
text-embedding-3-large64.63072$0.13
text-embedding-3-small62.31536$0.02
text-embedding-ada-00261.01536$0.10
Cohere embed-v364.51024$0.10
Voyage-367.11024$0.06

text-embedding-3-small beats ada-002 at one-fifth the price - for most RAG use cases, it is the obvious default.

Team workspace

Ship faster with chat, meetings, and projects in one place — Zlyqor.

Start free

Matryoshka Representation Learning

The headline technical feature is Matryoshka embeddings: the model is trained so that the first N dimensions of a 3072-dimension vector are nearly as useful as the full vector. This means you can truncate dimensions at query time without retraining.

python
from openai import OpenAI
import numpy as np

client = OpenAI()

def get_embedding(text: str, dimensions: int = 1536) -> list[float]:
    response = client.embeddings.create(
        model="text-embedding-3-small",
        input=text,
        dimensions=dimensions,  # truncate here, not post-hoc
    )
    return response.data[0].embedding

def cosine_similarity(a: list[float], b: list[float]) -> float:
    a_arr = np.array(a)
    b_arr = np.array(b)
    return float(np.dot(a_arr, b_arr) / (np.linalg.norm(a_arr) * np.linalg.norm(b_arr)))

query_emb = get_embedding("How do transformers handle long sequences?", dimensions=256)
doc_emb = get_embedding("Attention mechanisms scale quadratically with sequence length.", dimensions=256)

print(f"Similarity: {cosine_similarity(query_emb, doc_emb):.4f}")

Using 256 dimensions instead of 1536 reduces vector storage by 6x while retaining roughly 92% of retrieval quality on most benchmarks.

Migration from Ada-002

The embeddings are not backward compatible - ada-002 vectors and text-embedding-3 vectors live in different spaces and cannot be compared. If you are migrating a production vector database:

  1. Keep ada-002 running for existing queries
  2. Re-embed your entire corpus with text-embedding-3-small
  3. Update your vector store index
  4. Cut over traffic and deprecate ada-002

For Pinecone, create a new index with the new dimension count (1536 for small, 3072 for large). For pgvector, alter the column or create a new one.

When Voyage or Cohere Beat OpenAI

  • Voyage-3 consistently leads MTEB for English retrieval tasks - if maximum retrieval accuracy is the priority and you can afford slightly more complex integration, Voyage is worth testing.
  • Cohere embed-multilingual-v3 dominates when you need 100+ languages - OpenAI's multilingual performance is good but not best-in-class.
  • OpenAI wins on simplicity (one SDK, one billing account) and latency (well-optimized inference infrastructure).
#openai-embeddings#text-embedding-3#rag#mteb#vector-search

// discussion

Comments

0/4000
Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.

PristrenZlyqor