Sentence Transformers: The Go-To Library for Text Embeddings in 2026

The Sentence Transformers library provides a unified interface for generating text embeddings, enabling semantic search, clustering, and fine-tuning on custom similarity tasks with minimal code.

Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

March 21, 2026
7 min read
Sentence Transformers: The Go-To Library for Text Embeddings in 2026

Why Sentence Transformers

One AI engineering post, weekly

LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.

Raw BERT produces token embeddings, not sentence embeddings. Averaging BERT token outputs gives poor semantic representations - similar sentences get dissimilar vectors. Sentence Transformers (SBERT) fixes this by fine-tuning BERT-style models with siamese networks on natural language inference pairs, producing embeddings where cosine similarity directly correlates with semantic similarity.

The HuggingFace Sentence Transformers collection hosts 200+ pre-trained models covering different size/quality tradeoffs.

Encoding and Cosine Similarity

python
from sentence_transformers import SentenceTransformer, util
import torch

model = SentenceTransformer("all-MiniLM-L6-v2")

sentences = [
    "The quick brown fox jumps over the lazy dog",
    "A fast auburn fox leaps above a sleepy canine",
    "The stock market closed higher today",
]

embeddings = model.encode(sentences, convert_to_tensor=True)

# Pairwise cosine similarity
cos_sim = util.cos_sim(embeddings, embeddings)
print(f"Sentences 0 and 1 similarity: {cos_sim[0][1]:.4f}")  # ~0.72 (semantically similar)
print(f"Sentences 0 and 2 similarity: {cos_sim[0][2]:.4f}")  # ~0.05 (unrelated)

Team workspace

Ship faster with chat, meetings, and projects in one place — Zlyqor.

Start free
python
from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("all-mpnet-base-v2")

corpus = [
    "Python is a high-level programming language",
    "Machine learning requires large datasets",
    "Neural networks are inspired by the human brain",
    "Flask is a lightweight web framework",
]

corpus_embeddings = model.encode(corpus, convert_to_tensor=True)
query = "web development frameworks"
query_embedding = model.encode(query, convert_to_tensor=True)

hits = util.semantic_search(query_embedding, corpus_embeddings, top_k=2)
for hit in hits[0]:
    print(f"Score: {hit['score']:.4f} | {corpus[hit['corpus_id']]}")

Fine-Tuning on Custom Pairs

Use MultipleNegativesRankingLoss when you have (anchor, positive) pairs without explicit negatives - the other items in the batch serve as negatives:

python
from sentence_transformers import SentenceTransformer, InputExample, losses
from torch.utils.data import DataLoader

model = SentenceTransformer("all-MiniLM-L6-v2")

train_examples = [
    InputExample(texts=["What is machine learning?", "ML is a type of AI that learns from data"]),
    InputExample(texts=["How do I fix a bug?", "Debugging requires isolating the failing component"]),
]

train_dataloader = DataLoader(train_examples, shuffle=True, batch_size=16)
train_loss = losses.MultipleNegativesRankingLoss(model)

model.fit(
    train_objectives=[(train_dataloader, train_loss)],
    epochs=3,
    warmup_steps=100,
)
model.save("my-finetuned-model")

Best Models Comparison

ModelDimensionsSpeedQuality
all-MiniLM-L6-v2384Very fastGood
all-mpnet-base-v2768FastBetter
multi-qa-mpnet-base-dot-v1768FastBest for QA
BGE-M31024ModerateBest overall

The GitHub repository includes pretrained model benchmarks on STS, QA, and retrieval tasks. For most production RAG use cases, all-mpnet-base-v2 is the baseline to beat before reaching for larger models.

#sentence-transformers#embeddings#semantic-search#sbert#fine-tuning

// discussion

Comments

0/4000
Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.

PristrenZlyqor