Cohere Embed v3: Multilingual Embeddings Built for Enterprise RAG

Cohere's Embed v3 introduces a critical input_type parameter that tells the model whether it's encoding a query or a document - a distinction that meaningfully improves retrieval precision in production RAG pipelines.

Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

March 20, 2026
7 min read
Cohere Embed v3: Multilingual Embeddings Built for Enterprise RAG

What Makes Embed v3 Different

One AI engineering post, weekly

LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.

Most embedding models treat all text the same way. Cohere Embed v3 does not. It takes an explicit input_type parameter that changes how the model encodes text based on its intended use. This is not a minor API detail - it is the primary reason Embed v3 outperforms competitors on asymmetric retrieval tasks (where queries are short and documents are long).

The input_type Parameter

python
import cohere

co = cohere.Client("YOUR_COHERE_KEY")

# Encoding a user search query
query_emb = co.embed(
    texts=["What is the capital of France?"],
    model="embed-english-v3.0",
    input_type="search_query",
    embedding_types=["float"],
).embeddings.float[0]

# Encoding documents to store in a vector database
doc_embs = co.embed(
    texts=[
        "Paris is the capital and largest city of France.",
        "France is a country in Western Europe.",
    ],
    model="embed-english-v3.0",
    input_type="search_document",
    embedding_types=["float"],
).embeddings.float

The four valid values are:

  • search_query - for user queries at retrieval time
  • search_document - for documents being indexed
  • classification - for text classification tasks
  • clustering - for topic clustering and deduplication

Always use the matching type at index time and query time, otherwise you are leaving retrieval quality on the table.

Team workspace

Ship faster with chat, meetings, and projects in one place — Zlyqor.

Start free

Model Variants

  • embed-english-v3.0 - 1024 dimensions, English only, highest English performance
  • embed-multilingual-v3.0 - 1024 dimensions, 108 languages, within 2-3 points of English-only on most tasks

Both output 1024-dimensional float32 vectors. Pricing is $0.10/1M tokens for both.

Compressed Embedding Formats

Cohere Embed v3 is one of the first production embedding APIs to natively support compressed output formats:

python
# int8 compression  -  4x storage reduction, ~1% quality loss
doc_embs_int8 = co.embed(
    texts=["Your document text here"],
    model="embed-english-v3.0",
    input_type="search_document",
    embedding_types=["int8"],
).embeddings.int8

# binary compression  -  32x storage reduction, ~3-5% quality loss
doc_embs_binary = co.embed(
    texts=["Your document text here"],
    model="embed-english-v3.0",
    input_type="search_document",
    embedding_types=["ubinary"],
).embeddings.ubinary

For a corpus of 10 million documents at 1024 dimensions:

  • float32: ~40GB
  • int8: ~10GB
  • binary: ~1.25GB

The binary format fits a 10M-document index in the RAM of a standard cloud VM, enabling in-memory vector search without specialized hardware.

MTEB Comparison

ModelMTEB English RetrievalMultilingual
Cohere embed-english-v3.055.0English only
Cohere embed-multilingual-v3.054.1108 languages
OpenAI text-embedding-3-large55.4Good
Voyage-358.1Limited

OpenAI edges ahead on raw MTEB retrieval, but Cohere's multilingual coverage (108 languages) and compressed format support make it the stronger choice for international enterprise deployments.

Weaviate Integration

python
import weaviate
from weaviate.classes.init import Auth

client = weaviate.connect_to_weaviate_cloud(
    cluster_url="YOUR_WEAVIATE_URL",
    auth_credentials=Auth.api_key("YOUR_WEAVIATE_KEY"),
    headers={"X-Cohere-Api-Key": "YOUR_COHERE_KEY"},
)

# Weaviate handles Cohere embedding automatically with text2vec-cohere
collection = client.collections.get("Document")
results = collection.query.near_text(
    query="machine learning optimization techniques",
    limit=5,
)
#cohere-embed#multilingual#embeddings#rag#enterprise

// discussion

Comments

0/4000
Mahmudul Haque Qudrati — CEO & ML Engineer at Pristren

Mahmudul Haque Qudrati

CEO & ML Engineer

Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.

PristrenZlyqor