Cohere Released Embed 5 Multimodal Model Family

The new embedding models allow developers to decouple high-precision indexing from high-speed retrieval.

Updated on Sept. 30, 2026 in Artificial Intelligence

Bold flat-color editorial illustration showing three sculptural spheres suspended within a geometric metal cube, representing a technical vector space.
Cohere has launched its Embed 5 model family, an architecture supporting text and image inputs designed to balance high-precision indexing with low-latency retrieval for RAG systems. AI Illustration. Upload story photo >

Live Poll

Would you prioritize cost savings over performance in your own AI system deployments?

Cohere has released the Embed 5 model family, an architecture designed to support text, image, and fused inputs across 100 languages. These models are now available for deployment via API, Model Vault, Microsoft Foundry, and Amazon SageMaker.

Why it matters

By providing two distinct models that share an embedding space, this release enables developers to optimize Retrieval-Augmented Generation (RAG) and agent systems for both indexing accuracy and inference speed. This approach prevents the need for costly re-indexing when scaling query throughput.

The Fast model costs $0.08 per million tokens compared to $0.12 for the Pro model, while both support a 128K-token context window. The system allows for vector dimensions ranging from 256 to 2,048.

The players

Cohere

An enterprise AI platform provider specializing in large language models and embedding services for RAG applications.

The details

Embed 5 utilizes a shared embedding space, a mathematical coordinate system that maps different data types into vectors, enabling the Pro and Fast models to operate interchangeably on the same dataset. Developers index data using the Pro model for higher precision and query that same index with the Fast model for lower latency. The platform supports fused text-image inputs, allowing for complex multimodal retrieval tasks across the specified 128K-token context window.

Timeline

  1. September 30, 2026: Cohere released the Embed 5 model family.

The Tech Race

This release follows the trend of decoupling indexing and retrieval to lower the cost of large-scale search operations. It positions Cohere against other proprietary embedding providers by emphasizing interoperability between different performance-tier models within the same stack.

Developers and companies can immediately access Embed 5 via API or cloud partners like Microsoft Foundry and Amazon SageMaker. This update enables immediate cost reductions for applications needing high query speeds, as users can switch to the $0.08 per million token Fast model.

The takeaway

The Embed 5 family signals an industry move toward specialized model pairings that prioritize architectural efficiency over monolithic scaling. Watch for future benchmarks regarding image-fused retrieval quality as the system gains wider adoption.

Further reading

For broader trends in enterprise search and vector database infrastructure, visit Artificial Intelligence.

Live Poll

Would you prioritize cost savings over performance in your own AI system deployments?