AI NewsModels & agentsAnnouncement
Cohere released Embed 5 with a Pro and a Fast model that share one vector space, so teams can index with Pro and query with Fast on the same corpus
Cohere released Embed 5 on Wednesday. The Pro and Fast models share one vector space, so a team can index a corpus with Pro and serve queries with Fast against the same index. Cohere says a Fast query against a Pro index scores 98.4 on its 40-dataset suite, against a Pro-to-Pro baseline of 100.

Image: Cohere
Why it mattersA RAG pipeline that reads its index far more often than it writes to it can now spend the expensive model on the write path and the cheap one on the read path, without maintaining two copies of the vectors or re-embedding when a decision changes.
Cohere released Embed 5 on Wednesday, and the detail worth reading is the one under the benchmarks. Embed 5 Pro and Embed 5 Fast share the same vector space at the same dimensions, so a team can index a corpus with Pro and serve queries with Fast against the same index. No second copy of the vectors, no re-embed when a model changes on either side.
Cohere puts a number on the trade. Across 40 datasets covering text, images, fused documents and parsed documents, a Fast query against a Pro index scored 98.4 against a Pro-to-Pro baseline of 100. Fast on both sides dropped that to 96.6. Cohere says no individual dataset in the suite showed a large drop when Pro and Fast were mixed.
What the two models cost
Cohere lists Pro at $0.12 per million tokens and Fast at $0.08, with Fast running an average of 2.4 times the document throughput in its own tests. For a system that ingests documents far less often than it searches them, Pro can handle each document as it enters the index and Fast can handle the much heavier query traffic. Both models expose six vector dimensions from 256 to 2,048, with float32, int8 and binary formats, and both support Matryoshka truncation and int8 quantization, so a corpus stored in one representation can be searched at another without a re-embed.
Cohere sets out the storage difference at scale in the announcement. A 2,048-dimensional float32 vector is 8 KB, which comes to 819 GB for 100 million chunks. A 1,024-dimensional int8 vector cuts that to 102 GB, and a 256-dimensional binary vector brings the same corpus down to 3.2 GB. Cohere recommends 1,024-dimensional int8 for most deployments, and treats binary as a first-stage filter before a higher-precision rerank.
Where Cohere claims the leadership
The launch names three evaluations where Pro comes out ahead of the rivals Cohere tested. On its five-dataset fused text-image benchmark, Pro averaged 82.3 against 81.2 for Fast and 61.3 for Google's Gemini Embedding 2. On its parsed-PDF evaluation, which covers service documentation, corporate reports, SEC filings, product manuals and privacy policies, Pro averaged 84.8 against 83.6 for Voyage 4 Large, 83.4 for Fast, 80.8 for Gemini Embedding 2 and 78.6 for Cohere's own Embed 4. On ViDoRe V3, evaluated using parsed text rather than page images, Pro averaged 85.8 and Fast 84.5.
Multilingual results are less one-sided. Pro leads Cohere's five-language European average at 77 against 76 for Voyage 4 Large and 73 for Gemini Embedding 2, and Cohere reports Gemini Embedding 2 ahead on nine of its multilingual tests.
Read the fine print
Cohere flags one caveat in the announcement itself. Embed 5 is the first Cohere family evaluated with RCP-nDCG@10, which uses query-specific relevance criteria and measures reranking over a fixed candidate set rather than first-stage retrieval from the full corpus. First-stage retrieval is evaluated separately using standard nDCG and Recall, and the fused text-image, page-image and cross-model tests use standard nDCG@10, so the reported scores are not directly comparable across evaluations. Cohere also says production teams will still need to benchmark Pro-to-Fast against Pro-to-Pro on their own corpus and query distribution, because retrieval errors carry through multi-step agent workflows.
Embed 5 Pro and Embed 5 Fast are available through Cohere's API and Model Vault, Microsoft Foundry and Amazon SageMaker, with private VPC and on-premises deployment through vLLM.
Source
- Primary: Introducing Embed 5, A New Family of Frontier Embedding Models, Cohere, 2026-09-30.
- Reporting: Cohere's faster query model barely dents retrieval quality in its tests, The New Stack, by Amanda Caswell, 2026-09-30.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


