AI NewsModels & agentsAnnouncement

Google released EmbeddingGemma 2, a 740M open-weights multimodal embedding model

Google released EmbeddingGemma 2 as open weights under Apache 2.0. The full model is 740M parameters and encodes text, code, images, video and audio into a shared 768-dimensional space. A 270M text and code variant ships alongside, and the embeddings can be truncated from 768 to 128 dimensions through Matryoshka Representation Learning.

AI News

Editorial2 min read

LinkedInX
Banner image for EmbeddingGemma 2 from the Google for Developers blog

Image: Google

Why it mattersA team that wants cross-modal search without a sequence of model calls now has a single model it can run on a laptop or a phone. Google says the 256-dimensional embeddings keep about 95 percent of retrieval quality, which cuts a one-million-vector index from 1.5 GB to 250 MB.

A team building retrieval over mixed content has been running one encoder for text, another for images, and often a third to join the two together. Google just released a single open-weights model that handles all of them, and it is small enough to run on a phone.

Maarten Grootendorst and Ian Ballantyne at Google announced EmbeddingGemma 2 on the Google for Developers blog. The model encodes text, code, images, video and audio into a shared 768-dimensional vector space, so a search across any of those modalities is a single model call followed by a vector lookup.

Model shape and licence

The release ships two sizes. The text-and-code variant is 270M parameters. The full multimodal model is 740M parameters, and the vision and audio towers can be disabled at load time to drop back to the smaller footprint when a workload does not need them. The context window is 8,192 tokens. Both are released under Apache 2.0.

Google says it used Matryoshka Representation Learning, which trains the model so that the first 128, 256 or 512 dimensions of each 768-dimensional vector are usable on their own. In the company's own evaluation, 256 dimensions keep about 95 percent of retrieval quality on image, video and speech, and the same technique drops a million-vector index from 1.5 GB to 250 MB.

What the numbers show, in Google's words

Every performance figure in the announcement is Google's own. Compared with EmbeddingGemma 1, the company reports a 14 percent gain on code-focused benchmarks. Google frames the 256-dimensional truncation as a quality-for-storage trade-off and gives the 95 percent figure for how much retrieval quality survives it. The announcement does not include a direct comparison to any third-party multimodal embedding model, and there is no independent benchmark to cite yet.

Google also published a companion post on edge deployment, and Hugging Face shipped support in the Transformers 5.19.0 release on the same day, so the model is available through the standard AutoModel API without extra setup.

What this replaces

A retrieval system over mixed assets today is usually a sequence: one model for text, one for images, one for audio, and a hand-written step that reconciles their outputs. Each adds a round trip, a cache, a slot in the latency budget, and a slot in the API bill. One multimodal model shared across every modality removes the sequence. For teams running retrieval on a device without a reliable network, the 270M text-only variant is small enough to ship inside a mobile app, and the full 740M model runs on an ordinary laptop with local inference.

Open weights under Apache 2.0 also matter for anyone who has been blocked from embedding user content with a hosted API, because the data stays on the machine the model runs on. The limits are the ones that come with every multimodal release: Google's quality numbers are what Google measured, every benchmark figure is at the publication date, and specialised domains will still want a fine-tune.

Source

Google for Developers: EmbeddingGemma 2: The Developer Guide by Maarten Grootendorst and Ian Ballantyne. Model card and weights on Hugging Face.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX