Skip to content
AI Primer
release

Google releases 740M-parameter EmbeddingGemma 2 for on-device media search

Google's open 740M-parameter multimodal embedding model runs on-device for image, video, and voice search. A local demo indexed 300 photos and reached 82.5% precision at 10 on held-out animal queries.

3 min read
Google releases 740M-parameter EmbeddingGemma 2 for on-device media search
Google releases 740M-parameter EmbeddingGemma 2 for on-device media search

TL;DR

  • Google has shipped EmbeddingGemma 2, a 740M-parameter model that maps text, code, images, audio, and video into one 768-dimensional embedding space, according to GoogleDeepMind's launch post.
  • The model is built for local inference, with roughly 191 MB of active RAM for text-only use and 567 MB for the full multimodal model on a Pixel 11 Pro, according to Google's Edge deployment guide and GoogleDeepMind's launch post.
  • A local demo indexed 300 photos and reached 82.5% precision at 10 on 200 held-out retrieval results, stevibe's local demo reports.
  • Voice search can retrieve a timestamped video moment from a spoken query, using raw audio and video frames without transcripts, minchoi's reply says. Google released the weights under Apache 2.0, with links to Hugging Face and Kaggle in GoogleDeepMind's launch post.

A spoken request such as “someone playing a guitar” can pull up the matching video segment, as minchoi's voice-search demo shows. Google's developer post describes the local workflow as a way to avoid chaining separate captioning, speech-to-text, and text-embedding models.

One shared embedding space

EmbeddingGemma 2 is built on the Gemma 4 architecture and supports cross-modal retrieval through a single representation rather than separate indexes for each media type, according to Google's announcement.

The model maps these inputs into a shared 768-dimensional vector space:

  • Text, including source code.
  • Still images and video frames.
  • Audio.
  • Combinations such as a voice query against visual content.

Google also positions it as a component for private, on-device RAG when paired with Gemma 4, GoogleDeepMind's launch post says.

Modular on-device footprint

The 740M parameters are split across a text model and optional modality encoders. The Hugging Face model card describes the layout as:

  • 270M parameters for the text model.
  • 170M for vision.
  • 300M for audio.

The model card says those encoders can be loaded selectively, while Google's Edge documentation reports about 191 MB of active RAM for text-only weights and 567 MB for the full multimodal configuration on a Pixel 11 Pro. It supports an 8K context window and more than 100 languages, according to the same model card.

stevibe's local test indexed 300 photos across 30 animal categories, then used 20 separate photos as queries, stevibe's local demo reports.

The evaluation used a straightforward retrieval loop:

  • Each query became a 768-dimensional vector.
  • Cosine similarity selected the 10 nearest indexed images.
  • Animal labels were used only to check the results.
  • 165 of 200 retrieved images matched the query animal, for 82.5% precision at 10.

Zebra, tiger, giant panda, and polar bear queries scored 10 out of 10. Horse and leopard were harder, scoring 4 out of 10 and 5 out of 10 respectively, stevibe's local demo found.

Voice-to-video retrieval

The demo turns a spoken search for “someone playing a guitar” into a timestamped video result. minchoi clarified that the model listens to raw audio and examines video frames in the same search space, with no transcript step, minchoi's reply says.

768 dimensions, or fewer

The default output is 768 dimensions, but EmbeddingGemma 2 supports Matryoshka truncation to 512, 256, or 128 dimensions, according to the model card.

Google says truncation can reduce vector storage by up to six times. Its guidance positions 256 dimensions as the balanced option and 128 dimensions for text-only workloads, while retaining the same model for indexing and retrieval, according to the model card.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 4 threads
TL;DR2 posts
One shared embedding space1 post
Voice-to-video retrieval1 post
Share on X