Skip to content
AI Primer
release

Google releases EmbeddingGemma 2, a 740M-parameter local retrieval model

Google released EmbeddingGemma 2 under Apache 2.0 to map text, code, images, video, and audio into one embedding space. The 740M-parameter version supports 8K context and on-device retrieval.

6 min read
Google releases EmbeddingGemma 2, a 740M-parameter local retrieval model
Google releases EmbeddingGemma 2, a 740M-parameter local retrieval model

TL;DR

  • Google released EmbeddingGemma 2 as an Apache 2.0 model that maps text, code, images, video, and audio into one shared embedding space, according to Google's announcement.
  • The model scales from 270M parameters for text and code to 740M for all modalities, with independently loadable vision and audio encoders, as tomaarsen's modular breakdown details.
  • Its 8,192-token window is shared across modalities, while Google's reported active-RAM figures range from about 191MB for text-only to 567MB for the full quantized model, Google's model-size note says.
  • The biggest text benchmark gain is code, which rose from 68.76 to 78.68 on MTEB Code, while multilingual text barely moved, according to tomaarsen's benchmark note.
  • The model had day-one support in llama.cpp and vLLM, with Ollama, browser, and Sentence Transformers paths also available through llama.cpp support and vLLM's announcement.

A voice memo finding a video clip is the launch demo in Google's post. The model card hides two details engineers will care about: all modalities consume one context budget, and FP16 can produce NaNs or silently degraded vectors. The developer guide also shows a single product listing, with text, photos, and video, becoming one searchable vector.

One vector space

EmbeddingGemma 2 is Google's open embedding model built on the Gemma 4 architecture. It produces 768-dimensional vectors for text, code, images, video, and audio, letting a text query rank media directly instead of requiring separate captioning, speech-to-text, and text-embedding stages, according to Google's AI Edge post.

The parameter budget is modular:

  • 270M: text and code backbone.
  • 440M: text, images, and video, adding a 170M vision encoder.
  • 570M: text and audio, adding a 300M audio encoder.
  • 740M: the full multimodal configuration.

All four configurations project into the same vector space. The developer guide says a text-only query can be compared with vectors produced by the full model, and adding an encoder later does not require recomputing existing embeddings.

Memory and context

The 8,192-token window is shared by every input type. The model card lists these default costs:

  • Text: one token per subword.
  • Images: 280 tokens each, or about 29 images in an otherwise empty context.
  • Video: 140 tokens per frame, sampled at one frame per second by default, or about 58 frames.
  • Audio: 25 tokens per second of mono 16 kHz audio, or about 327 seconds without accompanying text.

Mixed inputs draw from the same budget, so a product description, images, and video reduce one another's available context. Vision detail is configurable from 70 to 1,120 soft tokens per image or frame, trading latency and context capacity for visual detail, as tomaarsen's input-budget note reports.

Google measured about 191MB of active RAM for quantized text-only weights and 567MB for the full multimodal model on a Pixel 11 Pro, Google's model-size note says. The model card adds a deployment hazard: use bfloat16 or float32, because float16 can return NaNs or silently degraded embeddings, a warning repeated by tomaarsen's precision warning.

Vector storage

Matryoshka Representation Learning lets the same model emit 768-, 512-, 256-, or 128-dimensional vectors. The model card reports the following tradeoffs:

  • 512d: 1.5x storage reduction; MMEB overall falls from 59.01 to 58.38.
  • 256d: 3x reduction; multilingual MTEB moves from 61.36 to 60.41, while MMEB overall moves to 56.24.
  • 128d: 6x reduction; MMEB overall falls to 45.65 and is positioned for text-only workloads.

Vectors shortened below 768 dimensions must be re-normalized before cosine search, and queries and documents must use the same dimension. tomaarsen's truncation test measured the sixfold storage reduction alongside the sharper multimodal quality loss at 128d.

Queries, documents, and media

Text retrieval uses task-specific prefixes rather than one generic embedding instruction. The model card defines separate query and document formats for search, question answering, fact checking, and code retrieval, plus symmetric prompts for classification, clustering, and sentence similarity.

The important split is:

  • A search query uses SearchQuery, while corpus text uses Document.
  • Code search uses CodeRetrieval, with a title or filename attached to the code.
  • Media receives no text prefix, according to tomaarsen's task-prefix note.
  • Interleaved inputs place <|image|>, <|video|>, or <|audio|> markers in the text and return one vector for the combined object, as tomaarsen's interleaving example shows.

That last path supports a listing with a description, two photos, and a demo video, then compares it with a text-only query. A separate text query can also retrieve photos, recordings, and clips through the same API, tomaarsen's API note says.

Benchmarks and caveats

The model card reports these full-precision results at 768 dimensions:

  • MTEB Code: 68.76 for EmbeddingGemma 1 to 78.68 for EmbeddingGemma 2, a +9.92-point change.
  • MTEB multilingual: 61.15 to 61.36, a +0.21-point change.
  • MTEB English: 69.67 to 68.46, a -1.21-point change.
  • MMEB: 57.28 on image retrieval, 67.84 on visual documents, and 50.67 on video.
  • Audio: 69.54 on MSEB retrieval and 49.39 on MAEB.

The code result is the clear text-side improvement. The multilingual result is nearly flat, and the English score is lower than the first generation. An independent launch-day analysis by S5 Labs reported that the published numbers were Google self-reported, with no independent quality benchmark available at launch and the MTEB submission still unaccepted.

Local runtimes

Weights are available through Hugging Face and Kaggle under Apache 2.0, while Google's launch post lists Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, LM Studio, LiteRT, and browser deployment through transformers.js or WebGPU.

The ecosystem moved quickly:

  • llama.cpp shipped day-zero support, the llama.cpp support post said.
  • vLLM announced day-zero support for its nightly build, the vLLM announcement said.
  • Ollama exposed the model as ollama pull embeddinggemma-2, according to Ollama's release note.
  • Sentence Transformers 6.1.0 or later provides the Python path, while Google AI Edge's MediaPipe and LiteRT stack targets mobile, desktop, and web deployments.
  • Google's Android path through ML Kit is planned for the coming weeks, with NPU acceleration when available, according to the AI Edge deployment post.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR2 posts
Memory and context1 post
Queries, documents, and media1 post
Benchmarks and caveats2 posts
Local runtimes2 posts
Share on X