Google's EmbeddingGemma 2 Maps Text, Photos, Video, and Audio on Device
Google DeepMind launched EmbeddingGemma 2, a 740-million-parameter open-weight model that maps text, images, video frames, and audio into a single 768-dimensional vector space. It runs on phones with about 191MB of RAM for text and 567MB fully multimodal, embeds vision in 37 milliseconds on a laptop GPU, and slices to 128 dimensions via Matryoshka training for smaller indexes.
On this page
One vector space for four modalities
Google DeepMind launched EmbeddingGemma 2 on 6 October, a 740-million-parameter open-weight embedding model that natively maps text, images, video frames, and audio into a single unified vector space, producing 768-dimensional vectors that developers can slice down to 128 or 512 dimensions through Matryoshka Representation Learning, which Google says cuts local storage and index footprints by up to eight times.
The architectural idea is modular per-modality encoders feeding one space, so a query in any modality lands near its matches in every other. That is what makes the launch a search-infrastructure story rather than a model story: the same embedding answers a typed phrase about a photo, finds the moment in a video where something was said, and routes an assistant's intent, without per-task fine-tuning. Google's demonstration includes zero-shot intent routing that matches inputs to labels with no training data at all.
Sized for the edge
The memory budget is the specification that matters. Text-only embedding needs about 191MB of active RAM; the full multimodal configuration runs in about 567MB, which Google measured on a Pixel 11 Pro. Vision embeddings take as little as 37.3 milliseconds per image, about 27 images per second, on a MacBook M5 Pro GPU, and approximate-nearest-neighbor retrieval returns ranked matches in single-digit milliseconds.
Deployment paths are unusually complete for launch day: LiteRT bundles, MediaPipe Tasks including a UniversalEmbedder and a Decision Task that scores 500 options in under 100 milliseconds, ML Kit integration on Android with NPU acceleration arriving in weeks, and a WASM path for the web. Weights are on Hugging Face's LiteRT community, with quantization-aware training to INT4 and INT8 for further compression.
What it unlocks for local search
The timing is not accidental. Open tools that index personal media entirely on-device shipped this same week, and the embedding layer they depend on is exactly what EmbeddingGemma 2 now provides as a maintained, multilingual, four-modality model under an open-weight release. A local photo library, a folder of lecture recordings, and a phone's camera roll all become semantically searchable without any content leaving the device, with vector sizes small enough that the index itself fits in memory.
The honest caveats are that Google quotes edge latency on its chosen hardware, that the license is described only as open-weight on the launch page, and that multimodal embedding quality across all four modalities will need third-party benchmarks to confirm. But the direction is unambiguous: the embedding models that power local semantic search are getting smaller, broader, and maintained by the same company that ships the runtimes they run on. The gap between what a phone can search and what a cloud service can search narrowed again this week.