
EmbeddingGemma 2: Build Text and Image Search in Python
Summary
Run Google's new open multimodal embedder locally and search text, images and video with one index.
On October 6, Google DeepMind released EmbeddingGemma 2, an Apache 2.0 embedding model that puts text, code, images, video and audio into the same 768-dimensional vector space. It is built on a Gemma 4 text encoder and ships in sizes from 270M to 740M parameters, which means a laptop CPU can run it. The Hugging Face card and Google's developer guide both went up the same day, and the model has been climbing the trending lists since.
Why it matters: most RAG stacks today use one embedding model for text and a second, completely separate pipeline for images (usually OCR plus captioning). EmbeddingGemma 2 lets you skip that. A text query like 'how do I get my money back' can land directly on a screenshot of a refund policy slide, with no OCR step in between. And because the encoders are modular, you can load only the parts you need.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment