Google Launches EmbeddingGemma 2 for On-Device Multimodal AI
Google’s EmbeddingGemma 2 brings multimodal embeddings to phones and laptops, enabling developers to process text, code, images, video and audio locally while reducing cloud dependency and improving...
Google has introduced EmbeddingGemma 2, an open multimodal embedding model designed to bring advanced AI retrieval capabilities directly to local devices.
With 740 million parameters, the model can represent text, code, images, video and audio in a shared embedding space. This allows developers to build applications that understand relationships across different types of content rather than treating each modality separately.
A key advantage is its relatively small footprint. EmbeddingGemma 2 can operate with around 191MB for text-only workloads and approximately 567MB when fully loaded, making it suitable for phones, laptops and other resource-constrained devices. It also supports an 8K context window for handling longer inputs.
The model can be useful for applications such as semantic search, recommendation systems and retrieval-augmented generation (RAG). Running inference locally can also help reduce latency, cloud costs and the need to send sensitive information to external servers.
Google has made the model weights available through Hugging Face, giving developers an easier path to experiment and integrate the technology.
For AI developers, EmbeddingGemma 2 signals a continued shift toward smaller, capable multimodal AI models that can run directly on devices.


