Google releases EmbeddingGemma 2, an open multimodal AI model for on-device use
Google DeepMind released EmbeddingGemma 2, an open-source multimodal embedding model under the Apache 2.0 license. Built on the Gemma 4 architecture, it maps text, code, images, video, and audio into a unified 768-dimensional vector space. The model has up to 740 million parameters with modular encoders, an 8K-token context window, and is designed for on-device efficiency, enabling privacy-first retrieval-augmented generation without sending data to a server.
IllustrationEditorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
Google releases EmbeddingGemma 2 for on-device multimodal AI under Apache 2.0 license
Google has released EmbeddingGemma 2, a multimodal AI model for on-device use, under an Apache 2.0 license. The model integrates text, code, images, audio, and video into a single searchable space on phones, enabling users to search their own data without sending it to a server. It has 740 million parameters and uses the Gemma 4 architecture. Its modular design allows a text-only app to use just 270 million parameters, while a 170 million vision encoder and a 300 million audio encoder load only when needed. On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB. The context window has grown 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images, or 58 video frames in one input. Code search performance improved significantly, with the MTEB Code score rising from 68.76 to 78.68, while multilingual text scores held level. Google also claims top sub-1B results on audio and vision benchmarks, outperforming some specialist models twice its size.
Read sourceGoogle DeepMind Releases Open-Source Lightweight Multimodal Embedding Model EmbeddingGemma 2
Google DeepMind has announced the release of EmbeddingGemma 2, a lightweight multimodal embedding model based on the Gemma 4 architecture. The model is open-sourced under the Apache 2.0 license and is designed to map text, code, images, video, and audio into a unified embedding space. The official release includes specific details on parameters, memory usage, dimension pruning, and benchmark scores, allowing developers to evaluate the feasibility of on-device multimodal embedding solutions. This release aims to advance accessible AI by providing a versatile, efficient model for various multimodal tasks.
Read sourceGoogle releases EmbeddingGemma 2 open-weight multimodal model under Apache 2 license
Google has released EmbeddingGemma 2, an open-weight multimodal embedding model, under the Apache 2 license. The model is described as lightweight and capable of mapping text, code, images, video, and audio into a single unified embedding space. It features a 740 million parameter form factor with modular encoders and an 8,000-token context window. The announcement was made via a post on X, which included a link for embedded testing. This release represents a significant step in making versatile, multimodal AI embedding technology openly available to developers and researchers.
Read sourceShow 3 older updatesHide older updates
GoogleDeepMind unveils EmbeddingGemma 2, its first natively multimodal open model
GoogleDeepMind announced EmbeddingGemma 2, its first natively multimodal open model designed for on-device embeddings. The model expands beyond text to unify code, images, audio, and video in a shared embedding space. This release represents a significant step in making multimodal AI capabilities accessible for on-device applications, allowing developers to process and compare different types of data within a single framework. The model is open, suggesting it is available for public use and modification, which could accelerate research and development in areas such as cross-modal search, content understanding, and edge AI. The announcement was made via a social media thread, highlighting the model's ability to handle diverse data types natively, without requiring separate processing pipelines for each modality.
Read sourceGoogle introduces EmbeddingGemma 2, an open multimodal model for on-device efficiency
Sundar Pichai announced the release of EmbeddingGemma 2, Google's first open, natively multimodal embedding model. The model is designed for on-device efficiency, handling text, code, image, video, and audio tasks within a lightweight, modular 740 million parameter form factor. It is described as ideal for offline, privacy-first retrieval-augmented generation (RAG) when paired with Gemma 4. The announcement claims it outperforms some specialist models more than twice its size. Model weights are available immediately on Hugging Face.
Read sourceGoogle Releases EmbeddingGemma 2, Open-Source Multimodal Embedding Model Based on Gemma 4
Google has announced the release of EmbeddingGemma 2, an open-source embedding model based on the Gemma 4 architecture. The model, distributed under the Apache 2.0 license, is designed to map multiple data modalities—including text, code, images, video, and audio—into a unified 768-dimensional vector space. The model's parameter count ranges from 270 million for text and code-only configurations to 740 million for the full multimodal version, allowing developers to load only the necessary components for their specific use case. The Google Developers Blog post provides detailed configuration options, dimension compression storage figures, and selection recommendations to help developers plan local multimodal retrieval solutions. This release aims to advance accessible multimodal AI capabilities for the developer community.