EmbeddingGemma 2 — A lightweight model for on-device multimodal embeddings.
EmbeddingGemma 2 unifies text, images, audio, and video into a single embedding space, optimized for on-device performance. This model is designed for developers looking to enhance applications with robust multimodal capabilities while ensuring data privacy.
- Supports a wide range of media types, enabling tasks like searching video clips from audio queries or finding images based on text descriptions (E009).
- Built on the Gemma 4 architecture, it features 740 million parameters and is optimized for efficient on-device inference, requiring minimal RAM (E008).
- Offers an extended context window of 8K tokens, allowing for processing of longer audio and video inputs directly on local hardware (E016).

