Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google
Google DeepMind has introduced EmbeddingGemma 2, an open multimodal embedding model. This model maps text, images, video, and audio inputs into a unified 768-dimensional vector space. With 740M parameters, it combines a 270M parameter text model with modular vision (170M) and audio (300M) encoders. Designed for consumer hardware, EmbeddingGemma 2 provides low-latency semantic representations for on-device applications such as search, RAG, classification, and clustering.
This model is the first to natively support multimodal embeddings across text, images, video, and audio within a single 768-dimensional vector space.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 6, 2026, 23:00 UTC
- Ingested
- Oct 6, 2026, 23:00
- Source type
- Dev community
Full text isn't available here.
Read at source →