EmbeddingGemma 2 is an open multimodal embedding model from Google DeepMind, released on October 6, 2026. It maps text, code, images, video, and audio into a single 768-dimensional vector space and is built to run on phones and laptops, needing about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model with quantization on a Google Pixel 11 Pro. It has 740M total parameters: a 270M text backbone plus separately loadable 170M vision and 300M audio encoders. It ships under Apache 2.0 with an 8K token context window, support for 100+ languages, and Matryoshka truncation to 512, 256, or 128 dimensions.
A workable 0.74B-parameter dense embedding model from Google. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.0 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.0 GB |
| AMD Instinct MI300XAMD | SS | 1.0 GB |
| AMD Instinct MI325XAMD | SS | 1.0 GB |
| AMD Instinct MI355XAMD | SS | 1.0 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.07 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4070Vast.ai · Spot · 12 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3080Vast.ai · Spot · 10 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Google DeepMind's EmbeddingGemma 2 is a dense, 0.74-billion-parameter open embedding model released under the Apache 2.0 license. Built on the Gemma 4 architecture, it maps text, source code, images, audio, and video into a shared 768-dimensional vector space. Rather than locking developers into a massive monolith, the model uses a modular design totaling 740M parameters: a core 270M parameter text backbone paired with separately loadable 170M parameter vision and 300M parameter audio encoders.
This design makes EmbeddingGemma 2 one of the most versatile sub-1B embedding models available for local deployment. It solves a long-standing headache in privacy-first, on-device AI engineering: indexing and retrieving cross-modal data without relying on fragmented, disparate encoders or cloud endpoints. Whether you are building an offline semantic code search tool, an on-device personal photo and video vault, or an audio transcript search engine, EmbeddingGemma 2 delivers state-of-the-art retrieval accuracy at edge-tier resource consumption.
For practitioners looking to run EmbeddingGemma 2 locally, the model demands negligible system overhead. When quantized, the text-only configuration requires as little as 191MB of active RAM, while the fully loaded multimodal pipeline runs within approximately 567MB on mobile hardware like the Google Pixel 11 Pro. On consumer desktop GPUs and unified-memory Apple Silicon chips, it executes with near-instantaneous latency, setting a high standard for local AI model 0.74B parameters 2026 deployments.
EmbeddingGemma 2 relies on a dense transformer architecture derived directly from Gemma 4, structured cleanly into decoupled components:
The model provides an 8,192-token context window, a 4x increase over the original EmbeddingGemma. For local retrieval pipelines, an 8K window accommodates large technical documents, full code files, up to 5.5 minutes of continuous audio, or sequences of up to 58 video frames and 29 images.
Crucially for local vector stores, EmbeddingGemma 2 natively integrates Matryoshka Representation Learning (MRL). Developers can truncate the native 768-dimensional embeddings down to 512, 256, or 128 dimensions without retraining. Truncating to 128 dimensions reduces downstream storage and RAM consumption inside local vector databases like Qdrant, LanceDB, or Chroma by up to 6x, with only marginal loss in retrieval precision.
EmbeddingGemma 2 is engineered specifically for semantic search, asymmetric retrieval, zero-shot classification, clustering, and routing. Because it shares the Gemma 4 tokenizer and architectural conventions, it serves as an ideal retrieval frontend for local Retrieval-Augmented Generation (RAG) pipelines powered by small language models.
Traditional local setups require running CLIP for images, a Whisper-based transcription pipeline for audio, and a BERT-style model for text. EmbeddingGemma 2 eliminates this complexity. You can embed an entire mixed-media library and query it using plain natural language:
With a specialized training emphasis on source code, EmbeddingGemma 2 jumps from 68.76 to 78.68 on the MTEB Code benchmark compared to its predecessor. This makes it an exceptional engine for developer tooling, allowing coding agents to perform semantic code navigation, dependency linking, and symbol search across multi-language repositories entirely offline.
The model utilizes lightweight instruction prefixes to tune its vector projections based on the target workload. By prepending prompts with specific task instructions (such as querying for document retrieval, semantic textual similarity, or clustering), engineers can direct the vector geometry to optimize specifically for retrieval recall or semantic clustering separation.
EmbeddingGemma 2 is small enough to run on almost any modern consumer device, from mid-tier laptops to high-end development workstations. Because the encoders can be loaded independently, hardware requirements scale based on which modalities you deploy.
Memory requirements depend on the loaded encoders and execution precision:
| Configuration | Parameters | FP16 / BF16 VRAM | INT8 (Q8_0) VRAM | INT4 (Q4_K_M) VRAM |
|---|---|---|---|---|
| Text & Code Only | 270M | ~0.6 GB | ~0.35 GB | ~0.20 GB |
| Text + Vision | 440M | ~0.95 GB | ~0.55 GB | ~0.32 GB |
| Full Multimodal | 740M | ~1.6 GB | ~0.95 GB | ~0.55 GB |
These ultra-low memory requirements mean that running out of VRAM is virtually impossible on modern consumer cards. Even when embedding batches of large 8K-token documents, total system memory usage rarely exceeds 2 GB.
When evaluating how to run 0.74B model on consumer GPU setups, virtually any hardware platform succeeds:
llama.cpp with near-zero power draw, processing text and multimodal tokens instantly in unified memory.For typical text-only tasks, running the model in unquantized FP16 or BF16 is recommended if you have any modern discrete GPU or Apple Silicon Mac. At 270M parameters, the model is so compact that quantization saves negligible physical space on a workstation while FP16 preserves exact vector fidelity.
However, if you are targeting embedded hardware, mobile devices, or running dense multimodal workloads alongside large generative LLMs, Q4_K_M (or INT4 via LiteRT) is the best quantization for EmbeddingGemma 2. It preserves benchmark retrieval accuracy within 1-2% of full precision while shrinking the active multimodal footprint down to roughly 567MB.
EmbeddingGemma 2 performance is exceptionally fast. On an Nvidia RTX 4090 running through vLLM or optimized C++ backends, EmbeddingGemma 2 tokens per second can easily exceed 20,000 to 35,000 tokens per second under parallel batching for text ingestion. On an Apple M4 Max using MLX, document chunking and embedding latency typically measures in single-digit milliseconds per passage.
The quickest way to run EmbeddingGemma 2 locally for text workloads is via Ollama:
1ollama run embeddinggemma-2
For Python-based production pipelines, use Hugging Face transformers or sentence-transformers:
1from sentence_transformers import SentenceTransformer23# Load the modular text backbone4model = SentenceTransformer("google/embeddinggemma-2")56# Embed documentation chunks or queries7embeddings = model.encode([8 "def calculate_vector_distance(a, b):",9 "Optimizing vector indices for edge devices."10], prompt_name="query")
For cross-platform on-device deployment including vision and audio modalities, Google's LiteRT and MediaPipe Task APIs provide pre-compiled packages targeting mobile, WebGPU, and desktop environments.
EmbeddingGemma 1 was restricted strictly to text, featured an 8K-token limit disadvantage with its 2K context window, and struggled with complex technical repositories. EmbeddingGemma 2 expands context to 8,192 tokens, integrates vision and audio modalities into the exact same vector space, and provides a massive leap in MTEB Code performance (jumping nearly 10 points to 78.68). Unless constrained by a legacy pipeline, there is no technical justification to remain on version 1.
Nomic Embed Text v1.5 is a standard workhorse for lightweight local text embeddings, featuring 137M parameters, an 8K context window, and Matryoshka dimension truncation. While Nomic Embed is smaller and highly competent for simple English text search, EmbeddingGemma 2 outperforms it significantly on multi-language corpora, advanced code retrieval, and cross-modal operations. If your application handles strictly plain English prose, Nomic remains an ultra-lean choice; for multimodal search, complex codebases, or multilingual workloads, EmbeddingGemma 2 is the superior model.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Google model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.