Imagine you run a post office that sorts letters by zip code, packages by weight, and postcards by destination — three separate sorting lines, three crews, three sets of bins. Now imagine someone builds a single conveyor belt that reads all three formats and drops everything into one unified grid of slots. That's what EmbeddingGemma 2 does for on-device multimodal search: it replaces separate text, image, and audio embedding models with a single 740M-parameter encoder that maps all modalities into one shared vector space. The committed claim: EmbeddingGemma 2 is the best-in-class multimodal embedding model under 1 billion parameters, achieving leading scores on MTEB Code and MAEB while matching or beating specialist models more than twice its size. The text backbone starts at 270M parameters, with optional vision (170M) and audio (300M) encoders bolted on — a modular design that lets developers pay only for the modalities they need. Matryoshka Representation Learning allows truncating output vectors from 768 to 128 dimensions, yielding up to 6x storage savings for on-device vector databases. On the ladder, the numbers that matter are the 9.92-point improvement over EmbeddingGemma 1 on MTEB Code (68.76 → 78.68), and the claim of outperforming specialist models more than 2x its size on image, video, document, and audio tasks. But the blog post conspicuously omits naming those larger models or providing head-to-head numbers beyond code. The model card is referenced but not reproduced here, so the reader is asked to take the headline claims on trust. For text, they say performance 'matches' v1 — which means multimodal capability came at zero text regression, a real engineering achievement if the model card confirms it. Architecturally, this is a Gemma 4 backbone with modality-specific encoder heads — a now-standard design pattern where a shared transformer processes tokenized representations from each modality. The audio encoder is shared with Gemma 4 itself, reducing combined memory footprint when both models run in a RAG pipeline. The 8K token context window (4x v1) translates to 5.5 minutes of audio or 29 images or 58 video frames — serious on-device throughput. Quantized, the full multimodal model fits in ~567MB of active RAM on a Pixel 11 Pro. Integrity is mixed. They cite community benchmarks (MTEB, MAEB), which is good — these are standardized and externally maintained. But the blog post is a product announcement, not a peer-reviewed paper. No ablation studies, no training data disclosure, no pre-registration, and the competitive comparisons are vague ('outperforms models more than twice its size' without naming them). The model weights are Apache 2.0 on Hugging Face and Kaggle, which means independent replication is at least possible. The 20 million downloads figure for v1 is a real adoption signal but tells you nothing about quality. The milestone to watch is whether this model — or its successors — can close the gap with server-side multimodal embedders like those behind Google's own Gemini API. Right now, sub-1B on-device models accept meaningful quality trade-offs for latency and privacy. The question is whether the next generation (likely ~2B parameters, running on 2027-era phone silicon) erases enough of that gap to make cloud embedding APIs redundant for most consumer use cases. We're probably 2-3 hardware generations away. The obvious experiment not run: cross-modal retrieval benchmarks at scale against named server-side models. How much quality do you actually sacrifice going on-device? Google almost certainly has these numbers internally — the Gemini Embedding models are the parent technology. The omission reads as strategic: this is a product launch aimed at developers who've already decided on-device is worth the trade-off. Publishing the gap would undercut the pitch. Fair enough, but it means the reader can't make a fully informed architecture decision from this announcement alone.