EmbeddingGemma 2 Brings Multimodal Search Into Transformers.js
Transformers.js 4.3.1 can embed text, images, audio, and video in one space, but the encoder you load determines the browser cost.
Transformers.js 4.3.1, released October 7, added EmbeddingGemma 2. I care less about another model name than the new search shape: a text query can rank text, images, audio, and video in the same 768-dimensional embedding space, inside the browser. The release’s example embeds a query and three media types, then scores them with a dot product.
The model is 740 million parameters in total, split across a 270M text model, a 170M vision encoder, and a 300M audio encoder. Video frames use the vision encoder. You do not need all three parts for every app. A code search UI needs text; a photo library needs text and vision; a voice archive needs text and audio.
For text-only search, the short path is pipeline('feature-extraction', 'onnx-community/embeddinggemma-2-ONNX', { device: 'webgpu', dtype: 'q4' }). Encode queries with the prefix task: search result | query: and documents as title: {title} | text: {content}. Ask for pooling: 'mean' and normalize: true. Then rank the resulting vectors by dot product. Those prefixes are part of the model’s retrieval setup, not decorative prompt text.
There is a catch in that short path: the feature-extraction pipeline loads every encoder. The ONNX model card shows a lower-level route for a smaller deployment. Load AutoConfig, set vision_config and audio_config to null for text-only use, and pass that config to AutoModel.from_pretrained. Keep vision_config for images or video; keep audio_config for sound. This is the difference between choosing a task and choosing the actual model footprint.
Quantization is another choice with visible cost. The model card lists the full multimodal q4 weights at 473 MB versus 2,929 MB for fp32. Those are weight sizes, not a promise about total memory or load time on your users’ devices. The card recommends q4 on WebGPU for browser use, and suggests q8 for audio when quality matters more than download size.
For mixed-media search, use AutoProcessor with AutoModel instead of the text pipeline. The processor accepts text, images, audio, and video in separate argument positions and returns inputs for one sentence_embedding output. The release example decodes audio to mono 16 kHz and samples video at one frame per second. That preprocessing is part of your search design: a long clip does not become a useful index just because it fits through one model call.
I would start with a small local index and measure retrieval quality on the actual collection before adding every modality. Keep only the encoders your product uses, choose a quantization level with real examples, and remember that embeddings retrieve candidates; they do not verify the answer built from them. Sources: Transformers.js 4.3.1 release and EmbeddingGemma 2 ONNX model card.