Until now, building a search system that understands both what a document says and what its images, videos, or audio convey could be complex and limited. It meant stitching together multiple embedding models and writing code to reconcile results across modalities. That pipeline complexity may have been the single biggest barrier to production-grade multimodal retrieval.