At 10,000 queries per second with ~30-token queries, you’re pushing ~18 million tokens per minute through your embedding API. At $0.02 per million tokens, that’s over $15,000/month — just for query embeddings. Documents are embedded once. Queries are embedded forever. What if you could drop that to $0? That’s the promise of asymmetric retrieval : embed your documents with the best model money can