LiteLLM now supports semantic prompt caching on Valkey clusters and AWS ElastiCache without requiring RediSearch or Qdrant. The feature lets the gateway reuse responses for semantically similar prompts to cut latency and cost.
Semantic Caching on Valkey and AWS ElastiCache
calendar_today
June 17, 2026
domain
litellm