Learn how to configure speculative decoding on vLLM — including draft model selection, memory budgeting, quantization tradeoffs, and when to disable it based on your query rate and acceptance rate data.
Need help?
Contact usLearn how to configure speculative decoding on vLLM — including draft model selection, memory budgeting, quantization tradeoffs, and when to disable it based on your query rate and acceptance rate data.