Requests-per-minute is the wrong meter for LLM endpoints. One call can be 50 tokens or 50,000. Rate limit on input and output tokens with Zuplo’s complex-rate-limit-inbound policy and the real counts from each upstream response.
Need help?
Contact usRequests-per-minute is the wrong meter for LLM endpoints. One call can be 50 tokens or 50,000. Rate limit on input and output tokens with Zuplo’s complex-rate-limit-inbound policy and the real counts from each upstream response.