Skip to content

LLM Rate-Limit Planning Calculator

Find the sustainable request rate and the in-flight concurrency a provider rate limit allows, from its tokens-per-minute and requests-per-minute caps, your per-request token use, and average request latency.

Inputs

Provider limits

≥ 1
≥ 1

Workload

≥ 1
≥ 0.1

Results

Enter a value to see results.

Details

Your token limit binds first: the tokens-per-minute cap permits fewer requests than the requests-per-minute cap. To raise throughput, reduce tokens per request (shorter prompts, smaller completions, prompt caching) or request a higher TPM allocation.

Embed this calculator

Preview

Paste this code into your page to show the calculator.

Share this calculation

Anyone who opens this link sees your values already filled in.

200+ calculators · 10 languages · 100% free

Was this calculator helpful?

No

How can we improve this calculator?