Skip to content

Inference Latency Calculator

Estimate the end-to-end response time of a streaming LLM completion from time to first token, time per output token, and the number of output tokens.

Inputs

Serving Profile

≥ 1

Results

Enter a value to see results.

Details

Embed this calculator

Preview

Paste this code into your page to show the calculator.

Share this calculation

Anyone who opens this link sees your values already filled in.

200+ calculators · 10 languages · 100% free

Was this calculator helpful?

No

How can we improve this calculator?