Skip to content

Inference Throughput Calculator

Estimate the memory-bandwidth limit on LLM decoding speed from model size, weight precision, and accelerator memory bandwidth — the tokens-per-second roofline for a single stream.

Inputs

Model

Hardware

Results

Enter a value to see results.

Details

Embed this calculator

Preview

Paste this code into your page to show the calculator.

Share this calculation

Anyone who opens this link sees your values already filled in.

200+ calculators · 10 languages · 100% free

Was this calculator helpful?

No

How can we improve this calculator?