Skip to content

LLM Inference VRAM Calculator

Estimate the GPU VRAM needed to serve a large language model for inference from its parameter count, weight precision, and runtime overhead for the KV cache, activations, and fragmentation.

Inputs

Model

Runtime

%
0 – 200 %

Results

Enter a value to see results.

Details

Embed this calculator

Preview

Paste this code into your page to show the calculator.

Share this calculation

Anyone who opens this link sees your values already filled in.

200+ calculators · 10 languages · 100% free

Was this calculator helpful?

No

How can we improve this calculator?