Find out what AI could save you — calculate your automation ROI for free in minutes
Yowox.

Free developer tool

Local LLM Hardware Calculator

Pick your GPU or Mac, see which open-weight models actually run — at what context, and how fast.

  • ✓ No signup
  • ✓ Runs in your browser
  • ✓ Sizes read from HuggingFace, refreshed weekly

Your hardware

What can this machine run?

Usable memory

Filter by what the model can do Only models that can: Read from each model's own configuration, not from marketing claims.
More options

96 of 116 models run on this

Model Quant Weights KV cache Total Max context Tokens/sec
apache-2.0 · file sizes Q4_K_M 17.4 GB 0.8 GB 19.5 GB 35k 103–309
see model card · file sizes Q4_K_M 17.4 GB 0.8 GB 19.5 GB 35k 103–309
see model card · file sizes Q4_K_M 17.4 GB 0.8 GB 19.5 GB 35k 103–309
Qwen3.5 27B Top 4
apache-2.0 · file sizes
Q4_K_M 16.8 GB 2.0 GB 20.1 GB 15k 27–57
Qwen3.6 27B Top 5
apache-2.0 · file sizes
Q4_K_M 16.8 GB 2.0 GB 20.1 GB 15k 27–57
Gemma · file sizes Q4_K_M 15.5 GB 2.9 GB 19.6 GB 8k 30–61
gemma · file sizes Q4_K_M 15.5 GB 2.9 GB 19.6 GB 8k 30–61
Gemma · file sizes Q4_K_M 15.4 GB 3.9 GB 20.6 GB 11k 30–62
gemma · file sizes Q4_K_M 15.4 GB 3.9 GB 20.6 GB 11k 30–62
other · file sizes Q4_K_M 15.4 GB 3.9 GB 20.6 GB 11k 30–62
apache-2.0 · file sizes Q4_K_M 15.9 GB 1.9 GB 19.0 GB 20k 70–210
apache-2.0 · file sizes Q4_K_M 13.3 GB 1.3 GB 15.8 GB 48k 34–71
apache-2.0 · file sizes Q4_K_M 13.3 GB 1.3 GB 15.8 GB 32k 34–71
apache-2.0 · file sizes Q4_K_M 13.3 GB 1.3 GB 15.8 GB 48k 34–71
apache-2.0 · file sizes Q4_K_M 13.3 GB 1.3 GB 15.8 GB 48k 34–71
other · file sizes Q4_K_M 12.4 GB 1.8 GB 15.3 GB 32k 37–77
see model card · file sizes Q4_K_M 10.9 GB 0.4 GB 12.3 GB 128k 42–87
see model card · file sizes Q8_0 14.6 GB 1.5 GB 17.4 GB 33k 31–65
see model card · file sizes Q8_0 14.6 GB 1.3 GB 17.1 GB 39k 31–65
apache-2.0 · file sizes Q8_0 14.6 GB 1.5 GB 17.4 GB 32k 31–65
apache-2.0 · file sizes Q8_0 14.6 GB 1.5 GB 17.4 GB 33k 31–65
apache-2.0 · file sizes Q8_0 14.6 GB 1.5 GB 17.4 GB 32k 31–65
apache-2.0 · file sizes Q8_0 14.6 GB 1.5 GB 17.4 GB 32k 31–65
see model card · file sizes Q8_0 14.6 GB 1.3 GB 17.1 GB 39k 31–65
apache-2.0 · file sizes Q8_0 11.8 GB 3.0 GB 15.9 GB 24k 39–81
Gemma · file sizes Q8_0 11.7 GB 3.0 GB 15.7 GB 24k 39–82
gemma · file sizes Q8_0 11.7 GB 3.0 GB 15.7 GB 24k 39–82
mit · file sizes F16 17.5 GB 0.3 GB 19.2 GB 32k 26–54
mit · file sizes F16 17.5 GB 0.3 GB 19.2 GB 32k 26–54
apache-2.0 · file sizes F16 17.1 GB 1.0 GB 19.5 GB 28k 27–55
Gemma · file sizes Q8_0 9.2 GB 2.6 GB 12.7 GB 8k 50–104
gemma · file sizes Q8_0 9.2 GB 2.6 GB 12.7 GB 8k 50–104
see model card · file sizes F16 16.6 GB 1.8 GB 19.6 GB 19k 28–57
apache-2.0 · file sizes F16 15.2 GB 1.3 GB 17.7 GB 35k 30–62
apache-2.0 · file sizes F16 15.3 GB 1.1 GB 17.6 GB 39k 30–62
see model card · file sizes F16 15.3 GB 1.1 GB 17.6 GB 39k 30–62
other · file sizes F16 14.9 GB 1.1 GB 17.3 GB 32k 31–64
see model card · file sizes F16 15.0 GB 1.0 GB 17.2 GB 46k 31–63
llama3 · file sizes Q8_0 8.0 GB 1.0 GB 9.8 GB 105k 58–120
Llama 3.1 · file sizes Q8_0 8.0 GB 1.0 GB 9.8 GB 105k 58–120
llama3.1 · file sizes Q8_0 8.0 GB 1.0 GB 9.8 GB 105k 58–120
see model card · file sizes F16 14.2 GB 0.4 GB 15.8 GB 122k 32–67
mit · file sizes F16 14.2 GB 0.4 GB 15.8 GB 122k 32–67
see model card · file sizes F16 14.2 GB 0.4 GB 15.8 GB 122k 32–67
apache-2.0 · file sizes F16 14.2 GB 0.4 GB 15.8 GB 32k 32–67
apache-2.0 · file sizes F16 14.2 GB 0.4 GB 15.8 GB 32k 32–67
apache-2.0 · file sizes F16 14.2 GB 0.4 GB 15.8 GB 32k 32–67
apache-2.0 · file sizes F16 14.2 GB 0.4 GB 15.8 GB 122k 32–67
apache-2.0 · file sizes F16 14.2 GB 0.4 GB 15.8 GB 32k 32–67
apache-2.0 · file sizes Q4_K_M 4.4 GB 0.4 GB 5.5 GB 4k 105–218
apache-2.0 · file sizes F16 14.0 GB 0.7 GB 15.9 GB 83k 33–68
apache-2.0 · file sizes Q4_K_M 4.2 GB 4.0 GB 8.9 GB 4k 110–228
apache-2.0 · file sizes F16 8.7 GB 0.3 GB 9.9 GB 128k 53–110
apache-2.0 · file sizes F16 8.1 GB 1.0 GB 10.0 GB 104k 57–118
other · file sizes F16 7.4 GB 1.3 GB 9.6 GB 84k 62–128
see model card · file sizes F16 7.5 GB 1.1 GB 9.5 GB 40k 61–127
see model card · file sizes F16 7.5 GB 1.1 GB 9.5 GB 97k 61–127
see model card · file sizes F16 7.5 GB 1.1 GB 9.5 GB 97k 61–127
Gemma · file sizes F16 7.2 GB 1.1 GB 9.2 GB 105k 63–131
see model card · file sizes F16 7.2 GB 1.1 GB 9.2 GB 105k 63–131
apache-2.0 · file sizes F16 6.3 GB 0.6 GB 7.8 GB 128k 72–150
apache-2.0 · file sizes F16 6.4 GB 0.8 GB 8.0 GB 146k 72–149
llama3 · file sizes F16 6.0 GB 0.9 GB 7.7 GB 128k 76–159
Llama 3.2 · file sizes F16 6.0 GB 0.9 GB 7.7 GB 128k 76–159
llama3.2 · file sizes F16 6.0 GB 0.9 GB 7.7 GB 128k 76–159
apache-2.0 · file sizes F16 5.7 GB 0.6 GB 7.1 GB 64k 80–166
other · file sizes F16 5.8 GB 0.3 GB 6.8 GB 32k 80–165
other · file sizes F16 5.8 GB 0.3 GB 6.8 GB 32k 80–165
other · file sizes F16 5.8 GB 0.3 GB 6.8 GB 32k 80–165
gemma · file sizes Q8_0 2.6 GB 0.8 GB 4.0 GB 8k 177–367
see model card · file sizes F16 3.8 GB 0.9 GB 5.4 GB 40k 121–251
apache-2.0 · file sizes F16 3.6 GB 0.4 GB 4.7 GB 256k 126–262
see model card · file sizes F16 3.3 GB 0.2 GB 4.2 GB 128k 138–286
cc-by-nc-4.0 · file sizes F16 3.3 GB 0.2 GB 4.2 GB 128k 138–286
apache-2.0 · file sizes F16 3.2 GB 1.5 GB 5.3 GB 8k 143–298
see model card · file sizes F16 3.2 GB 0.9 GB 4.7 GB 166k 143–296
apache-2.0 · file sizes F16 2.9 GB 0.2 GB 3.7 GB 32k 159–330
apache-2.0 · file sizes F16 2.9 GB 0.2 GB 3.7 GB 32k 159–330
apache-2.0 · file sizes F16 2.9 GB 0.2 GB 3.7 GB 32k 159–330
apache-2.0 · file sizes Q4_K_M 0.9 GB 0.2 GB 1.7 GB 4k 497–1033
Llama 3.2 · file sizes F16 2.3 GB 0.3 GB 3.2 GB 128k 198–411
llama3.2 · file sizes F16 2.3 GB 0.3 GB 3.2 GB 128k 198–411
Gemma · file sizes F16 1.9 GB 0.2 GB 2.7 GB 32k 245–508
gemma · file sizes F16 1.9 GB 0.2 GB 2.7 GB 32k 245–508
see model card · file sizes F16 1.4 GB 0.9 GB 2.9 GB 40k 325–674
apache-2.0 · file sizes F16 1.4 GB 0.4 GB 2.4 GB 256k 316–656
apache-2.0 · file sizes F16 0.9 GB 0.1 GB 1.6 GB 32k 492–1022
mit · file sizes Q4_K_M 18.7 GB 1.9 GB 22.0 GB 8k 25–51
apache-2.0 · file sizes Q4_K_M 18.4 GB 2.0 GB 21.8 GB 9k 25–52
see model card · file sizes Q4_K_M 18.4 GB 2.0 GB 21.8 GB 9k 25–52
see model card · file sizes Q4_K_M 18.5 GB 2.0 GB 21.9 GB 8k 25–51
apache-2.0 · file sizes Q4_K_M 18.5 GB 2.0 GB 21.9 GB 8k 25–51
apache-2.0 · file sizes Q4_K_M 18.5 GB 2.0 GB 21.9 GB 8k 25–51
apache-2.0 · file sizes Q4_K_M 18.5 GB 2.0 GB 21.9 GB 8k 25–51
apache-2.0 · file sizes Q4_K_M 18.5 GB 2.0 GB 21.9 GB 8k 25–51
apache-2.0 · file sizes Q4_K_M 18.1 GB 2.0 GB 21.5 GB 10k 25–52

Share or save this shortlist

The link and the text summary are free to copy. The PDF is a one-page report of what this machine runs, with the assumptions attached — a free newsletter subscription unlocks that once.

Share via

How the estimate works

Weights come from the actual files

Every size is the real byte count of the GGUF file on HuggingFace, read through its API and summed across shards where a quantisation is split. Nothing is estimated from parameter counts, and nothing is typed in by hand.

The KV cache is computed, not looked up

Cache cost per token is 2 × layers × KV heads × head dimension × bytes, taken from each model's own config.json. That is why grouped-query models hold far more context than their size suggests, and why the number here updates when a model does.

Speed follows memory bandwidth

At batch 1 the bottleneck is reading the weights once per token, so throughput is bandwidth divided by bytes read, at roughly three quarters of the vendor's peak figure. The result is always shown as a range.

Hardware figures are cross-checked

Bandwidth is checked against the vendor's published memory interface width: dividing one by the other has to land on a real memory speed grade. A transcription error lands between grades and is rejected before it reaches this page.

What can you run on 12, 16, 24, 48 and 64 GB?

The largest models that fit each memory size at 8k of context, with the best quantisation that still fits. Speed is a range, at batch 1.

12 GB — 70 of 116 models run

RTX 3060 12 GB, RTX 4070 and similar. Estimated from 360 GB/s of memory bandwidth.

Model Quant Total memory Max context Tokens/sec
Gemma 2 9B Q4_K_M 8.8 GB 8k 30–63
gemma 2 9b Instruct Q4_K_M 8.8 GB 8k 30–63
Qwen3.5 9B Q4_K_M 7.5 GB 36k 28–59
NVIDIA Nemotron Nano 9B v2 Q4_K_M 8.6 GB 18k 27–56
granite 3.1 8b instruct Q8_0 10.2 GB 13k 20–42

16 GB — 71 of 116 models run

RTX 5080, RTX 4080 Super, RTX 4070 Ti Super. Estimated from 896 GB/s of memory bandwidth.

Model Quant Total memory Max context Tokens/sec
gpt oss 20b Q4_K_M 12.3 GB 59k 37–78
DeepSeek R1 Distill Qwen 14B Q4_K_M 10.8 GB 28k 49–101
Qwen2.5 14B Instruct Q4_K_M 10.8 GB 28k 49–101
Qwen2.5 14B Instruct 1M Q4_K_M 10.8 GB 28k 49–101
Qwen2.5 Coder 14B Q4_K_M 10.8 GB 28k 49–101

24 GB — 96 of 116 models run

RTX 4090, RTX 3090, Radeon RX 7900 XTX. Estimated from 1008 GB/s of memory bandwidth.

Model Quant Total memory Max context Tokens/sec
Qwen3 30B A3B Q4_K_M 19.5 GB 35k 103–309
Qwen3 30B A3B Instruct 2507 Q4_K_M 19.5 GB 35k 103–309
Qwen3 VL 30B A3B Instruct Q4_K_M 19.5 GB 35k 103–309
Qwen3.5 27B Q4_K_M 20.1 GB 15k 27–57
Qwen3.6 27B Q4_K_M 20.1 GB 15k 27–57

48 GB — 101 of 116 models run

L40S, RTX A6000. Estimated from 864 GB/s of memory bandwidth.

Model Quant Total memory Max context Tokens/sec
Hermes 4.3 36B Q8_0 40.1 GB 24k 11–23
Qwen3.5 35B A3B Q8_0 38.1 GB 85k 51–152
Qwen3.6 35B A3B Q8_0 38.1 GB 85k 51–152
GLM Z1 Rumination 32B 0414 Q8_0 36.9 GB 38k 12–25
DeepSeek R1 Distill Qwen 32B Q8_0 36.6 GB 38k 12–25

64 GB unified — 115 of 116 models run

Apple Silicon with 64 GB or more. Estimated from 546 GB/s of memory bandwidth.

Model Quant Total memory Max context Tokens/sec
Mistral Medium 3.5 128B Q4_K_M 80.0 GB 54k 3–6
Qwen3.5 122B A10B Q4_K_M 77.3 GB 207k 14–43
gpt oss 120b Q4_K_M 62.6 GB 128k 4–8
Qwen3 Next 80B A3B Instruct Q8_0 84.3 GB 132k 28–83
Qwen3 Next 80B A3B Thinking Q8_0 84.3 GB 132k 28–83

Frequently asked questions

How much VRAM do I need to run an LLM? +

Enough for three things at once: the weight file, the KV cache, and about half a gigabyte of runtime buffers. A 32B model at Q4_K_M is roughly 18 GiB of weights, and its KV cache adds about 0.25 GiB for every 1,024 tokens of context. On a 24 GiB card that leaves room for roughly 16k tokens of context. Doubling the context doubles the cache, not the weights.

Can I run a 70B model locally? +

At Q4_K_M a 70B model is close to 40 GiB of weights, so it does not fit a single 24 GiB consumer card. It fits comfortably on 48 GiB and above, or on Apple Silicon with 64 GB or more of unified memory, where the memory is shared with the operating system. Below that the practical options are a 32B model at Q4_K_M or a mixture-of-experts model with far fewer active parameters.

How much memory does the KV cache use at long context? +

The cache holds a key and a value for every layer and every KV head, for every token in the window. That is 2 × layers × KV heads × head dimension × bytes per element. For Qwen3 32B, with 64 layers, 8 KV heads and a head dimension of 128, one token costs 262,144 bytes at fp16 — a quarter of a gibibyte per 1,024 tokens. At 128k context the cache alone is over 30 GiB, larger than the quantised weights.

Is Q4_K_M good enough, or should I use Q8_0? +

Q4_K_M costs measurable quality, but the loss is proportionally smaller on larger models. In the same memory budget a 32B model at Q4_K_M usually beats a 14B model at Q8_0 on general tasks. Q8_0 is worth it when the model is small, when output is parsed by another program, or when the task is arithmetic or code where small errors compound.

How many tokens per second will I actually get? +

Generation at batch 1 is limited by memory bandwidth: the runtime reads the weights once per token, so speed is roughly bandwidth divided by the bytes it must read. The figures here are a range rather than a point, because real throughput also depends on the runtime, flash attention, context length and thermal headroom. Prompt processing is a separate, compute-bound step and is not included.

Does Apple unified memory work as well as a discrete GPU? +

It trades speed for capacity. An M-series chip with 64 GB or more can hold models no consumer graphics card can, because the memory is shared rather than soldered to one device. But its bandwidth is lower than a high-end discrete card, so the same model generates more slowly. Unified memory is also shared with the operating system, so a smaller share is available than the headline figure suggests.

What about multiple GPUs or offloading to CPU? +

Neither is modelled here. Splitting a model across two cards adds interconnect cost that depends heavily on the link between them, and offloading layers to system RAM slows generation by an order of magnitude because those layers are read across PCIe. Both make the answer depend on details this calculator does not ask for, so it reports what a single device holds.

Running this in production rather than on a desk?

Choosing hardware is the easy half. We build and operate the agents that run on it — evaluation, guardrails, cost control and the parts that decide whether an AI workflow survives contact with real work.

See what we do