Calculator
How much model fits on my GPU?
Pick a computer or enter its memory and bandwidth. The list shows which open models fit, at what quantization, and how fast each one generates text.
- Fits and speeds are estimates made from published specs and the rules below. How they're made
- It runs in your browser. Nothing you pick or type is sent.
Calculator
Graphics card: all of its VRAM. Mac: 3/4 of its memory, or 2/3 at 32 GB or less. Bandwidth is on the spec sheet.
Fits needs up to 90% of usable memory. Tight needs 90% to 100%. Each speed is for generating the answer, first with an empty context, then with the whole context in use. How they're made
Figwick is an AI assistant you control. Pick the models and set the rules.
Join the waitlistEvery model on common machines
Each cell is the estimated speed in tokens per second with an empty context, at the best quantization that fits a 32K context. A dash means the model doesn't fit. The calculator above has every machine and context length.
| Model | Laptop, no GPU, 32 GB24 GB usable, 136.5 GB/s | RTX 3060 12 GB12 GB usable, 360 GB/s | RTX 5060 Ti 16 GB16 GB usable, 448 GB/s | RTX 4090 24 GB24 GB usable, 1008 GB/s | RTX 5090 32 GB32 GB usable, 1792 GB/s | Mac mini M5 Pro 64 GB48 GB usable, 307 GB/s | RTX PRO 6000 96 GB96 GB usable, 1792 GB/s | DGX Spark 128 GB112 GB usable, 273 GB/s | Ryzen AI Max+ 395 128 GB116 GB usable, 256 GB/s | Mac Studio M5 Ultra 512 GB384 GB usable, 1200 GB/s |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3.8 Flash Next180B, 6B activeReleased Aug 2026 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 359Q3, tight | 55Q3 | 51Q3 | 113Q8 |
| Qwen3.8 27B27BReleased Aug 2026 | 5.0Q4 | Won't fit | Won't fit | 37Q4 | 49Q6 | 6.4Q8 | 37Q8 | 5.7Q8 | 5.4Q8 | 25Q8 |
| Gemma 4 12B11.95BReleased Jun 2026 | 6.4Q8 | 25Q5 | 21Q8 | 48Q8 | 85Q8 | 15Q8 | 85Q8 | 13Q8 | 12Q8 | 57Q8 |
| Qwen3.6 35B A3B35B, 3B activeReleased Apr 2026 | 55Q3 | Won't fit | Won't fit | 404Q3 | 503Q5 | 58Q8 | 337Q8 | 51Q8 | 48Q8 | 226Q8 |
| Gemma 4 26B A4B25.2B, 3.8B activeReleased Apr 2026 | 30Q5 | Won't fit | 142Q3 | 223Q5 | 266Q8 | 46Q8 | 266Q8 | 41Q8 | 38Q8 | 178Q8 |
| Gemma 4 31B30.7BReleased Apr 2026 | 5.3Q3 | Won't fit | Won't fit | 39Q3 | 49Q5 | 5.6Q8 | 33Q8 | 5.0Q8 | 4.7Q8 | 22Q8 |
| Gemma 4 E4B8B, 4.5B activeReleased Apr 2026 | 17Q8 | 45Q8 | 56Q8 | 126Q8 | 225Q8 | 39Q8 | 225Q8 | 34Q8 | 32Q8 | 151Q8 |
| Mistral Small 4 119B119B, 6.5B activeReleased Mar 2026 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 232Q5 | 31Q6 | 29Q6 | 104Q8 |
| Qwen3.5 4B4BReleased Mar 2026 | 19Q8 | 51Q8 | 63Q8 | 142Q8 | 253Q8 | 43Q8 | 253Q8 | 39Q8 | 36Q8 | 169Q8 |
| Qwen3.5 9B9BReleased Mar 2026 | 8.6Q8 | 29Q6 | 28Q8 | 63Q8 | 112Q8 | 19Q8 | 112Q8 | 17Q8 | 16Q8 | 75Q8 |
| Qwen3.5 122B A10B122B, 10B activeReleased Feb 2026 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 176Q4 | 20Q6 | 19Q6 | 68Q8 |
| GLM-4.7 Flash30B, 3B activeReleased Jan 2026 | 45Q4 | Won't fit | Won't fit | 330Q4 | 437Q6 | 58Q8 | 337Q8 | 51Q8 | 48Q8 | 226Q8 |
| Nemotron 3 Nano 30B A3B30B, 3.5B activeReleased Dec 2025 | 38Q4 | Won't fit | 154Q3, tight | 282Q4 | 374Q6 | 50Q8 | 289Q8 | 44Q8 | 41Q8 | 194Q8 |
| Devstral Small 2 24B24BReleased Dec 2025 | 5.6Q4 | Won't fit | Won't fit | 41Q4 | 55Q6 | 7.2Q8 | 42Q8 | 6.4Q8 | 6.0Q8 | 28Q8 |
| Ministral 3 14B13.9BReleased Dec 2025 | 5.5Q8 | Won't fit | 32Q4 | 41Q8 | 73Q8 | 12Q8 | 73Q8 | 11Q8 | 10Q8 | 49Q8 |
| gpt-oss 120B117B, 5.1B activeReleased Aug 2025 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 378MXFP4 | 58MXFP4 | 54MXFP4 | 253MXFP4 |
| gpt-oss 20B21B, 3.6B activeReleased Aug 2025 | 35MXFP4 | Won't fit | 114MXFP4, tight | 256MXFP4 | 456MXFP4 | 78MXFP4 | 456MXFP4 | 69MXFP4 | 65MXFP4 | 305MXFP4 |
| Qwen3 Coder 30B A3B30.5B, 3.3B activeReleased Jul 2025 | 50Q3 | Won't fit | Won't fit | 367Q3 | 397Q6 | 53Q8 | 307Q8 | 47Q8 | 44Q8 | 205Q8 |
| GLM-4.5 Air106B, 12B activeReleased Jul 2025 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 126Q5 | 17Q6 | 16Q6 | 56Q8 |
| Mistral Small 3.2 24B24BReleased Jun 2025 | 5.6Q4 | Won't fit | Won't fit | 41Q4 | 55Q6 | 7.2Q8 | 42Q8 | 6.4Q8 | 6.0Q8 | 28Q8 |
| Llama 4 Scout 17B-16E109B, 17B activeReleased Apr 2025 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 89Q5 | 12Q6 | 11Q6 | 40Q8 |
| Phi-4 mini3.8BReleased Feb 2025 | 20Q8 | 53Q8 | 67Q8 | 150Q8 | 266Q8 | 46Q8 | 266Q8 | 41Q8 | 38Q8 | 178Q8 |
| DeepSeek R1 Distill Llama 8B8BReleased Jan 2025 | 9.6Q8 | 38Q5 | 32Q8 | 71Q8 | 126Q8 | 22Q8 | 126Q8 | 19Q8 | 18Q8 | 85Q8 |
| DeepSeek R1 Distill Qwen 32B32.5BReleased Jan 2025 | Won't fit | Won't fit | Won't fit | Won't fit | 54Q4 | 5.3Q8 | 31Q8 | 4.7Q8 | 4.4Q8 | 21Q8 |
| Phi-414BReleased Dec 2024 | 5.5Q8 | 31Q3, tight | 27Q5 | 41Q8 | 72Q8 | 12Q8 | 72Q8 | 11Q8 | 10Q8 | 48Q8 |
| Llama 3.3 70B Instruct70BReleased Dec 2024 | Won't fit | Won't fit | Won't fit | Won't fit | Won't fit | 5.3Q3, tight | 14Q8 | 2.2Q8 | 2.1Q8 | 9.7Q8 |
| Llama 3.1 8B Instruct8BReleased Jul 2024 | 9.6Q8 | 38Q5 | 32Q8 | 71Q8 | 126Q8 | 22Q8 | 126Q8 | 19Q8 | 18Q8 | 85Q8 |
"Tight" means the model needs 90% to 100% of usable memory. How they're made
How the estimates work
Every fit and speed on this page is an estimate, calculated from published specs and the rules below. Real speeds depend on the engine, its settings and the prompt, and are often lower.
Memory
A model fits when its weights, its KV cache and an overhead for working space fit in usable memory:
weights = parameters x bits per weight / 8
KV cache = 2 bytes x tokens x layers that keep every token x 2 x KV heads x head size
(+ the same for sliding-window layers, up to the window)
overhead = 1 GiB + 5% of the weights
needed = weights + KV cache + overhead
- A model fits when it needs at most 90% of usable memory, which leaves room for the system and longer prompts. It’s tight at 90% to 100% and won’t fit above that.
- Bits per weight are llama.cpp’s measured averages for its quantization types: Q8_0 8.50, Q6_K 6.56, Q5_K_M 5.70, Q4_K_M 4.89 and Q3_K_M 4.00 (llama.cpp quantize README). gpt-oss ships only in MXFP4. Its bits per weight are its checkpoint size divided by its parameter count.
- “Best that fits” picks the highest of those types that fits, or else the highest that’s tight. It skips 16-bit, which needs twice the memory of 8-bit for little gain.
- Layer counts, KV heads and head sizes come from each model’s
config.jsonon Hugging Face. Layers with linear attention or a state-space design keep no KV cache, and models with latent attention cache one compressed vector per token. - The context is capped at the model’s own maximum.
Usable memory
Usable memory is the part of a computer’s memory a model can use:
| Machine | Usable memory |
|---|---|
| Graphics card (NVIDIA or AMD) | All of its VRAM. The GPU’s own needs are in the overhead |
| Mac | 2/3 of memory at 32 GB or less, 3/4 above. This is macOS’s default limit for the GPU, and sudo sysctl iogpu.wired_limit_mb raises it |
| DGX Spark (GB10) | Memory minus 16 GB, for the firmware and the system |
| Ryzen AI Max+ 395 (Strix Halo) | Memory minus 12 GB under Linux, for the firmware and the system |
| No GPU | Memory minus 8 GB, for the system and other apps |
For a custom machine, you enter this figure.
Speed
Generating a token reads every active weight from memory once, plus the KV cache so far, so memory bandwidth sets the speed:
tokens per second = bandwidth x 0.6 / active weight bytes
tokens per second, window full = bandwidth x 0.6 / (active weight bytes + KV cache)
- 0.6 is the assumed share of the stated bandwidth that an engine reaches. It’s a round figure, not a measurement.
- A mixture-of-experts model reads only its active parameters for each token, so a 35B model with 3B active can be faster than a dense 8B model.
- The first figure is for an empty cache, at the start of a chat. The second is for the whole chosen context in the cache, the slowest case.
- Both are speeds for generating the answer. Reading the prompt isn’t counted.
What’s left out
- Running part of a model from system memory when it doesn’t fit on the graphics card.
- Prompt processing, batching and several users at once.
- Differences between engines (llama.cpp, MLX, vLLM and others), speculative decoding, and quantization formats other than the llama.cpp types above.
- Vision encoders loaded beside a model, unless the model’s stated size includes them.
For memory, GB means GiB (2³⁰ bytes), as vendors use it for RAM and VRAM. For bandwidth, GB/s means 10⁹ bytes per second.
Sources (91)
Memory and bandwidth are from the vendors' spec pages. Model sizes and context lengths are from the model cards. Release dates are from the vendors' announcements and model cards.
- Apple, Mac mini - Technical Specifications - Apple
- ggml-org/llama.cpp (GitHub Discussions), Adjust VRAM/RAM split on Apple Silicon (Discussion #2182)
- Apple, MacBook Air - Tech Specs - Apple
- Apple, Mac mini - Technical Specifications - Apple
- Apple, Mac mini - Technical Specifications - Apple
- Apple, MacBook Pro - Tech Specs - Apple
- Apple, MacBook Pro - Tech Specs - Apple
- Apple, Mac Studio - Technical Specifications - Apple
- Apple, Mac Studio - Technical Specifications - Apple
- NVIDIA, GeForce RTX 3060 Family | NVIDIA
- TechPowerUp, NVIDIA GeForce RTX 3060 12 GB Specs | TechPowerUp GPU Database
- NVIDIA, GeForce RTX 5070 Family Graphics Cards | NVIDIA
- NVIDIA, NVIDIA RTX Blackwell GPU Architecture (whitepaper V1.1)
- NVIDIA, GeForce RTX 4060 Ti & 4060 Graphics Cards | NVIDIA
- TechPowerUp, NVIDIA GeForce RTX 4060 Ti 16 GB Specs | TechPowerUp GPU Database
- NVIDIA, GeForce RTX 5060 Family Graphics Cards | NVIDIA
- TechPowerUp, NVIDIA GeForce RTX 5060 Ti 16 GB Specs | TechPowerUp GPU Database
- NVIDIA, GeForce RTX 5080 Graphics Cards | NVIDIA
- NVIDIA, NVIDIA RTX Blackwell GPU Architecture (whitepaper V1.1)
- NVIDIA, 3090 & 3090 Ti Graphics Cards | NVIDIA GeForce
- NVIDIA, NVIDIA Ampere GA102 GPU Architecture (whitepaper V2)
- NVIDIA, NVIDIA RTX Blackwell GPU Architecture (whitepaper V1.1)
- NVIDIA, GeForce RTX 4090 Graphics Cards for Gaming | NVIDIA
- NVIDIA, GeForce RTX 5090 Graphics Cards | NVIDIA
- NVIDIA, NVIDIA RTX PRO 4000 Blackwell datasheet
- NVIDIA, NVIDIA RTX A6000 Datasheet
- NVIDIA, NVIDIA RTX 6000 Ada Generation datasheet
- NVIDIA, NVIDIA RTX PRO 6000 Blackwell Workstation Edition datasheet
- AMD, AMD Radeon™ AI PRO R9700
- NVIDIA, Hardware Overview, DGX Spark User Guide
- Framework, Configure Framework Desktop DIY Edition (AMD Ryzen AI Max)
- AMD, AMD Ryzen AI Halo Developer Platform with Ryzen AI Max+ 395 processor
- AMD, AMD Ryzen AI MAX+ 395 Processor: Breakthrough AI Performance in Thin and Light
- Framework, Configure Framework Desktop DIY Edition (AMD Ryzen AI Max)
- AMD, FAQs: AMD Variable Graphics Memory, VRAM, AI Model Sizes, Quantization, MCP and More!
- Lenovo, ThinkPad X1 Carbon Gen 13 PSREF Product Specifications Reference
- Tom's Hardware, We benchmarked Intel's Lunar Lake GPU with Core Ultra 9
- Qwen, Qwen/Qwen3.8-27B model card
- Qwen, Qwen/Qwen3.8-Flash-Next model card
- Qwen, Qwen/Qwen3.6-35B-A3B model card
- Qwen, Qwen/Qwen3.5-122B-A10B model card
- Qwen, Qwen/Qwen3.5-9B model card
- Qwen, Qwen/Qwen3.5-4B model card
- Qwen, Qwen/Qwen3-Coder-30B-A3B-Instruct model card
- Google, google/gemma-4-31B-it model card
- Google, google/gemma-4-26B-A4B-it model card
- Google, google/gemma-4-12B-it model card
- Google, google/gemma-4-E4B-it model card
- OpenAI, openai/gpt-oss-20b model card
- OpenAI, gpt-oss-120b & gpt-oss-20b Model Card (arXiv 2508.10925)
- OpenAI, openai/gpt-oss-120b model card
- Meta, Llama 3.1 model card
- Meta, Llama 3.3 model card
- Meta, Llama 4 model card
- Mistral AI, mistralai/Mistral-Small-3.2-24B-Instruct-2506 model card
- Mistral AI, mistralai/Mistral-Small-3.1-24B-Instruct-2503 model card
- Mistral AI, mistralai/Ministral-3-14B-Instruct-2512 model card
- Mistral AI, mistralai/Devstral-Small-2-24B-Instruct-2512 model card
- Mistral AI, mistralai/Mistral-Small-4-119B-2603 model card
- DeepSeek, deepseek-ai/DeepSeek-R1-Distill-Qwen-32B model card
- Qwen, Qwen/Qwen2.5-32B model card
- DeepSeek, deepseek-ai/DeepSeek-R1-Distill-Llama-8B model card
- Microsoft, microsoft/phi-4 model card
- Microsoft, microsoft/Phi-4-mini-instruct model card
- Z.ai, zai-org/GLM-4.5-Air model card
- Z.ai, zai-org/GLM-4.5 GitHub README
- Z.ai, zai-org/GLM-4.7-Flash model card
- Z.ai, Z.ai developer docs: GLM-4.7 overview
- NVIDIA, nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 model card
- Qwen, QwenLM/Qwen3.8 GitHub README, news
- Qwen, QwenLM/Qwen3.8-Flash-Next GitHub README, news
- Qwen, QwenLM/Qwen3.8 GitHub README, news
- Qwen, QwenLM/Qwen3.8 GitHub README, news
- Qwen, QwenLM/Qwen3.8 GitHub README, news
- Qwen, QwenLM/Qwen3-Coder commit 1e8a9dc, "Update README.md" (dated 2025-07-31 by GitHub)
- Google, Gemma 4: Our most capable open models to date
- Google, Introducing Gemma 4 12B
- OpenAI, OpenAI News RSS feed, "Introducing gpt-oss"
- Meta, Llama 3.1 model card
- Meta, Llama 3.3 model card
- Meta, Llama 4 model card
- Mistral AI, Mistral Docs changelog
- Mistral AI, Introducing Mistral 3
- Mistral AI, Introducing: Devstral 2 and Mistral Vibe CLI
- Mistral AI, Introducing Mistral Small 4
- DeepSeek, deepseek-ai/DeepSeek-R1 commit 23807ce, "Release DeepSeek-R1" (dated 2025-01-20 by GitHub)
- Microsoft, microsoft/phi-4 model card
- Microsoft, Empowering innovation: The next generation of the Phi family
- Z.ai, Z.ai developer docs: new releases
- Z.ai, Z.ai developer docs: new releases
- NVIDIA, nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 model card
Hear when Figwick opens.
Figwick is an AI assistant you control, built on leading cloud and open models. Pick the models, set the rules and use it from wherever you work.
Prefer email? Write to hello@figwick.com.