Calculator

How much model fits on my GPU?

Pick a computer or enter its memory and bandwidth. The list shows which open models fit, at what quantization, and how fast each one generates text.

  • Fits and speeds are estimates made from published specs and the rules below. How they're made
  • It runs in your browser. Nothing you pick or type is sent.

Calculator

Figwick is an AI assistant you control. Pick the models and set the rules.

Join the waitlist

Every model on common machines

Each cell is the estimated speed in tokens per second with an empty context, at the best quantization that fits a 32K context. A dash means the model doesn't fit. The calculator above has every machine and context length.

ModelLaptop, no GPU, 32 GB24 GB usable, 136.5 GB/sRTX 3060 12 GB12 GB usable, 360 GB/sRTX 5060 Ti 16 GB16 GB usable, 448 GB/sRTX 4090 24 GB24 GB usable, 1008 GB/sRTX 5090 32 GB32 GB usable, 1792 GB/sMac mini M5 Pro 64 GB48 GB usable, 307 GB/sRTX PRO 6000 96 GB96 GB usable, 1792 GB/sDGX Spark 128 GB112 GB usable, 273 GB/sRyzen AI Max+ 395 128 GB116 GB usable, 256 GB/sMac Studio M5 Ultra 512 GB384 GB usable, 1200 GB/s
Qwen3.8 Flash Next180B, 6B activeReleased Aug 2026Won't fitWon't fitWon't fitWon't fitWon't fitWon't fit359Q3, tight55Q351Q3113Q8
Qwen3.8 27B27BReleased Aug 20265.0Q4Won't fitWon't fit37Q449Q66.4Q837Q85.7Q85.4Q825Q8
Gemma 4 12B11.95BReleased Jun 20266.4Q825Q521Q848Q885Q815Q885Q813Q812Q857Q8
Qwen3.6 35B A3B35B, 3B activeReleased Apr 202655Q3Won't fitWon't fit404Q3503Q558Q8337Q851Q848Q8226Q8
Gemma 4 26B A4B25.2B, 3.8B activeReleased Apr 202630Q5Won't fit142Q3223Q5266Q846Q8266Q841Q838Q8178Q8
Gemma 4 31B30.7BReleased Apr 20265.3Q3Won't fitWon't fit39Q349Q55.6Q833Q85.0Q84.7Q822Q8
Gemma 4 E4B8B, 4.5B activeReleased Apr 202617Q845Q856Q8126Q8225Q839Q8225Q834Q832Q8151Q8
Mistral Small 4 119B119B, 6.5B activeReleased Mar 2026Won't fitWon't fitWon't fitWon't fitWon't fitWon't fit232Q531Q629Q6104Q8
Qwen3.5 4B4BReleased Mar 202619Q851Q863Q8142Q8253Q843Q8253Q839Q836Q8169Q8
Qwen3.5 9B9BReleased Mar 20268.6Q829Q628Q863Q8112Q819Q8112Q817Q816Q875Q8
Qwen3.5 122B A10B122B, 10B activeReleased Feb 2026Won't fitWon't fitWon't fitWon't fitWon't fitWon't fit176Q420Q619Q668Q8
GLM-4.7 Flash30B, 3B activeReleased Jan 202645Q4Won't fitWon't fit330Q4437Q658Q8337Q851Q848Q8226Q8
Nemotron 3 Nano 30B A3B30B, 3.5B activeReleased Dec 202538Q4Won't fit154Q3, tight282Q4374Q650Q8289Q844Q841Q8194Q8
Devstral Small 2 24B24BReleased Dec 20255.6Q4Won't fitWon't fit41Q455Q67.2Q842Q86.4Q86.0Q828Q8
Ministral 3 14B13.9BReleased Dec 20255.5Q8Won't fit32Q441Q873Q812Q873Q811Q810Q849Q8
gpt-oss 120B117B, 5.1B activeReleased Aug 2025Won't fitWon't fitWon't fitWon't fitWon't fitWon't fit378MXFP458MXFP454MXFP4253MXFP4
gpt-oss 20B21B, 3.6B activeReleased Aug 202535MXFP4Won't fit114MXFP4, tight256MXFP4456MXFP478MXFP4456MXFP469MXFP465MXFP4305MXFP4
Qwen3 Coder 30B A3B30.5B, 3.3B activeReleased Jul 202550Q3Won't fitWon't fit367Q3397Q653Q8307Q847Q844Q8205Q8
GLM-4.5 Air106B, 12B activeReleased Jul 2025Won't fitWon't fitWon't fitWon't fitWon't fitWon't fit126Q517Q616Q656Q8
Mistral Small 3.2 24B24BReleased Jun 20255.6Q4Won't fitWon't fit41Q455Q67.2Q842Q86.4Q86.0Q828Q8
Llama 4 Scout 17B-16E109B, 17B activeReleased Apr 2025Won't fitWon't fitWon't fitWon't fitWon't fitWon't fit89Q512Q611Q640Q8
Phi-4 mini3.8BReleased Feb 202520Q853Q867Q8150Q8266Q846Q8266Q841Q838Q8178Q8
DeepSeek R1 Distill Llama 8B8BReleased Jan 20259.6Q838Q532Q871Q8126Q822Q8126Q819Q818Q885Q8
DeepSeek R1 Distill Qwen 32B32.5BReleased Jan 2025Won't fitWon't fitWon't fitWon't fit54Q45.3Q831Q84.7Q84.4Q821Q8
Phi-414BReleased Dec 20245.5Q831Q3, tight27Q541Q872Q812Q872Q811Q810Q848Q8
Llama 3.3 70B Instruct70BReleased Dec 2024Won't fitWon't fitWon't fitWon't fitWon't fit5.3Q3, tight14Q82.2Q82.1Q89.7Q8
Llama 3.1 8B Instruct8BReleased Jul 20249.6Q838Q532Q871Q8126Q822Q8126Q819Q818Q885Q8

"Tight" means the model needs 90% to 100% of usable memory. How they're made

How the estimates work

Every fit and speed on this page is an estimate, calculated from published specs and the rules below. Real speeds depend on the engine, its settings and the prompt, and are often lower.

Memory

A model fits when its weights, its KV cache and an overhead for working space fit in usable memory:

weights   = parameters x bits per weight / 8
KV cache  = 2 bytes x tokens x layers that keep every token x 2 x KV heads x head size
            (+ the same for sliding-window layers, up to the window)
overhead  = 1 GiB + 5% of the weights
needed    = weights + KV cache + overhead
  • A model fits when it needs at most 90% of usable memory, which leaves room for the system and longer prompts. It’s tight at 90% to 100% and won’t fit above that.
  • Bits per weight are llama.cpp’s measured averages for its quantization types: Q8_0 8.50, Q6_K 6.56, Q5_K_M 5.70, Q4_K_M 4.89 and Q3_K_M 4.00 (llama.cpp quantize README). gpt-oss ships only in MXFP4. Its bits per weight are its checkpoint size divided by its parameter count.
  • “Best that fits” picks the highest of those types that fits, or else the highest that’s tight. It skips 16-bit, which needs twice the memory of 8-bit for little gain.
  • Layer counts, KV heads and head sizes come from each model’s config.json on Hugging Face. Layers with linear attention or a state-space design keep no KV cache, and models with latent attention cache one compressed vector per token.
  • The context is capped at the model’s own maximum.

Usable memory

Usable memory is the part of a computer’s memory a model can use:

MachineUsable memory
Graphics card (NVIDIA or AMD)All of its VRAM. The GPU’s own needs are in the overhead
Mac2/3 of memory at 32 GB or less, 3/4 above. This is macOS’s default limit for the GPU, and sudo sysctl iogpu.wired_limit_mb raises it
DGX Spark (GB10)Memory minus 16 GB, for the firmware and the system
Ryzen AI Max+ 395 (Strix Halo)Memory minus 12 GB under Linux, for the firmware and the system
No GPUMemory minus 8 GB, for the system and other apps

For a custom machine, you enter this figure.

Speed

Generating a token reads every active weight from memory once, plus the KV cache so far, so memory bandwidth sets the speed:

tokens per second                 = bandwidth x 0.6 / active weight bytes
tokens per second, window full    = bandwidth x 0.6 / (active weight bytes + KV cache)
  • 0.6 is the assumed share of the stated bandwidth that an engine reaches. It’s a round figure, not a measurement.
  • A mixture-of-experts model reads only its active parameters for each token, so a 35B model with 3B active can be faster than a dense 8B model.
  • The first figure is for an empty cache, at the start of a chat. The second is for the whole chosen context in the cache, the slowest case.
  • Both are speeds for generating the answer. Reading the prompt isn’t counted.

What’s left out

  • Running part of a model from system memory when it doesn’t fit on the graphics card.
  • Prompt processing, batching and several users at once.
  • Differences between engines (llama.cpp, MLX, vLLM and others), speculative decoding, and quantization formats other than the llama.cpp types above.
  • Vision encoders loaded beside a model, unless the model’s stated size includes them.

For memory, GB means GiB (2³⁰ bytes), as vendors use it for RAM and VRAM. For bandwidth, GB/s means 10⁹ bytes per second.

Sources (91)

Memory and bandwidth are from the vendors' spec pages. Model sizes and context lengths are from the model cards. Release dates are from the vendors' announcements and model cards.

  1. Apple, Mac mini - Technical Specifications - Apple
  2. ggml-org/llama.cpp (GitHub Discussions), Adjust VRAM/RAM split on Apple Silicon (Discussion #2182)
  3. Apple, MacBook Air - Tech Specs - Apple
  4. Apple, Mac mini - Technical Specifications - Apple
  5. Apple, Mac mini - Technical Specifications - Apple
  6. Apple, MacBook Pro - Tech Specs - Apple
  7. Apple, MacBook Pro - Tech Specs - Apple
  8. Apple, Mac Studio - Technical Specifications - Apple
  9. Apple, Mac Studio - Technical Specifications - Apple
  10. NVIDIA, GeForce RTX 3060 Family | NVIDIA
  11. TechPowerUp, NVIDIA GeForce RTX 3060 12 GB Specs | TechPowerUp GPU Database
  12. NVIDIA, GeForce RTX 5070 Family Graphics Cards | NVIDIA
  13. NVIDIA, NVIDIA RTX Blackwell GPU Architecture (whitepaper V1.1)
  14. NVIDIA, GeForce RTX 4060 Ti & 4060 Graphics Cards | NVIDIA
  15. TechPowerUp, NVIDIA GeForce RTX 4060 Ti 16 GB Specs | TechPowerUp GPU Database
  16. NVIDIA, GeForce RTX 5060 Family Graphics Cards | NVIDIA
  17. TechPowerUp, NVIDIA GeForce RTX 5060 Ti 16 GB Specs | TechPowerUp GPU Database
  18. NVIDIA, GeForce RTX 5080 Graphics Cards | NVIDIA
  19. NVIDIA, NVIDIA RTX Blackwell GPU Architecture (whitepaper V1.1)
  20. NVIDIA, 3090 & 3090 Ti Graphics Cards | NVIDIA GeForce
  21. NVIDIA, NVIDIA Ampere GA102 GPU Architecture (whitepaper V2)
  22. NVIDIA, NVIDIA RTX Blackwell GPU Architecture (whitepaper V1.1)
  23. NVIDIA, GeForce RTX 4090 Graphics Cards for Gaming | NVIDIA
  24. NVIDIA, GeForce RTX 5090 Graphics Cards | NVIDIA
  25. NVIDIA, NVIDIA RTX PRO 4000 Blackwell datasheet
  26. NVIDIA, NVIDIA RTX A6000 Datasheet
  27. NVIDIA, NVIDIA RTX 6000 Ada Generation datasheet
  28. NVIDIA, NVIDIA RTX PRO 6000 Blackwell Workstation Edition datasheet
  29. AMD, AMD Radeon™ AI PRO R9700
  30. NVIDIA, Hardware Overview, DGX Spark User Guide
  31. Framework, Configure Framework Desktop DIY Edition (AMD Ryzen AI Max)
  32. AMD, AMD Ryzen AI Halo Developer Platform with Ryzen AI Max+ 395 processor
  33. AMD, AMD Ryzen AI MAX+ 395 Processor: Breakthrough AI Performance in Thin and Light
  34. Framework, Configure Framework Desktop DIY Edition (AMD Ryzen AI Max)
  35. AMD, FAQs: AMD Variable Graphics Memory, VRAM, AI Model Sizes, Quantization, MCP and More!
  36. Lenovo, ThinkPad X1 Carbon Gen 13 PSREF Product Specifications Reference
  37. Tom's Hardware, We benchmarked Intel's Lunar Lake GPU with Core Ultra 9
  38. Qwen, Qwen/Qwen3.8-27B model card
  39. Qwen, Qwen/Qwen3.8-Flash-Next model card
  40. Qwen, Qwen/Qwen3.6-35B-A3B model card
  41. Qwen, Qwen/Qwen3.5-122B-A10B model card
  42. Qwen, Qwen/Qwen3.5-9B model card
  43. Qwen, Qwen/Qwen3.5-4B model card
  44. Qwen, Qwen/Qwen3-Coder-30B-A3B-Instruct model card
  45. Google, google/gemma-4-31B-it model card
  46. Google, google/gemma-4-26B-A4B-it model card
  47. Google, google/gemma-4-12B-it model card
  48. Google, google/gemma-4-E4B-it model card
  49. OpenAI, openai/gpt-oss-20b model card
  50. OpenAI, gpt-oss-120b & gpt-oss-20b Model Card (arXiv 2508.10925)
  51. OpenAI, openai/gpt-oss-120b model card
  52. Meta, Llama 3.1 model card
  53. Meta, Llama 3.3 model card
  54. Meta, Llama 4 model card
  55. Mistral AI, mistralai/Mistral-Small-3.2-24B-Instruct-2506 model card
  56. Mistral AI, mistralai/Mistral-Small-3.1-24B-Instruct-2503 model card
  57. Mistral AI, mistralai/Ministral-3-14B-Instruct-2512 model card
  58. Mistral AI, mistralai/Devstral-Small-2-24B-Instruct-2512 model card
  59. Mistral AI, mistralai/Mistral-Small-4-119B-2603 model card
  60. DeepSeek, deepseek-ai/DeepSeek-R1-Distill-Qwen-32B model card
  61. Qwen, Qwen/Qwen2.5-32B model card
  62. DeepSeek, deepseek-ai/DeepSeek-R1-Distill-Llama-8B model card
  63. Microsoft, microsoft/phi-4 model card
  64. Microsoft, microsoft/Phi-4-mini-instruct model card
  65. Z.ai, zai-org/GLM-4.5-Air model card
  66. Z.ai, zai-org/GLM-4.5 GitHub README
  67. Z.ai, zai-org/GLM-4.7-Flash model card
  68. Z.ai, Z.ai developer docs: GLM-4.7 overview
  69. NVIDIA, nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 model card
  70. Qwen, QwenLM/Qwen3.8 GitHub README, news
  71. Qwen, QwenLM/Qwen3.8-Flash-Next GitHub README, news
  72. Qwen, QwenLM/Qwen3.8 GitHub README, news
  73. Qwen, QwenLM/Qwen3.8 GitHub README, news
  74. Qwen, QwenLM/Qwen3.8 GitHub README, news
  75. Qwen, QwenLM/Qwen3-Coder commit 1e8a9dc, "Update README.md" (dated 2025-07-31 by GitHub)
  76. Google, Gemma 4: Our most capable open models to date
  77. Google, Introducing Gemma 4 12B
  78. OpenAI, OpenAI News RSS feed, "Introducing gpt-oss"
  79. Meta, Llama 3.1 model card
  80. Meta, Llama 3.3 model card
  81. Meta, Llama 4 model card
  82. Mistral AI, Mistral Docs changelog
  83. Mistral AI, Introducing Mistral 3
  84. Mistral AI, Introducing: Devstral 2 and Mistral Vibe CLI
  85. Mistral AI, Introducing Mistral Small 4
  86. DeepSeek, deepseek-ai/DeepSeek-R1 commit 23807ce, "Release DeepSeek-R1" (dated 2025-01-20 by GitHub)
  87. Microsoft, microsoft/phi-4 model card
  88. Microsoft, Empowering innovation: The next generation of the Phi family
  89. Z.ai, Z.ai developer docs: new releases
  90. Z.ai, Z.ai developer docs: new releases
  91. NVIDIA, nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 model card

Hear when Figwick opens.

Figwick is an AI assistant you control, built on leading cloud and open models. Pick the models, set the rules and use it from wherever you work.

Prefer email? Write to hello@figwick.com.

We use your details to run the waitlist and to write to you about Figwick. See our privacy notice.