What can I run · updated Oct 11, 2026

Local LLMs that run on a 12 GB graphics card

18 of the 56 models we've tested fit, with room for a normal-length chat. The best one is Granite 4.2 8B, scoring 77 out of 100.

#ModelScoreCodingDocumentsDownloadMemory
1 Granite 4.2 8B 77 77 78 Q8 11.3 GB
2 Qwen3 14B
ollama run qwen3:14b
72 63 81 Q4 10.9 GB
3 Qwen3.5-9B 72 47 97 Q6 9.6 GB
4 Ministral 3 14B 2512 44 43 45 Q4 10.4 GB
5 Qwen3 VL 8B Instruct 39 37 41 Q8 11.1 GB
6 Ministral 3 8B 2512 37 30 43 Q8 11.2 GB
7 Gemma 3 12B
ollama run gemma3:12b
35 30 40 Q5 10.3 GB
8 Phi 4
ollama run phi4
25 23 26 Q4 11.2 GB
9 Qwen2.5 7B Instruct
ollama run qwen2.5:7b
24 10 37 Q8 9.2 GB
10 Llama 3.3 8B Instruct (Q4, Mac)
ollama run hf.co/bartowski/allura-forge_Llama-3.3-8B-Instruct-GGUF:Q4_K_M
23 13 33 Q4 6 GB
11 Mistral Nemo
ollama run mistral-nemo
23 10 36 Q5 10.7 GB
12 Granite 4.0 Micro (Q4, Mac)
ollama run granite4:micro-h
22 13 31 Q4 3.2 GB
13 Ministral 3 3B 2512 21 10 32 Q8 5.6 GB
14 Llama 3.1 8B Instruct
ollama run llama3.1:8b
15 3 26 Q8 9.6 GB
15 Gemma 3 4B
ollama run gemma3:4b
14 3 25 Q8 5.7 GB
16 Reka Edge 8 0 16 Q8 9.2 GB
17 Llama 3.2 3B Instruct
ollama run llama3.2:3b
6 0 13 Q8 4.5 GB
18 Llama 3.2 1B Instruct (Q4, Mac)
ollama run llama3.2:1b-instruct-q4_K_M
5 0 9 Q4 1.9 GB

"Download" is the biggest version that fits on 12 GB graphics card; "Memory" is roughly what that version needs. Smaller versions also work and leave more room. Q4 or Q8?

Top models that don't fit

  • Qwen3.6 27B (97/100) needs about 19.6 GB even in its smallest version.
  • gpt-oss-120b (96/100) needs about 63.3 GB even in its smallest version.
  • Muse Glimmer 30B (94/100) needs about 19.1 GB even in its smallest version.
  • Ling 3.0 Flash (94/100) needs about 83.5 GB even in its smallest version.
  • GLM 4.5V (93/100) needs about 67.4 GB even in its smallest version.

Not sure how much memory you have? Here's how to check.