What can I run · updated Oct 11, 2026

Local LLMs that run on a 8 GB graphics card

13 of the 56 models we've tested fit, with room for a normal-length chat. The best one is Granite 4.2 8B, scoring 77 out of 100.

#ModelScoreCodingDocumentsDownloadMemory
1 Granite 4.2 8B 77 77 78 Q4 7.3 GB
2 Qwen3.5-9B 72 47 97 Q4 7.5 GB
3 Qwen3 VL 8B Instruct 39 37 41 Q4 7.1 GB
4 Ministral 3 8B 2512 37 30 43 Q4 7.1 GB
5 Qwen2.5 7B Instruct
ollama run qwen2.5:7b
24 10 37 Q6 7.3 GB
6 Llama 3.3 8B Instruct (Q4, Mac)
ollama run hf.co/bartowski/allura-forge_Llama-3.3-8B-Instruct-GGUF:Q4_K_M
23 13 33 Q4 6 GB
7 Granite 4.0 Micro (Q4, Mac)
ollama run granite4:micro-h
22 13 31 Q4 3.2 GB
8 Ministral 3 3B 2512 21 10 32 Q8 5.6 GB
9 Llama 3.1 8B Instruct
ollama run llama3.1:8b
15 3 26 Q5 6.8 GB
10 Gemma 3 4B
ollama run gemma3:4b
14 3 25 Q8 5.7 GB
11 Reka Edge 8 0 16 Q6 7.5 GB
12 Llama 3.2 3B Instruct
ollama run llama3.2:3b
6 0 13 Q8 4.5 GB
13 Llama 3.2 1B Instruct (Q4, Mac)
ollama run llama3.2:1b-instruct-q4_K_M
5 0 9 Q4 1.9 GB

"Download" is the biggest version that fits on 8 GB graphics card; "Memory" is roughly what that version needs. Smaller versions also work and leave more room. Q4 or Q8?

Top models that don't fit

  • Qwen3.6 27B (97/100) needs about 19.6 GB even in its smallest version.
  • gpt-oss-120b (96/100) needs about 63.3 GB even in its smallest version.
  • Muse Glimmer 30B (94/100) needs about 19.1 GB even in its smallest version.
  • Ling 3.0 Flash (94/100) needs about 83.5 GB even in its smallest version.
  • GLM 4.5V (93/100) needs about 67.4 GB even in its smallest version.

Not sure how much memory you have? Here's how to check.