| # | Model | Score | Coding | Documents | Download | Memory |
|---|---|---|---|---|---|---|
| 1 | Granite 4.2 8B | 77 | 77 | 78 | Q4 | 7.3 GB |
| 2 | Qwen3.5-9B | 72 | 47 | 97 | Q4 | 7.5 GB |
| 3 | Qwen3 VL 8B Instruct | 39 | 37 | 41 | Q4 | 7.1 GB |
| 4 | Ministral 3 8B 2512 | 37 | 30 | 43 | Q4 | 7.1 GB |
| 5 | Qwen2.5 7B Instructollama run qwen2.5:7b | 24 | 10 | 37 | Q6 | 7.3 GB |
| 6 | Llama 3.3 8B Instruct (Q4, Mac)ollama run hf.co/bartowski/allura-forge_Llama-3.3-8B-Instruct-GGUF:Q4_K_M | 23 | 13 | 33 | Q4 | 6 GB |
| 7 | Granite 4.0 Micro (Q4, Mac)ollama run granite4:micro-h | 22 | 13 | 31 | Q4 | 3.2 GB |
| 8 | Ministral 3 3B 2512 | 21 | 10 | 32 | Q8 | 5.6 GB |
| 9 | Llama 3.1 8B Instructollama run llama3.1:8b | 15 | 3 | 26 | Q5 | 6.8 GB |
| 10 | Gemma 3 4Bollama run gemma3:4b | 14 | 3 | 25 | Q8 | 5.7 GB |
| 11 | Reka Edge | 8 | 0 | 16 | Q6 | 7.5 GB |
| 12 | Llama 3.2 3B Instructollama run llama3.2:3b | 6 | 0 | 13 | Q8 | 4.5 GB |
| 13 | Llama 3.2 1B Instruct (Q4, Mac)ollama run llama3.2:1b-instruct-q4_K_M | 5 | 0 | 9 | Q4 | 1.9 GB |
"Download" is the biggest version that fits on 8 GB graphics card; "Memory" is roughly what that version needs. Smaller versions also work and leave more room. Q4 or Q8?
Top models that don't fit
- Qwen3.6 27B (97/100) needs about 19.6 GB even in its smallest version.
- gpt-oss-120b (96/100) needs about 63.3 GB even in its smallest version.
- Muse Glimmer 30B (94/100) needs about 19.1 GB even in its smallest version.
- Ling 3.0 Flash (94/100) needs about 83.5 GB even in its smallest version.
- GLM 4.5V (93/100) needs about 67.4 GB even in its smallest version.
Not sure how much memory you have? Here's how to check.