Rankings · updated OCT 11 · 56 models

Best local LLMs: October 2026

A local LLM is an AI chatbot you download and run on your own computer. No subscription, and nothing you type leaves your machine. Here's the best one for the computer you have.

The short version
  • The best model we've tested is Qwen3.6 27B, with 97 out of 100.
  • Close behind: gpt-oss-120b (96).
  • Less memory? Find your graphics card or Mac below. We list the best model that fits and which version to download.

Pick by graphics card

Not sure how much memory your card has? On Windows, open Task Manager, go to Performance and click GPU. More help here.

8 GB graphics card
Get the Q4 version · scored 77
Also good: Qwen3.5-9B
12 GB graphics card
Get the Q8 version · scored 77
Also good: Qwen3 14B
16 GB graphics card
Get the standard version · scored 88
Also good: Granite 4.2 8B
24 GB graphics card
Get the Q5 version · scored 97
Also good: Muse Glimmer 30B
32 GB graphics card
Get the Q6 version · scored 97
Also good: Muse Glimmer 30B

Pick by Mac

Click the Apple menu, then About This Mac. The number next to Memory is what you need.

Mac with 16 GB
Get the Q6 version · scored 77
Also good: Qwen3.5-9B
Mac with 24 GB
Get the standard version · scored 88
Also good: Granite 4.2 8B
Mac with 32 GB
Get the Q4 version · scored 97
Also good: Muse Glimmer 30B
Mac with 48 GB
Get the Q8 version · scored 97
Also good: Muse Glimmer 30B
Mac with 64 GB
Get the Q8 version · scored 97
Also good: Muse Glimmer 30B
Mac with 96 GB
Get the Q8 version · scored 97
Also good: gpt-oss-120b
Mac with 128 GB
Get the Q8 version · scored 97
Also good: gpt-oss-120b

All models we've tested

# Model Score CodingDocuments Hard tasks Runs on (PC) Runs on (Mac)
1 Qwen3.6 27B
28B
97 9797 91 24 GB card (Q5) 32 GB Mac (Q4)
2 gpt-oss-120b
117B · gpt-oss:120b
96 9399 94 more than 32 GB 96 GB Mac (standard)
3 Muse Glimmer 30B
30B
94 9098 89 24 GB card (Q5) 32 GB Mac (Q4)
4 Ling 3.0 Flash
127B
94 9395 91 more than 32 GB 128 GB Mac (Q4)
5 GLM 4.5V
108B
93 8799 89 more than 32 GB 96 GB Mac (Q4)
6 Ling 3.0 Flash VL
125B
92 8798 83 more than 32 GB 128 GB Mac (Q5)
7 gpt-oss-20b
21B · gpt-oss:20b
88 9087 69 16 GB card (standard) 24 GB Mac (standard)
8 Qwen3.5-27B
28B
88 77100 78 24 GB card (Q5) 32 GB Mac (Q4)
9 Qwen3.8 27B
28B
88 7799 83 24 GB card (Q5) 32 GB Mac (Q4)
10 Qwen3.5-35B-A3B
36B
88 8095 73 24 GB card (Q4) 48 GB Mac (Q6)
11 Qwen3.6 35B A3B
36B
86 8093 73 24 GB card (Q4) 48 GB Mac (Q6)
12 Laguna XS 2.1
33B
82 7391 56 24 GB card (Q4) 48 GB Mac (Q6)
13 Nemotron 3 Nano 30B A3B
32B
80 7091 69 24 GB card (Q5) 32 GB Mac (Q4)
14 Gemma 4 31B
31B
79 9365 73 32 GB card (Q5) 48 GB Mac (Q6)
15 Granite 4.2 8B
8.8B
77 7778 55 8 GB card (Q4) 16 GB Mac (Q6)
16 Qwen3 30B A3B
31B · qwen3:30b
77 6787 50 24 GB card (Q5) 32 GB Mac (Q4)
17 Qwen3 VL 30B A3B Thinking
31B
77 6093 51 24 GB card (Q5) 32 GB Mac (Q4)
18 GLM 4.5 Air
110B
76 7082 52 more than 32 GB 96 GB Mac (Q4)
19 GLM 4.6V
108B
73 7373 53 more than 32 GB 96 GB Mac (Q4)
20 Qwen3 14B
15B · qwen3:14b
72 6381 42 12 GB card (Q4) 24 GB Mac (Q6)
21 GLM 4.7 Flash
31B
72 7370 43 24 GB card (Q4) 48 GB Mac (Q6)
22 Qwen3.5-9B
9.7B
72 4797 58 8 GB card (Q4) 16 GB Mac (Q6)
23 Gemma 4 26B A4B
26B
70 8358 50 24 GB card (Q5) 32 GB Mac (Q5)
24 Qwen3 32B
33B · qwen3:32b
64 5079 50 24 GB card (Q4) 48 GB Mac (Q6)
25 Laguna S 2.1
118B
64 7059 46 more than 32 GB 128 GB Mac (Q5)
26 Nemotron 3.5 Lightning
32B
57 4372 43 24 GB card (Q5) 32 GB Mac (Q4)
27 Nex-N2.5-Mini
35B
57 2788 40 24 GB card (Q4) 48 GB Mac (Q6)
28 Devstral 2 2512
125B
53 5353 24 more than 32 GB 128 GB Mac (Q5)
29 Hunyuan A13B Instruct
80B
52 3767 28 more than 32 GB 96 GB Mac (Q6)
30 Qwen3 Coder 30B A3B Instruct
31B · qwen3-coder:30b
51 5746 42 24 GB card (Q5) 32 GB Mac (Q4)
31 Qwen3 VL 30B A3B Instruct
31B
46 4349 33 24 GB card (Q5) 32 GB Mac (Q4)
32 Qwen2.5 72B Instruct
73B
46 4052 23 more than 32 GB 64 GB Mac (Q4)
33 Llama 3.3 70B Instruct
71B
46 4745 27 more than 32 GB 64 GB Mac (Q4)
34 Command A
111B
45 3356 27 more than 32 GB 96 GB Mac (Q4)
35 Ministral 3 14B 2512
14B
44 4345 24 12 GB card (Q4) 16 GB Mac (Q4)
36 Mistral Small 3.2 24B
24B
44 4047 21 24 GB card (Q6) 32 GB Mac (Q5)
37 Mistral Small 3
24B
41 3745 24 24 GB card (Q6) 32 GB Mac (Q5)
38 Gemma 3 27B
27B · gemma3:27b
40 3346 29 24 GB card (Q5) 32 GB Mac (Q4)
39 Qwen3 30B A3B Instruct 2507
31B
40 3346 25 24 GB card (Q5) 32 GB Mac (Q4)
40 Qwen3 VL 8B Instruct
8.8B
39 3741 25 8 GB card (Q4) 16 GB Mac (Q6)
41 Llama 3.1 70B Instruct
71B
39 3344 22 more than 32 GB 64 GB Mac (Q4)
42 Ministral 3 8B 2512
8.9B
37 3043 18 8 GB card (Q4) 16 GB Mac (Q6)
43 Gemma 3 12B
12B · gemma3:12b
35 3040 24 12 GB card (Q5) 16 GB Mac (Q5)
44 Phi 4
15B · phi4
25 2326 19 12 GB card (Q4) 24 GB Mac (Q6)
45 Qwen2.5 7B Instruct
7.6B · qwen2.5:7b
24 1037 11 8 GB card (Q6) 16 GB Mac (Q8)
46 Gemma 2 27B
27B · gemma2:27b
23 1334 12 24 GB card (Q5) 32 GB Mac (Q4)
47 Llama 3.3 8B Instruct (Q4, Mac)
8B · hf.co/bartowski/allura-forge_Llama-3.3-8B-Instruct-GGUF:Q4_K_M
23 1333 10 8 GB card (Q4) 16 GB Mac (Q4)
48 Mistral Nemo
12B · mistral-nemo
23 1036 13 12 GB card (Q5) 16 GB Mac (Q4)
49 Granite 4.0 Micro (Q4, Mac)
3.2B · granite4:micro-h
22 1331 14 8 GB card (Q4) 16 GB Mac (Q4)
50 Ministral 3 3B 2512
3.9B
21 1032 15 8 GB card (Q8) 16 GB Mac (Q8)
51 Llama 3.1 8B Instruct
8B · llama3.1:8b
15 326 3 8 GB card (Q5) 16 GB Mac (Q8)
52 Gemma 3 4B
4.3B · gemma3:4b
14 325 6 8 GB card (Q8) 16 GB Mac (Q8)
53 Qwen2.5 VL 72B Instruct
73B
14 1314 8 more than 32 GB 64 GB Mac (Q4)
54 Reka Edge
7.1B
8 016 3 8 GB card (Q6) 16 GB Mac (Q8)
55 Llama 3.2 3B Instruct
3.2B · llama3.2:3b
6 013 2 8 GB card (Q8) 16 GB Mac (Q8)
56 Llama 3.2 1B Instruct (Q4, Mac)
1.2B · llama3.2:1b-instruct-q4_K_M
5 09 3 8 GB card (Q4) 16 GB Mac (Q4)

Scores are out of 100. "Runs on" is the smallest graphics card or Mac that can hold the model, and which version to download for it. How we worked this out

How we test

Every model gets the same 110 questions. 30 are small coding jobs, and we run the code to check that it works. The rest ask the model to read things like bills, pay stubs and email threads and pull out the right numbers. Many need a little math, or catching a correction further down the page.

We keep most questions secret, so no model can memorize the answers. The rest are public, along with every answer each model gave. Full technical details.

Questions people ask

What's the best local LLM for an 8 GB graphics card?

Right now it's Granite 4.2 8B. Download the Q4 version. It scored 77 out of 100 in our tests.

What's the best local LLM for a 12 GB graphics card?

Right now it's Granite 4.2 8B. Download the Q8 version. It scored 77 out of 100 in our tests.

What's the best local LLM for a 16 GB graphics card?

Right now it's gpt-oss-20b. Download the standard version. It scored 88 out of 100 in our tests.

What's the best local LLM for a 24 GB graphics card?

Right now it's Qwen3.6 27B. Download the Q5 version. It scored 97 out of 100 in our tests.

What's the best local LLM for a Mac with 16 GB of memory?

Right now it's Granite 4.2 8B. Download the Q6 version. It scored 77 out of 100 in our tests.

What's the best local LLM for a Mac with 32 GB of memory?

Right now it's Qwen3.6 27B. Download the Q4 version. It scored 97 out of 100 in our tests.

What do Q4 and Q8 mean?

They're versions of the same model, squeezed down to different sizes. Q4 is smaller and runs on more computers. Q8 is bigger and a little more accurate. If both fit on your computer, pick the bigger one.

Why don't you say how fast each model runs?

Speed depends a lot on your exact computer and the app you use, so any number we gave would be wrong for most people. We tell you which models fit on your machine and how good their answers are.

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.