Rankings · updated OCT 11

Decision models vs. regular LLMs: 14 free models beat the best one

Decision models don't write text. You give them a situation and a few multiple-choice questions, and they return a probability for every option. We tested them, and regular LLMs, on the same real-world calls.

The short version
  • Best dedicated decision model: Solar Decide, right on 96% of calls.
  • Best regular model you can run at home: Qwen3.5-27B, 100%.
  • 14 free models you can run yourself beat every dedicated decision model, including the small Qwen3.5-9B.
  • The gap shows on the hard cases: decision models got 68–95% of them right; the best regular models got 100%.
  • "Calibration" shows how honest a model's confidence is: 100 means it is always right and always sure; low numbers mean it is often confidently wrong.

Best decision models

Only the dedicated decision models, ranked. 6 of 14 have open weights you can download; the rest are available only through an API.

#ModelWeightsAccuracyHard casesCalibration
1 Solar Decide
Upstage
Closed (API only) 96 95 96
2 Mercury Decide
Inception
Closed (API only) 94 90 95
3 Jev 1.13
Typesafe
Closed (API only) 92 79 94
4 Drex v1.5
Nace AI · 9B
Open
custom license
92 84 93
5 Decider V1.1 27B
Perplexity
Closed (API only) 91 84 94
6 Microsoft-Decision-1
Microsoft
Closed (API only) 90 87 92
7 GPT-6 Luna Decisions
OpenAI
Closed (API only) 89 79 95
8 Clef
Cloudflare · 27B
Open
Apache 2.0
88 76 91
9 d1
Liquid AI
Closed (API only) 87 74 91
10 Solar Decide Flash
Upstage
Closed (API only) 86 74 90
11 Tev1 4B Experimental
Together AI · 4.7B
Open
license not stated
85 74 88
12 Clef Flash
Cloudflare · 9.4B
Open
Apache 2.0
85 68 88
13 Kev 4B
Jared Palmer
Open
Apache 2.0
84 68 87
14 Clef Omni
Cloudflare · 35B
Open
Apache 2.0
83 68 86
– Span-01
Respan
Closed (API only) Couldn't take the test: it only answers yes/no questions about a chat transcript.
– Span-01 Lite
Respan
Closed (API only) Couldn't take the test: it only answers yes/no questions about a chat transcript.

Decision models vs. regular LLMs

Everyone who took the test, including free models you can run on your own computer.

#ModelTypeAccuracyHard casesCalibration
1 Qwen3.5-27B
qwen/qwen3.5-27b
Local LLM 100 100 100
2 Qwen3.5-35B-A3B
qwen/qwen3.5-35b-a3b
Local LLM 99 100 99
3 Qwen3.6 27B
qwen/qwen3.6-27b
Local LLM 99 100 99
4 Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b
Local LLM 99 100 99
5 Muse Glimmer 30B
meta/muse-glimmer-30b
Local LLM 99 100 99
6 gpt-oss-120b
openai/gpt-oss-120b
Local LLM 99 100 99
7 Ling 3.0 Flash
inclusionai/ling-3.0-flash
Local LLM 99 97 99
8 GLM 4.5V
z-ai/glm-4.5v
Local LLM 98 100 99
9 Qwen3.8 27B
qwen/qwen3.8-27b
Local LLM 98 100 98
10 Laguna XS 2.1
poolside/laguna-xs-2.1
Local LLM 97 100 98
11 Qwen3.5-9B
qwen/qwen3.5-9b
Local LLM 97 100 97
12 Qwen3 32B
qwen/qwen3-32b
Local LLM 96 97 98
13 Ling 3.0 Flash VL
inclusionai/ling-3.0-flash-vl
Local LLM 96 92 97
14 Qwen3 VL 30B A3B Thinking
qwen/qwen3-vl-30b-a3b-thinking
Local LLM 96 97 96
15 Solar Decide
upstage/solar-decide
Decision model 96 95 96
16 Qwen3 30B A3B
qwen/qwen3-30b-a3b
Local LLM 96 97 96
17 Qwen3 14B
qwen/qwen3-14b
Local LLM 95 95 95
18 Nex-N2.5-Mini
nex-agi/nex-n2.5-mini
Local LLM 95 95 95
19 gpt-oss-20b
openai/gpt-oss-20b
Local LLM 95 97 95
20 GLM 4.6V
z-ai/glm-4.6v
Local LLM 94 95 95
21 Mercury Decide
inception/mercury-decide
Decision model 94 90 95
22 Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b
Local LLM 94 100 94
23 Gemma 4 31B
google/gemma-4-31b-it
Local LLM 93 82 93
24 Granite 4.2 8B
ibm-granite/granite-4.2-8b
Local LLM 93 97 93
25 Jev 1.13
typesafe/jev-1.13
Decision model 92 79 94
26 Drex v1.5
nace-ai/drex-v1.5
Decision model 92 84 93
27 Devstral 2 2512
mistralai/devstral-2512
Local LLM 92 79 92
28 Gemma 4 26B A4B
google/gemma-4-26b-a4b-it
Local LLM 92 79 91
29 Decider V1.1 27B
perplexity/pplx-decider-v1.1-27b
Decision model 91 84 94
30 GLM 4.5 Air
z-ai/glm-4.5-air
Local LLM 91 90 92
31 Microsoft-Decision-1
microsoft/microsoft-decision-1
Decision model 90 87 92
32 GPT-6 Luna Decisions
openai/gpt-6-luna-decisions
Decision model 89 79 95
33 Qwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instruct
Local LLM 89 82 89
34 Gemma 3 27B
google/gemma-3-27b-it
Local LLM 88 79 91
35 Qwen2.5 72B Instruct
qwen/qwen-2.5-72b-instruct
Local LLM 88 82 89
36 Clef
cloudflare/clef
Decision model 88 76 91
37 Command A
cohere/command-a
Local LLM 88 68 89
38 d1
liquid/d1
Decision model 87 74 91
39 Solar Decide Flash
upstage/solar-decide-flash
Decision model 86 74 90
40 Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instruct
Local LLM 86 68 86
41 Ministral 3 14B 2512
mistralai/ministral-14b-2512
Local LLM 86 68 86
42 Tev1 4B Experimental
togethercomputer/tev1-4b-experimental
Decision model 85 74 88
43 Clef Flash
cloudflare/clef-flash
Decision model 85 68 88
44 Mistral Small 3
mistralai/mistral-small-24b-instruct-2501
Local LLM 85 68 87
45 GLM 4.7 Flash
z-ai/glm-4.7-flash
Local LLM 85 84 87
46 Kev 4B
jaredpalmer/kev-4b
Decision model 84 68 87
47 Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
Local LLM 84 71 86
48 Qwen3 30B A3B Instruct 2507
qwen/qwen3-30b-a3b-instruct-2507
Local LLM 84 68 86
49 Ministral 3 8B 2512
mistralai/ministral-8b-2512
Local LLM 84 74 85
50 Gemma 3 12B
google/gemma-3-12b-it
Local LLM 83 71 87
51 Clef Omni
cloudflare/clef-omni
Decision model 83 68 86
52 Mistral Small 3.2 24B
mistralai/mistral-small-3.2-24b-instruct
Local LLM 83 61 85
53 Laguna S 2.1
poolside/laguna-s-2.1
Local LLM 83 74 84
54 Qwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instruct
Local LLM 81 63 85
55 Nemotron 3.5 Lightning
nvidia/nemotron-3.5-lightning
Local LLM 80 68 83
56 Gemma 2 27B
google/gemma-2-27b-it
Local LLM 80 71 83
57 Qwen3 VL 8B Instruct
qwen/qwen3-vl-8b-instruct
Local LLM 79 63 81
58 Mistral Nemo
mistralai/mistral-nemo
Local LLM 78 71 83
59 Phi 4
microsoft/phi-4
Local LLM 75 53 78
60 Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b-instruct
Local LLM 75 42 76
61 Hunyuan A13B Instruct
tencent/hunyuan-a13b-instruct
Local LLM 74 61 74
62 Ministral 3 3B 2512
mistralai/ministral-3b-2512
Local LLM 72 53 75
63 Qwen2.5 7B Instruct
qwen/qwen-2.5-7b-instruct
Local LLM 71 63 76
64 Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instruct
Local LLM 68 53 71
65 Gemma 3 4B
google/gemma-3-4b-it
Local LLM 68 61 73
66 Llama 3.3 8B Instruct (Q4, Mac)
Local LLM 59 55 61
67 Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instruct
Local LLM 44 42 54
68 Reka Edge
rekaai/reka-edge
Local LLM 9 8 14

What this means

Decision models promise fast, cheap, well-calibrated answers to multiple-choice questions, and on the easy cases they deliver: most of them get the obvious calls right. Where they fall behind is when the answer depends on reading a rule carefully and doing a small calculation, like counting days in a refund window or converting a time zone. The best regular models handle those much better.

If you're choosing one for real work, test it on your own hardest cases, not your typical ones. And if you can run a model yourself, Qwen3.5-9B is a strong, free baseline to beat.

What we test

56 situations, each with one to three questions: which team should get a support ticket and how urgent it is, whether an email is phishing, which tool an assistant should use, what kind of document something is, whether a message contains sensitive data, how severe an incident is, and whether a code change needs a security review. Some are tricky on purpose: sarcasm, "this is NOT a billing question", criticism that isn't hate speech.

19 of the situations are hard cases: the answer follows from rules we spell out (a refund policy, an expense limit, a severity matrix, a contract clause), but getting it right means doing date or money math, converting time zones, or ignoring a detail that points the wrong way. They are part of the accuracy score and also shown on their own.

Decision models answer through OpenRouter's Decisions API. Regular LLMs get the same situations and options and are asked to give a probability for each option as JSON. Accuracy counts how often the most likely option is right; calibration uses the Brier score. Full technical details.

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.