- Best dedicated decision model: Solar Decide, right on 96% of calls.
- Best regular model you can run at home: Qwen3.5-27B, 100%.
- 14 free models you can run yourself beat every dedicated decision model, including the small Qwen3.5-9B.
- The gap shows on the hard cases: decision models got 68–95% of them right; the best regular models got 100%.
- "Calibration" shows how honest a model's confidence is: 100 means it is always right and always sure; low numbers mean it is often confidently wrong.
Best decision models
Only the dedicated decision models, ranked. 6 of 14 have open weights you can download; the rest are available only through an API.
| # | Model | Weights | Accuracy | Hard cases | Calibration |
|---|---|---|---|---|---|
| 1 | Solar Decide Upstage | Closed (API only) | 96 | 95 | 96 |
| 2 | Mercury Decide Inception | Closed (API only) | 94 | 90 | 95 |
| 3 | Jev 1.13 Typesafe | Closed (API only) | 92 | 79 | 94 |
| 4 | Drex v1.5 Nace AI · 9B | Open custom license | 92 | 84 | 93 |
| 5 | Decider V1.1 27B Perplexity | Closed (API only) | 91 | 84 | 94 |
| 6 | Microsoft-Decision-1 Microsoft | Closed (API only) | 90 | 87 | 92 |
| 7 | GPT-6 Luna Decisions OpenAI | Closed (API only) | 89 | 79 | 95 |
| 8 | Clef Cloudflare · 27B | Open Apache 2.0 | 88 | 76 | 91 |
| 9 | d1 Liquid AI | Closed (API only) | 87 | 74 | 91 |
| 10 | Solar Decide Flash Upstage | Closed (API only) | 86 | 74 | 90 |
| 11 | Tev1 4B Experimental Together AI · 4.7B | Open license not stated | 85 | 74 | 88 |
| 12 | Clef Flash Cloudflare · 9.4B | Open Apache 2.0 | 85 | 68 | 88 |
| 13 | Kev 4B Jared Palmer | Open Apache 2.0 | 84 | 68 | 87 |
| 14 | Clef Omni Cloudflare · 35B | Open Apache 2.0 | 83 | 68 | 86 |
| – | Span-01 Respan | Closed (API only) | Couldn't take the test: it only answers yes/no questions about a chat transcript. | ||
| – | Span-01 Lite Respan | Closed (API only) | Couldn't take the test: it only answers yes/no questions about a chat transcript. | ||
Decision models vs. regular LLMs
Everyone who took the test, including free models you can run on your own computer.
| # | Model | Type | Accuracy | Hard cases | Calibration |
|---|---|---|---|---|---|
| 1 | Qwen3.5-27Bqwen/qwen3.5-27b | Local LLM | 100 | 100 | 100 |
| 2 | Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Local LLM | 99 | 100 | 99 |
| 3 | Qwen3.6 27Bqwen/qwen3.6-27b | Local LLM | 99 | 100 | 99 |
| 4 | Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | Local LLM | 99 | 100 | 99 |
| 5 | Muse Glimmer 30Bmeta/muse-glimmer-30b | Local LLM | 99 | 100 | 99 |
| 6 | gpt-oss-120bopenai/gpt-oss-120b | Local LLM | 99 | 100 | 99 |
| 7 | Ling 3.0 Flashinclusionai/ling-3.0-flash | Local LLM | 99 | 97 | 99 |
| 8 | GLM 4.5Vz-ai/glm-4.5v | Local LLM | 98 | 100 | 99 |
| 9 | Qwen3.8 27Bqwen/qwen3.8-27b | Local LLM | 98 | 100 | 98 |
| 10 | Laguna XS 2.1poolside/laguna-xs-2.1 | Local LLM | 97 | 100 | 98 |
| 11 | Qwen3.5-9Bqwen/qwen3.5-9b | Local LLM | 97 | 100 | 97 |
| 12 | Qwen3 32Bqwen/qwen3-32b | Local LLM | 96 | 97 | 98 |
| 13 | Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | Local LLM | 96 | 92 | 97 |
| 14 | Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinking | Local LLM | 96 | 97 | 96 |
| 15 | Solar Decideupstage/solar-decide | Decision model | 96 | 95 | 96 |
| 16 | Qwen3 30B A3Bqwen/qwen3-30b-a3b | Local LLM | 96 | 97 | 96 |
| 17 | Qwen3 14Bqwen/qwen3-14b | Local LLM | 95 | 95 | 95 |
| 18 | Nex-N2.5-Mininex-agi/nex-n2.5-mini | Local LLM | 95 | 95 | 95 |
| 19 | gpt-oss-20bopenai/gpt-oss-20b | Local LLM | 95 | 97 | 95 |
| 20 | GLM 4.6Vz-ai/glm-4.6v | Local LLM | 94 | 95 | 95 |
| 21 | Mercury Decideinception/mercury-decide | Decision model | 94 | 90 | 95 |
| 22 | Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | Local LLM | 94 | 100 | 94 |
| 23 | Gemma 4 31Bgoogle/gemma-4-31b-it | Local LLM | 93 | 82 | 93 |
| 24 | Granite 4.2 8Bibm-granite/granite-4.2-8b | Local LLM | 93 | 97 | 93 |
| 25 | Jev 1.13typesafe/jev-1.13 | Decision model | 92 | 79 | 94 |
| 26 | Drex v1.5nace-ai/drex-v1.5 | Decision model | 92 | 84 | 93 |
| 27 | Devstral 2 2512mistralai/devstral-2512 | Local LLM | 92 | 79 | 92 |
| 28 | Gemma 4 26B A4B google/gemma-4-26b-a4b-it | Local LLM | 92 | 79 | 91 |
| 29 | Decider V1.1 27Bperplexity/pplx-decider-v1.1-27b | Decision model | 91 | 84 | 94 |
| 30 | GLM 4.5 Airz-ai/glm-4.5-air | Local LLM | 91 | 90 | 92 |
| 31 | Microsoft-Decision-1microsoft/microsoft-decision-1 | Decision model | 90 | 87 | 92 |
| 32 | GPT-6 Luna Decisionsopenai/gpt-6-luna-decisions | Decision model | 89 | 79 | 95 |
| 33 | Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instruct | Local LLM | 89 | 82 | 89 |
| 34 | Gemma 3 27Bgoogle/gemma-3-27b-it | Local LLM | 88 | 79 | 91 |
| 35 | Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | Local LLM | 88 | 82 | 89 |
| 36 | Clefcloudflare/clef | Decision model | 88 | 76 | 91 |
| 37 | Command Acohere/command-a | Local LLM | 88 | 68 | 89 |
| 38 | d1liquid/d1 | Decision model | 87 | 74 | 91 |
| 39 | Solar Decide Flashupstage/solar-decide-flash | Decision model | 86 | 74 | 90 |
| 40 | Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Local LLM | 86 | 68 | 86 |
| 41 | Ministral 3 14B 2512mistralai/ministral-14b-2512 | Local LLM | 86 | 68 | 86 |
| 42 | Tev1 4B Experimentaltogethercomputer/tev1-4b-experimental | Decision model | 85 | 74 | 88 |
| 43 | Clef Flashcloudflare/clef-flash | Decision model | 85 | 68 | 88 |
| 44 | Mistral Small 3mistralai/mistral-small-24b-instruct-2501 | Local LLM | 85 | 68 | 87 |
| 45 | GLM 4.7 Flashz-ai/glm-4.7-flash | Local LLM | 85 | 84 | 87 |
| 46 | Kev 4Bjaredpalmer/kev-4b | Decision model | 84 | 68 | 87 |
| 47 | Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Local LLM | 84 | 71 | 86 |
| 48 | Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | Local LLM | 84 | 68 | 86 |
| 49 | Ministral 3 8B 2512mistralai/ministral-8b-2512 | Local LLM | 84 | 74 | 85 |
| 50 | Gemma 3 12Bgoogle/gemma-3-12b-it | Local LLM | 83 | 71 | 87 |
| 51 | Clef Omnicloudflare/clef-omni | Decision model | 83 | 68 | 86 |
| 52 | Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | Local LLM | 83 | 61 | 85 |
| 53 | Laguna S 2.1poolside/laguna-s-2.1 | Local LLM | 83 | 74 | 84 |
| 54 | Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | Local LLM | 81 | 63 | 85 |
| 55 | Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning | Local LLM | 80 | 68 | 83 |
| 56 | Gemma 2 27Bgoogle/gemma-2-27b-it | Local LLM | 80 | 71 | 83 |
| 57 | Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instruct | Local LLM | 79 | 63 | 81 |
| 58 | Mistral Nemomistralai/mistral-nemo | Local LLM | 78 | 71 | 83 |
| 59 | Phi 4microsoft/phi-4 | Local LLM | 75 | 53 | 78 |
| 60 | Qwen3 VL 30B A3B Instructqwen/qwen3-vl-30b-a3b-instruct | Local LLM | 75 | 42 | 76 |
| 61 | Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | Local LLM | 74 | 61 | 74 |
| 62 | Ministral 3 3B 2512mistralai/ministral-3b-2512 | Local LLM | 72 | 53 | 75 |
| 63 | Qwen2.5 7B Instructqwen/qwen-2.5-7b-instruct | Local LLM | 71 | 63 | 76 |
| 64 | Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | Local LLM | 68 | 53 | 71 |
| 65 | Gemma 3 4Bgoogle/gemma-3-4b-it | Local LLM | 68 | 61 | 73 |
| 66 | Llama 3.3 8B Instruct (Q4, Mac) | Local LLM | 59 | 55 | 61 |
| 67 | Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instruct | Local LLM | 44 | 42 | 54 |
| 68 | Reka Edgerekaai/reka-edge | Local LLM | 9 | 8 | 14 |
What this means
Decision models promise fast, cheap, well-calibrated answers to multiple-choice questions, and on the easy cases they deliver: most of them get the obvious calls right. Where they fall behind is when the answer depends on reading a rule carefully and doing a small calculation, like counting days in a refund window or converting a time zone. The best regular models handle those much better.
If you're choosing one for real work, test it on your own hardest cases, not your typical ones. And if you can run a model yourself, Qwen3.5-9B is a strong, free baseline to beat.
What we test
56 situations, each with one to three questions: which team should get a support ticket and how urgent it is, whether an email is phishing, which tool an assistant should use, what kind of document something is, whether a message contains sensitive data, how severe an incident is, and whether a code change needs a security review. Some are tricky on purpose: sarcasm, "this is NOT a billing question", criticism that isn't hate speech.
19 of the situations are hard cases: the answer follows from rules we spell out (a refund policy, an expense limit, a severity matrix, a contract clause), but getting it right means doing date or money math, converting time zones, or ignoring a detail that points the wrong way. They are part of the accuracy score and also shown on their own.
Decision models answer through OpenRouter's Decisions API. Regular LLMs get the same situations and options and are asked to give a probability for each option as JSON. Accuracy counts how often the most likely option is right; calibration uses the Brier score. Full technical details.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.