- Qwen3 14B is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's decent, with trade-offs: 72 out of 100, #20 of 56.
- It solved 19 of 30 coding jobs and scored 81 on reading documents. On our hardest tasks it scored 42.
- Runs on a 12 GB graphics card or a Mac with 24 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen3 14B got 19 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen3 14B scored 81; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 63 | 57 | 65 |
| Reading documents | 81 | 82 | 81 |
| Decisions | 95 | 100 | 93 |
On the 18 hardest tasks (included in the scores above) it scored 42. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Qwen3 14B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | Yes | Q4 |
| Graphics card, 16 GB | Yes | Q6 |
| Graphics card, 24 GB | Yes | Q8 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | No | – |
| Mac, 24 GB | Yes | Q6 |
| Mac, 32 GB | Yes | Q8 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run qwen3:14b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Qwen3 14B wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.stats = {
'hits': 0,
'misses': 0,
'evictions': 0,
'expirations': 0
}
def put(self, key, value):
if self.capacity == 0:
return
current_time = self.clock()
if key in self.cache:
self.cache[key] = (value, current_time + self.ttl)
self.cache.move_to_end(key)
return
self._remove_expired()
if len(self.cache) >= self.capacity:
self._evict_lru()
self.cache[key] = (value, current_time + self.ttl)
def get(self, key, default=None):
if self.capacity == 0:
self.stats['misses'] += 1
return default
current_time = self.clock()
if key not in self.cache:
self.stats['misses'] += 1
return default
value, expiry = self.cache[key]
if current_time >= expiry:
del self.cache[key]
self.stats['expirations'] += 1
return default
self.cache.move_to_end(key)
self.stats['hits'] += 1
return value
def __len__(self):
if self.capacity == 0:
return 0
current_time = self.clock()
count = 0
for key in self.cache:
_, expiry = self.cache[key]
if current_time < expiry:
count += 1
return count
def stats(self):
return self.stats.copy()
def _remove_expired(self):
current_time = self.clock()
expired_keys = []
for key in list(self.cache.keys()):
_, expiry = self.cache[key]
if current_time >= expiry:
expired_keys.append(key)
for key in expired_keys:
del self.cache[key]
self.stats['expirations'] += 1
def _evict_lru(self):
if self.cache:
key_to_evict, _ = self.cache.popitem(last=False)
self.stats['evictions'] += 1
``` Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
s = s.strip()
if not s:
raise ValueError("Empty string")
tokens = s.split()
units = []
total_seconds = 0
for token in tokens:
if len(token) < 2:
raise ValueError("Invalid token")
unit_char = token[-1].lower()
if unit_char not in {'h', 'm', 's'}:
raise ValueError("Unknown unit")
number_part = token[:-1]
try:
num = int(number_part)
except ValueError:
raise ValueError("Invalid number part")
if num < 0:
raise ValueError("Negative number")
units.append(unit_char)
if unit_char == 'h':
total_seconds += num * 3600
elif unit_char == 'm':
total_seconds += num * 60
elif unit_char == 's':
total_seconds += num
ordered_units = ['h', 'm', 's']
for i in range(len(units)):
if i >= len(ordered_units):
raise ValueError("Units out of order")
if units[i] != ordered_units[i]:
raise ValueError("Units out of order")
return total_seconds
``` Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
sorted_nums = sorted(set(nums))
ranges = []
current_start = sorted_nums[0]
current_end = sorted_nums[0]
for num in sorted_nums[1:]:
if num == current_end + 1:
current_end = num
else:
ranges.append((current_start, current_end))
current_start = num
current_end = num
ranges.append((current_start, current_end))
parts = []
for start, end in ranges:
if start == end:
parts.append(str(start))
elif end - start == 1:
parts.append(f"{start},{end}")
else:
parts.append(f"{start}..{end}")
return ",".join(parts)
``` Coding: Top customers Worked
```python
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
first_valid_name = {}
customer_totals = {}
for row in rows:
customer = row.get('customer', '').strip()
if not customer:
continue
amount_str = row.get('amount', '').replace('$', '').replace(',', '')
try:
amount = float(amount_str)
except ValueError:
continue
lower_name = customer.lower()
if lower_name not in first_valid_name:
first_valid_name[lower_name] = customer
if lower_name not in customer_totals:
customer_totals[lower_name] = 0.0
customer_totals[lower_name] += amount
result = []
for lower_name in customer_totals:
name = first_valid_name[lower_name]
total = round(customer_totals[lower_name], 2)
result.append((name, total))
result.sort(key=lambda x: (-x[1], x[0]))
return result[:n]
``` Coding: Slugify Worked
```python
import unicodedata
import re
def slugify(title: str, max_len: int = 50) -> str:
# Step 1: Transliterate accented characters to ASCII and drop non-ASCII
title = unicodedata.normalize('NFKD', title)
title = ''.join([c for c in title if ord(c) < 128])
# Step 2: Lowercase
title = title.lower()
# Step 3: Replace non-alphanumeric with '-', then strip leading/trailing '-'
title = re.sub(r'[^a-z0-9]+', '-', title)
title = title.strip('-')
# Step 4: Shorten if necessary
if len(title) <= max_len:
return title
parts = title.split('-')
if not parts:
return ''
first_part = parts[0]
if len(first_part) > max_len:
return first_part[:max_len]
selected_parts = [first_part]
current_length = len(first_part)
for part in parts[1:]:
if current_length + 1 + len(part) <= max_len:
current_length += 1 + len(part)
selected_parts.append(part)
else:
break
return '-'.join(selected_parts)
``` Coding: Token bucket Worked
```python
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.clock = clock
self.last_refill_time = clock()
self.current_tokens = capacity
def allow(self, cost: float = 1.0) -> bool:
now = self.clock()
delta = now - self.last_refill_time
self.current_tokens = min(self.current_tokens + delta * self.refill_per_sec, self.capacity)
self.last_refill_time = now
if self.current_tokens >= cost:
self.current_tokens -= cost
return True
return False
@property
def tokens(self) -> float:
now = self.clock()
delta = now - self.last_refill_time
self.current_tokens = min(self.current_tokens + delta * self.refill_per_sec, self.capacity)
self.last_refill_time = now
return self.current_tokens
``` Decisions: Refund window (hard) 100% right
{
"outcome": {
"full_refund": 0.0,
"partial_refund": 0.0,
"store_credit": 1.0,
"nothing": 0.0
},
"defective": {
"true": 0.0,
"false": 1.0
}
} Decisions: Incident matrix (hard) 100% right
{
"severity": {
"0": 0.0,
"1": 0.95,
"2": 0.05,
"3": 0.0
},
"page": {
"true": 0.0,
"false": 1.0
}
} Decisions: Tool followup (hard) 100% right
{
"tool": {
"calendar": 1.0
},
"confirm": {
"true": 1.0
}
} Decisions: Legit security alert (hard) 100% right
{
"phishing": {
"true": 0.1,
"false": 0.9
},
"action_needed": {
"true": 0.2,
"false": 0.8
}
} Decisions: Meeting slot (hard) 100% right
{
"slot": {
"A": 0,
"B": 1,
"C": 0,
"D": 0
},
"raj_last": {
"true": 1,
"false": 0
}
} Decisions: Review mixed (hard) 100% right
{"hardware": {"true": 1.0, "false": 0.0}, "support": {"true": 1.0, "false": 0.0}} Decisions: Support checkout down 100% right
{
"department": {
"technical": 0.95,
"billing": 0.05,
"account": 0.0,
"shipping": 0.0,
"sales": 0.0
},
"urgency": {
"3": 0.99,
"2": 0.01,
"1": 0.0,
"0": 0.0
},
"outage": {
"true": 0.99,
"false": 0.01
}
} Decisions: Refund wrong plan 100% right
{
"department": {
"billing": 1.0,
"technical": 0.0,
"account": 0.0,
"shipping": 0.0,
"sales": 0.0
},
"refund": {
"true": 1.0,
"false": 0.0
},
"tone": {
"frustrated": 0.0,
"calm": 1.0
}
} Decisions: Moderation doxxing 100% right
{
"policy": {
"none": 0.2,
"harassment": 0.8,
"hate": 0.0,
"spam": 0.0,
"self_harm": 0.0
},
"personal_info": {
"true": 1.0,
"false": 0.0
}
} Decisions: Route calendar 100% right
{
"tool": {
"calendar": 0.95,
"email": 0.05,
"web_search": 0.0,
"calculator": 0.0,
"none": 0.0
},
"confirm": {
"true": 0.85,
"false": 0.15
}
} Decisions: Doc invoice missing due 100% right
{
"doc_type": {
"invoice": 0.95,
"resume": 0.01,
"contract": 0.01,
"bank_statement": 0.01,
"other": 0.02
},
"missing_due_date": {
"true": 0.9,
"false": 0.1
}
} Decisions: Phishing paypal 100% right
{
"phishing": {
"true": 0.95,
"false": 0.05
},
"risk": {
"0": 0.01,
"1": 0.05,
"2": 0.10,
"3": 0.84
}
} Decisions: Pii ssn email 100% right
{
"data_kind": {
"government_id": 0.9,
"contact": 0.1,
"none": 0,
"financial": 0,
"health": 0
},
"sensitive": {
"true": 1.0,
"false": 0
}
} Decisions: Review mixed 100% right
{
"sentiment": {
"positive": 0.1,
"neutral": 0.1,
"negative": 0.8
},
"defect": {
"true": 1.0,
"false": 0.0
},
"recommend": {
"true": 0.0,
"false": 1.0
}
} Documents: Saas escalator (hard) 30% right
{
"year2_price_per_seat_month": 46.53,
"year3_price_per_seat_month": 48.86,
"year1_invoice": 64800.00,
"year2_invoice": 60302.88,
"addon_months_billed": 6,
"addon_invoice": 37968.48,
"year3_invoice": 139544.16,
"year3_discount_percent": 15.0,
"total_contract_value": 302615.52,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 100% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{"date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60},
{"date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88},
{"date": "2025-02-25", "category": "meals", "amount_usd": 229.39},
{"date": "2025-02-26", "category": "lodging", "amount_usd": 466.56},
{"date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.82}
],
"rejected_item_count": 1,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 2159.25,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 100% right
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150,
"monthly_rent_from_2025_06_01": 2236,
"late_fee_from_2025_06_01": 111.8,
"security_deposit": 2150,
"total_pet_deposits": 800,
"total_monthly_payment_july_2025": 2306,
"move_in_payment": 4700
} Documents: Ticket SLA 92% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address", "invoice_pdf"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-15T15:30",
"sla_due_utc": "2025-09-15T20:30:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 89% right
{
"q3_total_usd": 15346000,
"q2_total_usd": 14464000,
"q2_central_originally_reported_usd": 3047000,
"q2_to_q3_change_pct": 6.1,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "West",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731000,
"west_excluding_mountain_q3_usd": 4201000
} Size: 15B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.