Review · updated OCT 11

Qwen3 14B review: decent, with trade-offs

It scored 72 out of 100, #20 of 56. It solved 19 of 30 coding jobs and scored 81 on reading documents. Runs on a 12 GB graphics card or a Mac with 24 GB.

The short version
  • Qwen3 14B is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's decent, with trade-offs: 72 out of 100, #20 of 56.
  • It solved 19 of 30 coding jobs and scored 81 on reading documents. On our hardest tasks it scored 42.
  • Runs on a 12 GB graphics card or a Mac with 24 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen3 14B got 19 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen3 14B scored 81; the best model scored 100.

TestScorePublic questionsSecret questions
Coding635765
Reading documents818281
Decisions9510093

On the 18 hardest tasks (included in the scores above) it scored 42. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen3 14B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBYesQ4
Graphics card, 16 GBYesQ6
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBNo–
Mac, 24 GBYesQ6
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run qwen3:14b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen3 14B wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {
            'hits': 0,
            'misses': 0,
            'evictions': 0,
            'expirations': 0
        }

    def put(self, key, value):
        if self.capacity == 0:
            return
        current_time = self.clock()
        if key in self.cache:
            self.cache[key] = (value, current_time + self.ttl)
            self.cache.move_to_end(key)
            return
        self._remove_expired()
        if len(self.cache) >= self.capacity:
            self._evict_lru()
        self.cache[key] = (value, current_time + self.ttl)

    def get(self, key, default=None):
        if self.capacity == 0:
            self.stats['misses'] += 1
            return default
        current_time = self.clock()
        if key not in self.cache:
            self.stats['misses'] += 1
            return default
        value, expiry = self.cache[key]
        if current_time >= expiry:
            del self.cache[key]
            self.stats['expirations'] += 1
            return default
        self.cache.move_to_end(key)
        self.stats['hits'] += 1
        return value

    def __len__(self):
        if self.capacity == 0:
            return 0
        current_time = self.clock()
        count = 0
        for key in self.cache:
            _, expiry = self.cache[key]
            if current_time < expiry:
                count += 1
        return count

    def stats(self):
        return self.stats.copy()

    def _remove_expired(self):
        current_time = self.clock()
        expired_keys = []
        for key in list(self.cache.keys()):
            _, expiry = self.cache[key]
            if current_time >= expiry:
                expired_keys.append(key)
        for key in expired_keys:
            del self.cache[key]
            self.stats['expirations'] += 1

    def _evict_lru(self):
        if self.cache:
            key_to_evict, _ = self.cache.popitem(last=False)
            self.stats['evictions'] += 1
```
Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
    s = s.strip()
    if not s:
        raise ValueError("Empty string")
    tokens = s.split()
    units = []
    total_seconds = 0

    for token in tokens:
        if len(token) < 2:
            raise ValueError("Invalid token")
        unit_char = token[-1].lower()
        if unit_char not in {'h', 'm', 's'}:
            raise ValueError("Unknown unit")
        number_part = token[:-1]
        try:
            num = int(number_part)
        except ValueError:
            raise ValueError("Invalid number part")
        if num < 0:
            raise ValueError("Negative number")
        units.append(unit_char)
        if unit_char == 'h':
            total_seconds += num * 3600
        elif unit_char == 'm':
            total_seconds += num * 60
        elif unit_char == 's':
            total_seconds += num

    ordered_units = ['h', 'm', 's']
    for i in range(len(units)):
        if i >= len(ordered_units):
            raise ValueError("Units out of order")
        if units[i] != ordered_units[i]:
            raise ValueError("Units out of order")

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    sorted_nums = sorted(set(nums))
    ranges = []
    current_start = sorted_nums[0]
    current_end = sorted_nums[0]
    for num in sorted_nums[1:]:
        if num == current_end + 1:
            current_end = num
        else:
            ranges.append((current_start, current_end))
            current_start = num
            current_end = num
    ranges.append((current_start, current_end))
    parts = []
    for start, end in ranges:
        if start == end:
            parts.append(str(start))
        elif end - start == 1:
            parts.append(f"{start},{end}")
        else:
            parts.append(f"{start}..{end}")
    return ",".join(parts)
```
Coding: Top customers Worked
```python
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    first_valid_name = {}
    customer_totals = {}
    for row in rows:
        customer = row.get('customer', '').strip()
        if not customer:
            continue
        amount_str = row.get('amount', '').replace('$', '').replace(',', '')
        try:
            amount = float(amount_str)
        except ValueError:
            continue
        lower_name = customer.lower()
        if lower_name not in first_valid_name:
            first_valid_name[lower_name] = customer
        if lower_name not in customer_totals:
            customer_totals[lower_name] = 0.0
        customer_totals[lower_name] += amount
    result = []
    for lower_name in customer_totals:
        name = first_valid_name[lower_name]
        total = round(customer_totals[lower_name], 2)
        result.append((name, total))
    result.sort(key=lambda x: (-x[1], x[0]))
    return result[:n]
```
Coding: Slugify Worked
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    # Step 1: Transliterate accented characters to ASCII and drop non-ASCII
    title = unicodedata.normalize('NFKD', title)
    title = ''.join([c for c in title if ord(c) < 128])
    # Step 2: Lowercase
    title = title.lower()
    # Step 3: Replace non-alphanumeric with '-', then strip leading/trailing '-'
    title = re.sub(r'[^a-z0-9]+', '-', title)
    title = title.strip('-')
    
    # Step 4: Shorten if necessary
    if len(title) <= max_len:
        return title
    
    parts = title.split('-')
    if not parts:
        return ''
    
    first_part = parts[0]
    if len(first_part) > max_len:
        return first_part[:max_len]
    
    selected_parts = [first_part]
    current_length = len(first_part)
    
    for part in parts[1:]:
        if current_length + 1 + len(part) <= max_len:
            current_length += 1 + len(part)
            selected_parts.append(part)
        else:
            break
    
    return '-'.join(selected_parts)
```
Coding: Token bucket Worked
```python
class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.last_refill_time = clock()
        self.current_tokens = capacity

    def allow(self, cost: float = 1.0) -> bool:
        now = self.clock()
        delta = now - self.last_refill_time
        self.current_tokens = min(self.current_tokens + delta * self.refill_per_sec, self.capacity)
        self.last_refill_time = now
        if self.current_tokens >= cost:
            self.current_tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        now = self.clock()
        delta = now - self.last_refill_time
        self.current_tokens = min(self.current_tokens + delta * self.refill_per_sec, self.capacity)
        self.last_refill_time = now
        return self.current_tokens
```
Decisions: Refund window (hard) 100% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.0,
    "store_credit": 1.0,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0.0,
    "1": 0.95,
    "2": 0.05,
    "3": 0.0
  },
  "page": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.1,
    "false": 0.9
  },
  "action_needed": {
    "true": 0.2,
    "false": 0.8
  }
}
Decisions: Meeting slot (hard) 100% right
{
  "slot": {
    "A": 0,
    "B": 1,
    "C": 0,
    "D": 0
  },
  "raj_last": {
    "true": 1,
    "false": 0
  }
}
Decisions: Review mixed (hard) 100% right
{"hardware": {"true": 1.0, "false": 0.0}, "support": {"true": 1.0, "false": 0.0}}
Decisions: Support checkout down 100% right
{
  "department": {
    "technical": 0.95,
    "billing": 0.05,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "3": 0.99,
    "2": 0.01,
    "1": 0.0,
    "0": 0.0
  },
  "outage": {
    "true": 0.99,
    "false": 0.01
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 1.0,
    "technical": 0.0,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "refund": {
    "true": 1.0,
    "false": 0.0
  },
  "tone": {
    "frustrated": 0.0,
    "calm": 1.0
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.2,
    "harassment": 0.8,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.0
  },
  "personal_info": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "calendar": 0.95,
    "email": 0.05,
    "web_search": 0.0,
    "calculator": 0.0,
    "none": 0.0
  },
  "confirm": {
    "true": 0.85,
    "false": 0.15
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 0.95,
    "resume": 0.01,
    "contract": 0.01,
    "bank_statement": 0.01,
    "other": 0.02
  },
  "missing_due_date": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Phishing paypal 100% right
{
  "phishing": {
    "true": 0.95,
    "false": 0.05
  },
  "risk": {
    "0": 0.01,
    "1": 0.05,
    "2": 0.10,
    "3": 0.84
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "government_id": 0.9,
    "contact": 0.1,
    "none": 0,
    "financial": 0,
    "health": 0
  },
  "sensitive": {
    "true": 1.0,
    "false": 0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.1,
    "neutral": 0.1,
    "negative": 0.8
  },
  "defect": {
    "true": 1.0,
    "false": 0.0
  },
  "recommend": {
    "true": 0.0,
    "false": 1.0
  }
}
Documents: Saas escalator (hard) 30% right
{
  "year2_price_per_seat_month": 46.53,
  "year3_price_per_seat_month": 48.86,
  "year1_invoice": 64800.00,
  "year2_invoice": 60302.88,
  "addon_months_billed": 6,
  "addon_invoice": 37968.48,
  "year3_invoice": 139544.16,
  "year3_discount_percent": 15.0,
  "total_contract_value": 302615.52,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 100% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {"date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60},
    {"date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88},
    {"date": "2025-02-25", "category": "meals", "amount_usd": 229.39},
    {"date": "2025-02-26", "category": "lodging", "amount_usd": 466.56},
    {"date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.82}
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 2159.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2236,
  "late_fee_from_2025_06_01": 111.8,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2306,
  "move_in_payment": 4700
}
Documents: Ticket SLA 92% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address", "invoice_pdf"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T15:30",
  "sla_due_utc": "2025-09-15T20:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 89% right
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "West",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 15B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.