- Qwen2.5 7B Instruct is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's not one we'd recommend right now: 24 out of 100, #45 of 56.
- It solved 3 of 30 coding jobs and scored 37 on reading documents. On our hardest tasks it scored 11.
- Runs on an 8 GB graphics card or a Mac with 16 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen2.5 7B Instruct got 3 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen2.5 7B Instruct scored 37; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 10 | 0 | 13 |
| Reading documents | 37 | 57 | 32 |
| Decisions | 71 | 75 | 70 |
On the 18 hardest tasks (included in the scores above) it scored 11. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
The Q4 download
Most people don't run the full-size model; they download a smaller Q4 version. We ran that version with Ollama on our Mac mini (M4, 16 GB) and gave it the same questions. It scored 24, about the same as the full-size model (24). You lose almost nothing by downloading it.
| Test | Full size (online) | Q4 (our Mac) |
|---|---|---|
| Coding | 10 | 13 |
| Reading documents | 37 | 34 |
| Hardest tasks | 11 | 9 |
| Overall | 24 | 24 |
Can your computer run it?
You can download Qwen2.5 7B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | Yes | Q6 |
| Graphics card, 12 GB | Yes | Q8 |
| Graphics card, 16 GB | Yes | Q8 |
| Graphics card, 24 GB | Yes | Q8 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | Yes | Q8 |
| Mac, 24 GB | Yes | Q8 |
| Mac, 32 GB | Yes | Q8 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run qwen2.5:7b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Qwen2.5 7B Instruct wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
from collections import defaultdict
def parse_ini(text: str) -> dict[str, dict[str, str]]:
lines = text.splitlines()
sections = defaultdict(dict)
defaults = {}
current_section = 'DEFAULT'
for line in lines:
line = line.strip()
if not line or line.startswith(';') or line.startswith('#'):
continue
if line.startswith('[') and line.endswith(']'):
current_section = line[1:-1].strip()
if current_section == 'DEFAULT':
defaults = sections[current_section]
continue
sections[current_section] = {}
else:
key, value = line.split(maxsplit=1)
key = key.strip().lower()
value = value.strip()
if value.startswith('"') and value.endswith('"'):
value = value[1:-1]
while line.endswith('\\'):
next_line = next((l for l in lines[lines.index(line) + 1:] if l.strip()), '').strip()
value += '\n' + next_line
# Resolve interpolation
def resolve(key):
nonlocal sections, defaults
if key in defaults:
return defaults[key]
elif key in sections[current_section]:
return sections[current_section][key]
else:
raise KeyError(f"Key {key} not found")
value = re.sub(r'\$\{(\w+)\}', lambda m: resolve(m.group(1)), value)
value = re.sub(r'\$\{(\w+):(\w+)\}', lambda m: resolve(f"{m.group(1)}.{m.group(2)}"), value)
sections[current_section][key] = value
return {k: v for k, v in sections.items() if k != 'DEFAULT'}
``` Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from typing import Any, Callable
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock: Callable[[], float] = None):
self.capacity = capacity
self.ttl = ttl
self.clock = clock or (lambda: 0)
self.cache = OrderedDict()
self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
def put(self, key: Any, value: Any) -> None:
current_time = self.clock()
if key in self.cache:
self.cache.move_to_end(key)
self.cache[key] = value
self.cache[key]["expiry"] = current_time + self.ttl
else:
if len(self.cache) >= self.capacity:
while True:
oldest_key = next(iter(self.cache))
if self.cache[oldest_key]["expiry"] > current_time:
break
del self.cache[oldest_key]
self.stats["expirations"] += 1
self.stats["misses"] += 1
self.cache[key] = {"value": value, "expiry": current_time + self.ttl}
def get(self, key: Any, default: Any = None) -> Any:
current_time = self.clock()
if key not in self.cache:
self.stats["misses"] += 1
return default
entry = self.cache[key]
if entry["expiry"] <= current_time:
del self.cache[key]
self.stats["expirations"] += 1
self.stats["misses"] += 1
return default
self.cache.move_to_end(key)
entry["expiry"] = current_time + self.ttl
self.stats["hits"] += 1
return entry["value"]
def __len__(self) -> int:
current_time = self.clock()
return sum(1 for k in self.cache if self.cache[k]["expiry"] > current_time)
def stats(self) -> dict:
return self.stats
``` Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
import re
if not s.strip():
raise ValueError("Duration string cannot be empty or blank")
pattern = r"(\d+)(h|m|s)"
matches = re.findall(pattern, s.replace(" ", "").lower())
if not matches:
raise ValueError("Duration string must contain at least one valid unit with a number")
units = {"h": 3600, "m": 60, "s": 1}
total_seconds = 0
for amount, unit in matches:
if unit not in units:
raise ValueError(f"Unknown unit: {unit}")
total_seconds += int(amount) * units[unit]
return total_seconds
``` Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
nums.sort()
result = []
start = None
for i in range(len(nums)):
if start is None or nums[i] == nums[i - 1] + 1:
start = nums[i]
else:
if i > 0 and nums[i - 1] - nums[i - 2] >= 2:
result.append(f"{start}..{nums[i - 1]}")
start = nums[i]
else:
result.extend(map(str, range(nums[i - 1], nums[i])))
if start is not None:
if len(nums) - nums.index(start) >= 2:
result.append(f"{start}..{nums[-1]}")
else:
result.extend(map(str, range(start, nums[-1] + 1)))
return ",".join(result)
``` Coding: Top customers Didn't work
```python
from typing import List, Tuple
import re
def top_customers(rows: List[dict], n: int) -> List[Tuple[str, float]]:
def parse_amount(amount_str: str) -> float:
amount_str = amount_str.strip()
if not amount_str:
return 0.0
match = re.match(r'^-?\$?(\d+(,\d{3})*(\.\d+)?|\.\d+)$', amount_str.replace(',', ''))
if not match:
return 0.0
return round(float(match.group(1).replace(',', '')), 2)
customers = {}
for row in rows:
name = row.get('customer').strip().lower()
if not name:
continue
amount = parse_amount(row.get('amount'))
if amount == 0.0:
continue
if name not in customers:
customers[name] = amount
else:
customers[name] += amount
sorted_customers = sorted(customers.items(), key=lambda x: (-x[1], x[0]))
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import re
from unicodedata import normalize
def slugify(title: str, max_len: int = 50) -> str:
# Transliterate accented characters to ASCII and lowercase
ascii_title = normalize('NFKD', title).encode('ASCII', 'ignore').decode().lower()
# Replace every run of non-alphanumeric characters with a single "-"
slug = re.sub(r'[^a-z0-9]+', '-', ascii_title)
# Strip leading/trailing "-"
slug = slug.strip('-')
# If the slug is longer than max_len, shorten it without cutting a word
if len(slug) > max_len:
# Find the last space before max_len that allows us to keep whole words
last_space = max((i for i in range(max_len) if slug[i] == '-'), default=max_len)
slug = slug[:last_space]
# Ensure the result does not end with "-"
if slug.endswith('-'):
slug = slug[:-1]
return slug
``` Coding: Token bucket Didn't work
```python
from collections import deque
import time
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock=time.time):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.tokens = capacity
self.clock = clock
self.last_refill_time = clock()
self.refill_queue = deque()
def _refill(self):
current_time = self.clock()
if not self.refill_queue or current_time >= self.refill_queue[-1]:
elapsed_time = max(current_time - self.last_refill_time, 0)
tokens_to_add = min(elapsed_time * self.refill_per_sec, self.capacity - self.tokens)
self.tokens = min(self.capacity, self.tokens + tokens_to_add)
self.last_refill_time = current_time
if tokens_to_add > 0:
self.refill_queue.append(current_time + 1 / self.refill_per_sec)
def allow(self, cost: float = 1) -> bool:
self._refill()
if self.tokens >= cost:
self.tokens -= cost
return True
return False
@property
def tokens(self) -> float:
self._refill()
return self._tokens
@tokens.setter
def tokens(self, value: float):
self._tokens = min(value, self.capacity)
``` Decisions: Refund window (hard) 50% right
{
"outcome": {
"partial_refund": 1.0
},
"defective": {
"false": 1.0
}
} Decisions: Incident matrix (hard) 0% right
{
"severity": {
"0": 0.1,
"1": 0.2,
"2": 0.3,
"3": 0.4
},
"page": {
"true": 0.9,
"false": 0.1
}
} Decisions: Tool followup (hard) 100% right
{
"tool": {
"calendar": 1.0
},
"confirm": {
"true": 1.0
}
} Decisions: Legit security alert (hard) 100% right
{
"phishing": {
"true": 0.0,
"false": 1.0
},
"action_needed": {
"true": 0.2,
"false": 0.8
}
} Decisions: Meeting slot (hard) 0% right
{
"slot": {
"A": 0.0,
"B": 0.25,
"C": 0.5,
"D": 0.25
},
"raj_last": {
"true": 0.0,
"false": 1.0
}
} Decisions: Review mixed (hard) 100% right
{
"hardware": {
"true": 1,
"false": 0
},
"support": {
"true": 1,
"false": 0
}
} Decisions: Support checkout down 100% right
{
"department": {
"billing": 0,
"technical": 1,
"account": 0,
"shipping": 0,
"sales": 0
},
"urgency": {
"0": 0,
"1": 0,
"2": 0,
"3": 1
},
"outage": {
"true": 1,
"false": 0
}
} Decisions: Refund wrong plan 100% right
{
"department": {
"billing": 1.0
},
"refund": {
"true": 1.0
},
"tone": {
"calm": 1.0
}
} Decisions: Moderation doxxing 50% right
{
"policy": {
"none": 0.9,
"harassment": 0.1
},
"personal_info": {
"true": 1
}
} Decisions: Route calendar 100% right
{
"tool": {
"calendar": 1.0
},
"confirm": {
"true": 1.0
}
} Decisions: Doc invoice missing due 50% right
{
"doc_type": {
"invoice": 1.0
},
"missing_due_date": {
"true": 0.0,
"false": 1.0
}
} Decisions: Phishing paypal 100% right
{
"phishing": {
"true": 1,
"false": 0
},
"risk": {
"0": 0,
"1": 0,
"2": 0,
"3": 1
}
} Decisions: Pii ssn email 100% right
{
"data_kind": {
"none": 0,
"contact": 0,
"financial": 0.1,
"government_id": 1,
"health": 0
},
"sensitive": {
"true": 1,
"false": 0
}
} Decisions: Review mixed 100% right
{
"sentiment": {
"negative": 1
},
"defect": {
"true": 1
},
"recommend": {
"false": 1
}
} Documents: Saas escalator (hard) 30% right
{
"year2_price_per_seat_month": 47.68,
"year3_price_per_seat_month": 49.91,
"year1_invoice": 54000.00,
"year2_invoice": null,
"addon_months_billed": 6,
"addon_invoice": 23040.00,
"year3_invoice": null,
"year3_discount_percent": 15.00,
"total_contract_value": 133040.00,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 75% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{"date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60},
{"date": "2025-02-24", "category": "ground_transport", "amount_usd": 39.72},
{"date": "2025-02-25", "category": "meals", "amount_usd": 230.44},
{"date": "2025-02-26", "category": "lodging", "amount_usd": 469.08},
{"date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.84}
],
"rejected_item_count": 2,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 1883.84,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 67% right
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150,
"monthly_rent_from_2025_06_01": 2254,
"late_fee_from_2025_06_01": 112.7,
"security_deposit": 2150,
"total_pet_deposits": 800,
"total_monthly_payment_july_2025": 2289,
"move_in_payment": 4500
} Documents: Ticket SLA 82% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-13T17:00",
"sla_due_utc": "2025-09-13T22:00:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 33% right
{
"q3_total_usd": 1951000,
"q2_total_usd": 1622000,
"q2_central_originally_reported_usd": 3047000,
"q2_to_q3_change_pct": -15.6,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "West",
"regions_declining_q2_to_q3": ["Central"],
"international_q3_organic_usd": 1731000,
"west_excluding_mountain_q3_usd": 418000
} Size: 7.6B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.