- Qwen2.5 72B Instruct is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's not one we'd recommend right now: 46 out of 100, #31 of 56.
- It solved 12 of 30 coding jobs and scored 52 on reading documents. On our hardest tasks it scored 23.
- Runs on a Mac with 64 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen2.5 72B Instruct got 12 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen2.5 72B Instruct scored 52; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 40 | 43 | 39 |
| Reading documents | 52 | 64 | 49 |
| Decisions | 88 | 89 | 88 |
On the 18 hardest tasks (included in the scores above) it scored 23. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Qwen2.5 72B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | No | – |
| Graphics card, 32 GB | No | – |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | No | – |
| Mac, 48 GB | No | – |
| Mac, 64 GB | Yes | Q4 |
| Mac, 96 GB | Yes | Q6 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
Search for "Qwen2.5 72B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.
Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Qwen2.5 72B Instruct wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, List, Tuple
def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
def interpolate(value: str, section: str, seen: set) -> str:
while True:
match = re.search(r'\$\{([^}]+)\}', value)
if not match:
break
ref = match.group(1)
if ':' in ref:
ref_section, ref_key = ref.split(':', 1)
else:
ref_section = section
ref_key = ref.lower()
if (ref_section, ref_key) in seen:
raise ValueError(f"Reference cycle detected: {ref}")
seen.add((ref_section, ref_key))
if ref_section == 'DEFAULT':
if ref_key not in default_section:
raise KeyError(f"Missing key in DEFAULT: {ref_key}")
ref_value = default_section[ref_key]
elif ref_section in sections:
if ref_key not in sections[ref_section]:
if ref_key in default_section:
ref_value = default_section[ref_key]
else:
raise KeyError(f"Missing key in section {ref_section}: {ref_key}")
else:
ref_value = sections[ref_section][ref_key]
else:
raise KeyError(f"Missing section: {ref_section}")
value = value.replace(match.group(0), ref_value)
seen.discard((ref_section, ref_key))
return value
sections: Dict[str, Dict[str, str]] = {}
default_section: Dict[str, str] = {}
current_section: str = 'DEFAULT'
lines = text.splitlines()
buffer: List[str] = []
for line in lines:
line = line.strip()
if not line or line.startswith((';', '#')):
continue
if line.startswith(('[', ']')):
match = re.match(r'\[(\w+)\]', line)
if match:
if buffer:
sections[current_section][buffer[0].strip().lower()] = '\n'.join(buffer[1:])
buffer.clear()
current_section = match.group(1)
if current_section not in sections:
sections[current_section] = {}
else:
raise ValueError(f"Malformed section header: {line}")
elif '=' in line or ':' in line:
if buffer:
sections[current_section][buffer[0].strip().lower()] = '\n'.join(buffer[1:])
buffer.clear()
key, value = re.split(r'[=:]', line, maxsplit=1)
key = key.strip().lower()
if value.startswith('"') and value.endswith('"'):
value = value[1:-1]
else:
value = value.strip()
comment_start = value.find(' ;')
if comment_start == -1:
comment_start = value.find(' #')
if comment_start != -1:
value = value[:comment_start]
buffer.append(key)
buffer.append(value)
elif line.startswith((' ', '\t')):
if buffer:
buffer.append(line.strip())
else:
raise ValueError(f"Continuation line without preceding key: {line}")
else:
raise ValueError(f"Malformed key-value pair: {line}")
if buffer:
sections[current_section][buffer[0].strip().lower()] = '\n'.join(buffer[1:])
if current_section != 'DEFAULT':
default_section = sections.pop('DEFAULT', {})
for section in sections.values():
for key in section.keys():
section[key] = interpolate(section[key], current_section, set())
return sections
``` Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from time import time
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock=time):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
def put(self, key, value):
if self.capacity == 0:
return
if key in self.cache:
del self.cache[key]
elif len(self.cache) >= self.capacity:
self.expire_entries()
if len(self.cache) >= self.capacity:
self.cache.popitem(last=False)
self.stats["evictions"] += 1
self.cache[key] = (value, self.clock() + self.ttl)
def get(self, key, default=None):
if self.capacity == 0:
self.stats["misses"] += 1
return default
self.expire_entries()
if key in self.cache:
value, expiry = self.cache.pop(key)
self.cache[key] = (value, self.clock() + self.ttl)
self.stats["hits"] += 1
return value
self.stats["misses"] += 1
return default
def expire_entries(self):
now = self.clock()
while self.cache and next(iter(self.cache.values()))[1] <= now:
key, _ = self.cache.popitem(last=False)
self.stats["expirations"] += 1
def __len__(self):
self.expire_entries()
return len(self.cache)
def stats(self):
return self.stats
``` Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
if not s or not s.strip():
raise ValueError("Empty or blank string")
total_seconds = 0
used_units = set()
current_number = ""
last_unit = ""
s = s.strip() + " " # Add a space to process the last unit
for i, char in enumerate(s):
if char.isdigit():
current_number += char
elif char.isspace():
if current_number:
if last_unit:
raise ValueError("Numbers must be followed by a unit")
continue
elif char.lower() in {'h', 'm', 's'}:
if not current_number:
raise ValueError("Unit must be preceded by a number")
value = int(current_number)
unit = char.lower()
if unit in used_units:
raise ValueError(f"Unit '{char}' appears more than once")
if (last_unit and (last_unit == 'h' and unit != 'm') or
last_unit == 'm' and unit == 'h'):
raise ValueError("Units must appear in the order h, m, s")
if unit == 'h':
total_seconds += value * 3600
elif unit == 'm':
total_seconds += value * 60
else: # unit == 's'
total_seconds += value
used_units.add(unit)
last_unit = unit
current_number = ""
else:
raise ValueError(f"Unknown unit '{char}'")
return total_seconds
``` Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
nums = sorted(set(nums))
ranges = []
start = end = nums[0]
for num in nums[1:]:
if num == end + 1:
end = num
else:
if start == end:
ranges.append(f"{start}")
elif end == start + 1:
ranges.append(f"{start},{end}")
else:
ranges.append(f"{start}..{end}")
start = end = num
if start == end:
ranges.append(f"{start}")
elif end == start + 1:
ranges.append(f"{start},{end}")
else:
ranges.append(f"{start}..{end}")
return ",".join(ranges)
``` Coding: Top customers Didn't work
```python
from decimal import Decimal
from typing import List, Dict, Tuple
def parse_amount(amount: str) -> Decimal:
amount = amount.replace('$', '').replace(',', '')
return Decimal(amount)
def top_customers(rows: List[Dict[str, str]], n: int) -> List[Tuple[str, float]]:
customer_sums = {}
for row in rows:
customer = row['customer'].strip()
if not customer:
continue
try:
amount = parse_amount(row['amount'])
except (ValueError, Decimal):
continue
if customer.lower() not in customer_sums:
customer_sums[customer.lower()] = {'name': customer, 'total': Decimal(0)}
customer_sums[customer.lower()]['total'] += amount
sorted_customers = sorted(customer_sums.values(), key=lambda x: (-x['total'], x['name']))
result = [(customer['name'], round(float(customer['total']), 2)) for customer in sorted_customers[:n]]
return result
``` Coding: Slugify Worked
```python
import re
import unicodedata
def slugify(title: str, max_len: int = 50) -> str:
# Transliterate accented characters to ASCII and drop other non-ASCII characters
normalized_title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
# Lowercase
lowercase_title = normalized_title.lower()
# Replace every run of characters other than a-z and 0-9 with a single "-"
slug = re.sub(r'[^a-z0-9]+', '-', lowercase_title).strip('-')
# If the slug is longer than max_len, shorten it without cutting a word
if len(slug) > max_len:
words = slug.split('-')
shortened_slug = ''
for word in words:
if len(shortened_slug) + len(word) + (1 if shortened_slug else 0) <= max_len:
shortened_slug += ('-' if shortened_slug else '') + word
else:
break
if not shortened_slug:
shortened_slug = slug[:max_len].rstrip('-')
slug = shortened_slug
return slug
``` Coding: Token bucket Worked
```python
from time import time as system_time
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock):
self._capacity = capacity
self._refill_per_sec = refill_per_sec
self._clock = clock
self._tokens = capacity
self._last_refill_time = self._clock()
@property
def tokens(self) -> float:
self._refill()
return self._tokens
def allow(self, cost: float = 1) -> bool:
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
return False
def _refill(self):
current_time = self._clock()
time_since_last_refill = current_time - self._last_refill_time
new_tokens = time_since_last_refill * self._refill_per_sec
self._tokens = min(self._tokens + new_tokens, self._capacity)
self._last_refill_time = current_time
``` Decisions: Refund window (hard) 50% right
{
"outcome": {
"full_refund": 0.0,
"partial_refund": 0.7,
"store_credit": 0.3,
"nothing": 0.0
},
"defective": {
"true": 0.0,
"false": 1.0
}
} Decisions: Incident matrix (hard) 100% right
{
"severity": {
"0": 0,
"1": 0.6,
"2": 0.3,
"3": 0.1
},
"page": {
"true": 0.4,
"false": 0.6
}
} Decisions: Tool followup (hard) 100% right
{
"tool": {
"calendar": 0.8,
"none": 0.2,
"web_search": 0,
"calculator": 0,
"email": 0
},
"confirm": {
"true": 0.7,
"false": 0.3
}
} Decisions: Legit security alert (hard) 100% right
{
"phishing": {
"true": 0.05,
"false": 0.95
},
"action_needed": {
"true": 0.1,
"false": 0.9
}
} Decisions: Meeting slot (hard) 0% right
{
"slot": {
"A": 0.0,
"B": 0.33,
"C": 0.33,
"D": 0.34
},
"raj_last": {
"true": 0.0,
"false": 1.0
}
} Decisions: Review mixed (hard) 100% right
{
"hardware": {
"true": 0.7,
"false": 0.3
},
"support": {
"true": 0.9,
"false": 0.1
}
} Decisions: Support checkout down 100% right
{
"department": {
"billing": 0.1,
"technical": 0.6,
"account": 0.1,
"shipping": 0.05,
"sales": 0.15
},
"urgency": {
"0": 0.05,
"1": 0.1,
"2": 0.2,
"3": 0.65
},
"outage": {
"true": 0.9,
"false": 0.1
}
} Decisions: Refund wrong plan 100% right
{
"department": {
"billing": 0.9,
"technical": 0.05,
"account": 0.025,
"shipping": 0.0125,
"sales": 0.0125
},
"refund": {
"true": 0.85,
"false": 0.15
},
"tone": {
"frustrated": 0.1,
"calm": 0.9
}
} Decisions: Moderation doxxing 100% right
{
"policy": {
"none": 0.05,
"harassment": 0.85,
"hate": 0.05,
"spam": 0.02,
"self_harm": 0.03
},
"personal_info": {
"true": 0.9,
"false": 0.1
}
} Decisions: Route calendar 100% right
{
"tool": {
"web_search": 0.05,
"calculator": 0.0,
"calendar": 0.7,
"email": 0.2,
"none": 0.05
},
"confirm": {
"true": 0.9,
"false": 0.1
}
} Decisions: Doc invoice missing due 100% right
{
"doc_type": {
"invoice": 1.0,
"resume": 0.0,
"contract": 0.0,
"bank_statement": 0.0,
"other": 0.0
},
"missing_due_date": {
"true": 1.0,
"false": 0.0
}
} Decisions: Phishing paypal 100% right
{
"phishing": {
"true": 0.95,
"false": 0.05
},
"risk": {
"0": 0.05,
"1": 0.05,
"2": 0.10,
"3": 0.80
}
} Decisions: Pii ssn email 100% right
{
"data_kind": {
"none": 0.0,
"contact": 0.2,
"financial": 0.1,
"government_id": 0.6,
"health": 0.1
},
"sensitive": {
"true": 0.9,
"false": 0.1
}
} Decisions: Review mixed 100% right
{
"sentiment": {
"positive": 0.2,
"neutral": 0.2,
"negative": 0.6
},
"defect": {
"true": 0.9,
"false": 0.1
},
"recommend": {
"true": 0.1,
"false": 0.9
}
} Documents: Saas escalator (hard) 10% right
{
"year2_price_per_seat_month": 47.75,
"year3_price_per_seat_month": 47.75,
"year1_invoice": 54000.00,
"year2_invoice": 64260.00,
"addon_months_billed": 10,
"addon_invoice": 64500.00,
"year3_invoice": 135000.00,
"year3_discount_percent": 10.00,
"total_contract_value": 253260.00,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 71% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{ "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
{ "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
{ "date": "2025-02-25", "category": "meals", "amount_usd": 229.39 },
{ "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.72 },
{ "date": "2025-02-24", "category": "lodging", "amount_usd": 467.52 }
],
"rejected_item_count": 1,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 2350.01,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 92% right
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2236.00,
"late_fee_from_2025_06_01": 111.80,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2306.00,
"move_in_payment": 4730.00
} Documents: Ticket SLA 82% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-13T13:30",
"sla_due_utc": "2025-09-13T18:30:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 67% right
{
"q3_total_usd": 15346,
"q2_total_usd": 14464,
"q2_central_originally_reported_usd": 3047000,
"q2_to_q3_change_pct": 5.4,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731000,
"west_excluding_mountain_q3_usd": 4201000
} Size: 73B parameters. First tested OCT 11.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.