Review · updated OCT 11

Qwen2.5 72B Instruct review: not one we'd recommend right now

It scored 46 out of 100, #31 of 56. It solved 12 of 30 coding jobs and scored 52 on reading documents. Runs on a Mac with 64 GB.

The short version
  • Qwen2.5 72B Instruct is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's not one we'd recommend right now: 46 out of 100, #31 of 56.
  • It solved 12 of 30 coding jobs and scored 52 on reading documents. On our hardest tasks it scored 23.
  • Runs on a Mac with 64 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen2.5 72B Instruct got 12 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen2.5 72B Instruct scored 52; the best model scored 100.

TestScorePublic questionsSecret questions
Coding404339
Reading documents526449
Decisions888988

On the 18 hardest tasks (included in the scores above) it scored 23. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen2.5 72B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBYesQ4
Mac, 96 GBYesQ6
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Qwen2.5 72B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen2.5 72B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, List, Tuple

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    def interpolate(value: str, section: str, seen: set) -> str:
        while True:
            match = re.search(r'\$\{([^}]+)\}', value)
            if not match:
                break
            ref = match.group(1)
            if ':' in ref:
                ref_section, ref_key = ref.split(':', 1)
            else:
                ref_section = section
                ref_key = ref.lower()
            
            if (ref_section, ref_key) in seen:
                raise ValueError(f"Reference cycle detected: {ref}")
            seen.add((ref_section, ref_key))
            
            if ref_section == 'DEFAULT':
                if ref_key not in default_section:
                    raise KeyError(f"Missing key in DEFAULT: {ref_key}")
                ref_value = default_section[ref_key]
            elif ref_section in sections:
                if ref_key not in sections[ref_section]:
                    if ref_key in default_section:
                        ref_value = default_section[ref_key]
                    else:
                        raise KeyError(f"Missing key in section {ref_section}: {ref_key}")
                else:
                    ref_value = sections[ref_section][ref_key]
            else:
                raise KeyError(f"Missing section: {ref_section}")
            
            value = value.replace(match.group(0), ref_value)
            seen.discard((ref_section, ref_key))
        return value

    sections: Dict[str, Dict[str, str]] = {}
    default_section: Dict[str, str] = {}
    current_section: str = 'DEFAULT'

    lines = text.splitlines()
    buffer: List[str] = []

    for line in lines:
        line = line.strip()
        if not line or line.startswith((';', '#')):
            continue

        if line.startswith(('[', ']')):
            match = re.match(r'\[(\w+)\]', line)
            if match:
                if buffer:
                    sections[current_section][buffer[0].strip().lower()] = '\n'.join(buffer[1:])
                    buffer.clear()
                current_section = match.group(1)
                if current_section not in sections:
                    sections[current_section] = {}
            else:
                raise ValueError(f"Malformed section header: {line}")
        elif '=' in line or ':' in line:
            if buffer:
                sections[current_section][buffer[0].strip().lower()] = '\n'.join(buffer[1:])
                buffer.clear()
            key, value = re.split(r'[=:]', line, maxsplit=1)
            key = key.strip().lower()
            if value.startswith('"') and value.endswith('"'):
                value = value[1:-1]
            else:
                value = value.strip()
                comment_start = value.find(' ;')
                if comment_start == -1:
                    comment_start = value.find(' #')
                if comment_start != -1:
                    value = value[:comment_start]
            buffer.append(key)
            buffer.append(value)
        elif line.startswith((' ', '\t')):
            if buffer:
                buffer.append(line.strip())
            else:
                raise ValueError(f"Continuation line without preceding key: {line}")
        else:
            raise ValueError(f"Malformed key-value pair: {line}")

    if buffer:
        sections[current_section][buffer[0].strip().lower()] = '\n'.join(buffer[1:])

    if current_section != 'DEFAULT':
        default_section = sections.pop('DEFAULT', {})

    for section in sections.values():
        for key in section.keys():
            section[key] = interpolate(section[key], current_section, set())

    return sections
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from time import time

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock=time):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def put(self, key, value):
        if self.capacity == 0:
            return
        if key in self.cache:
            del self.cache[key]
        elif len(self.cache) >= self.capacity:
            self.expire_entries()
            if len(self.cache) >= self.capacity:
                self.cache.popitem(last=False)
                self.stats["evictions"] += 1
        self.cache[key] = (value, self.clock() + self.ttl)

    def get(self, key, default=None):
        if self.capacity == 0:
            self.stats["misses"] += 1
            return default
        self.expire_entries()
        if key in self.cache:
            value, expiry = self.cache.pop(key)
            self.cache[key] = (value, self.clock() + self.ttl)
            self.stats["hits"] += 1
            return value
        self.stats["misses"] += 1
        return default

    def expire_entries(self):
        now = self.clock()
        while self.cache and next(iter(self.cache.values()))[1] <= now:
            key, _ = self.cache.popitem(last=False)
            self.stats["expirations"] += 1

    def __len__(self):
        self.expire_entries()
        return len(self.cache)

    def stats(self):
        return self.stats
```
Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")

    total_seconds = 0
    used_units = set()
    current_number = ""
    last_unit = ""
    s = s.strip() + " "  # Add a space to process the last unit

    for i, char in enumerate(s):
        if char.isdigit():
            current_number += char
        elif char.isspace():
            if current_number:
                if last_unit:
                    raise ValueError("Numbers must be followed by a unit")
                continue
        elif char.lower() in {'h', 'm', 's'}:
            if not current_number:
                raise ValueError("Unit must be preceded by a number")
            value = int(current_number)
            unit = char.lower()
            if unit in used_units:
                raise ValueError(f"Unit '{char}' appears more than once")
            if (last_unit and (last_unit == 'h' and unit != 'm') or
                last_unit == 'm' and unit == 'h'):
                raise ValueError("Units must appear in the order h, m, s")
            if unit == 'h':
                total_seconds += value * 3600
            elif unit == 'm':
                total_seconds += value * 60
            else:  # unit == 's'
                total_seconds += value
            used_units.add(unit)
            last_unit = unit
            current_number = ""
        else:
            raise ValueError(f"Unknown unit '{char}'")
    
    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    
    nums = sorted(set(nums))
    ranges = []
    start = end = nums[0]
    
    for num in nums[1:]:
        if num == end + 1:
            end = num
        else:
            if start == end:
                ranges.append(f"{start}")
            elif end == start + 1:
                ranges.append(f"{start},{end}")
            else:
                ranges.append(f"{start}..{end}")
            start = end = num
    
    if start == end:
        ranges.append(f"{start}")
    elif end == start + 1:
        ranges.append(f"{start},{end}")
    else:
        ranges.append(f"{start}..{end}")
    
    return ",".join(ranges)
```
Coding: Top customers Didn't work
```python
from decimal import Decimal
from typing import List, Dict, Tuple

def parse_amount(amount: str) -> Decimal:
    amount = amount.replace('$', '').replace(',', '')
    return Decimal(amount)

def top_customers(rows: List[Dict[str, str]], n: int) -> List[Tuple[str, float]]:
    customer_sums = {}
    
    for row in rows:
        customer = row['customer'].strip()
        if not customer:
            continue
        
        try:
            amount = parse_amount(row['amount'])
        except (ValueError, Decimal):
            continue
        
        if customer.lower() not in customer_sums:
            customer_sums[customer.lower()] = {'name': customer, 'total': Decimal(0)}
        
        customer_sums[customer.lower()]['total'] += amount
    
    sorted_customers = sorted(customer_sums.values(), key=lambda x: (-x['total'], x['name']))
    
    result = [(customer['name'], round(float(customer['total']), 2)) for customer in sorted_customers[:n]]
    
    return result
```
Coding: Slugify Worked
```python
import re
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    # Transliterate accented characters to ASCII and drop other non-ASCII characters
    normalized_title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
    
    # Lowercase
    lowercase_title = normalized_title.lower()
    
    # Replace every run of characters other than a-z and 0-9 with a single "-"
    slug = re.sub(r'[^a-z0-9]+', '-', lowercase_title).strip('-')
    
    # If the slug is longer than max_len, shorten it without cutting a word
    if len(slug) > max_len:
        words = slug.split('-')
        shortened_slug = ''
        for word in words:
            if len(shortened_slug) + len(word) + (1 if shortened_slug else 0) <= max_len:
                shortened_slug += ('-' if shortened_slug else '') + word
            else:
                break
        if not shortened_slug:
            shortened_slug = slug[:max_len].rstrip('-')
        slug = shortened_slug
    
    return slug
```
Coding: Token bucket Worked
```python
from time import time as system_time

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._tokens = capacity
        self._last_refill_time = self._clock()

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False

    def _refill(self):
        current_time = self._clock()
        time_since_last_refill = current_time - self._last_refill_time
        new_tokens = time_since_last_refill * self._refill_per_sec
        self._tokens = min(self._tokens + new_tokens, self._capacity)
        self._last_refill_time = current_time
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.7,
    "store_credit": 0.3,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0,
    "1": 0.6,
    "2": 0.3,
    "3": 0.1
  },
  "page": {
    "true": 0.4,
    "false": 0.6
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "calendar": 0.8,
    "none": 0.2,
    "web_search": 0,
    "calculator": 0,
    "email": 0
  },
  "confirm": {
    "true": 0.7,
    "false": 0.3
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.1,
    "false": 0.9
  }
}
Decisions: Meeting slot (hard) 0% right
{
  "slot": {
    "A": 0.0,
    "B": 0.33,
    "C": 0.33,
    "D": 0.34
  },
  "raj_last": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 0.7,
    "false": 0.3
  },
  "support": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0.1,
    "technical": 0.6,
    "account": 0.1,
    "shipping": 0.05,
    "sales": 0.15
  },
  "urgency": {
    "0": 0.05,
    "1": 0.1,
    "2": 0.2,
    "3": 0.65
  },
  "outage": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.9,
    "technical": 0.05,
    "account": 0.025,
    "shipping": 0.0125,
    "sales": 0.0125
  },
  "refund": {
    "true": 0.85,
    "false": 0.15
  },
  "tone": {
    "frustrated": 0.1,
    "calm": 0.9
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.05,
    "harassment": 0.85,
    "hate": 0.05,
    "spam": 0.02,
    "self_harm": 0.03
  },
  "personal_info": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0.05,
    "calculator": 0.0,
    "calendar": 0.7,
    "email": 0.2,
    "none": 0.05
  },
  "confirm": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Phishing paypal 100% right
{
  "phishing": {
    "true": 0.95,
    "false": 0.05
  },
  "risk": {
    "0": 0.05,
    "1": 0.05,
    "2": 0.10,
    "3": 0.80
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.2,
    "financial": 0.1,
    "government_id": 0.6,
    "health": 0.1
  },
  "sensitive": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.2,
    "neutral": 0.2,
    "negative": 0.6
  },
  "defect": {
    "true": 0.9,
    "false": 0.1
  },
  "recommend": {
    "true": 0.1,
    "false": 0.9
  }
}
Documents: Saas escalator (hard) 10% right
{
  "year2_price_per_seat_month": 47.75,
  "year3_price_per_seat_month": 47.75,
  "year1_invoice": 54000.00,
  "year2_invoice": 64260.00,
  "addon_months_billed": 10,
  "addon_invoice": 64500.00,
  "year3_invoice": 135000.00,
  "year3_discount_percent": 10.00,
  "total_contract_value": 253260.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 71% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    { "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
    { "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
    { "date": "2025-02-25", "category": "meals", "amount_usd": 229.39 },
    { "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.72 },
    { "date": "2025-02-24", "category": "lodging", "amount_usd": 467.52 }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 2350.01,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 92% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 4730.00
}
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-13T13:30",
  "sla_due_utc": "2025-09-13T18:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 67% right
{
  "q3_total_usd": 15346,
  "q2_total_usd": 14464,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 5.4,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 73B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.