Review · updated OCT 11

Qwen2.5 7B Instruct review: not one we'd recommend right now

It scored 24 out of 100, #45 of 56. It solved 3 of 30 coding jobs and scored 37 on reading documents. Runs on an 8 GB graphics card or a Mac with 16 GB.

The short version
  • Qwen2.5 7B Instruct is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's not one we'd recommend right now: 24 out of 100, #45 of 56.
  • It solved 3 of 30 coding jobs and scored 37 on reading documents. On our hardest tasks it scored 11.
  • Runs on an 8 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen2.5 7B Instruct got 3 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen2.5 7B Instruct scored 37; the best model scored 100.

TestScorePublic questionsSecret questions
Coding10013
Reading documents375732
Decisions717570

On the 18 hardest tasks (included in the scores above) it scored 11. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

The Q4 download

Most people don't run the full-size model; they download a smaller Q4 version. We ran that version with Ollama on our Mac mini (M4, 16 GB) and gave it the same questions. It scored 24, about the same as the full-size model (24). You lose almost nothing by downloading it.

TestFull size (online)Q4 (our Mac)
Coding1013
Reading documents3734
Hardest tasks119
Overall2424

Can your computer run it?

You can download Qwen2.5 7B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBYesQ6
Graphics card, 12 GBYesQ8
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ8
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run qwen2.5:7b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen2.5 7B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from collections import defaultdict

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    lines = text.splitlines()
    sections = defaultdict(dict)
    defaults = {}
    current_section = 'DEFAULT'
    
    for line in lines:
        line = line.strip()
        if not line or line.startswith(';') or line.startswith('#'):
            continue
        
        if line.startswith('[') and line.endswith(']'):
            current_section = line[1:-1].strip()
            if current_section == 'DEFAULT':
                defaults = sections[current_section]
                continue
            sections[current_section] = {}
        else:
            key, value = line.split(maxsplit=1)
            key = key.strip().lower()
            value = value.strip()
            
            if value.startswith('"') and value.endswith('"'):
                value = value[1:-1]
            
            while line.endswith('\\'):
                next_line = next((l for l in lines[lines.index(line) + 1:] if l.strip()), '').strip()
                value += '\n' + next_line
            
            # Resolve interpolation
            def resolve(key):
                nonlocal sections, defaults
                if key in defaults:
                    return defaults[key]
                elif key in sections[current_section]:
                    return sections[current_section][key]
                else:
                    raise KeyError(f"Key {key} not found")
            
            value = re.sub(r'\$\{(\w+)\}', lambda m: resolve(m.group(1)), value)
            value = re.sub(r'\$\{(\w+):(\w+)\}', lambda m: resolve(f"{m.group(1)}.{m.group(2)}"), value)
            
            sections[current_section][key] = value
    
    return {k: v for k, v in sections.items() if k != 'DEFAULT'}
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from typing import Any, Callable

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock: Callable[[], float] = None):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock or (lambda: 0)
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def put(self, key: Any, value: Any) -> None:
        current_time = self.clock()
        if key in self.cache:
            self.cache.move_to_end(key)
            self.cache[key] = value
            self.cache[key]["expiry"] = current_time + self.ttl
        else:
            if len(self.cache) >= self.capacity:
                while True:
                    oldest_key = next(iter(self.cache))
                    if self.cache[oldest_key]["expiry"] > current_time:
                        break
                    del self.cache[oldest_key]
                    self.stats["expirations"] += 1
                    self.stats["misses"] += 1
            self.cache[key] = {"value": value, "expiry": current_time + self.ttl}

    def get(self, key: Any, default: Any = None) -> Any:
        current_time = self.clock()
        if key not in self.cache:
            self.stats["misses"] += 1
            return default
        entry = self.cache[key]
        if entry["expiry"] <= current_time:
            del self.cache[key]
            self.stats["expirations"] += 1
            self.stats["misses"] += 1
            return default
        self.cache.move_to_end(key)
        entry["expiry"] = current_time + self.ttl
        self.stats["hits"] += 1
        return entry["value"]

    def __len__(self) -> int:
        current_time = self.clock()
        return sum(1 for k in self.cache if self.cache[k]["expiry"] > current_time)

    def stats(self) -> dict:
        return self.stats
```
Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
    import re
    
    if not s.strip():
        raise ValueError("Duration string cannot be empty or blank")
    
    pattern = r"(\d+)(h|m|s)"
    matches = re.findall(pattern, s.replace(" ", "").lower())
    
    if not matches:
        raise ValueError("Duration string must contain at least one valid unit with a number")
    
    units = {"h": 3600, "m": 60, "s": 1}
    total_seconds = 0
    
    for amount, unit in matches:
        if unit not in units:
            raise ValueError(f"Unknown unit: {unit}")
        total_seconds += int(amount) * units[unit]
    
    return total_seconds
```
Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    
    nums.sort()
    result = []
    start = None
    
    for i in range(len(nums)):
        if start is None or nums[i] == nums[i - 1] + 1:
            start = nums[i]
        else:
            if i > 0 and nums[i - 1] - nums[i - 2] >= 2:
                result.append(f"{start}..{nums[i - 1]}")
                start = nums[i]
            else:
                result.extend(map(str, range(nums[i - 1], nums[i])))
    
    if start is not None:
        if len(nums) - nums.index(start) >= 2:
            result.append(f"{start}..{nums[-1]}")
        else:
            result.extend(map(str, range(start, nums[-1] + 1)))
    
    return ",".join(result)
```
Coding: Top customers Didn't work
```python
from typing import List, Tuple
import re

def top_customers(rows: List[dict], n: int) -> List[Tuple[str, float]]:
    def parse_amount(amount_str: str) -> float:
        amount_str = amount_str.strip()
        if not amount_str:
            return 0.0
        match = re.match(r'^-?\$?(\d+(,\d{3})*(\.\d+)?|\.\d+)$', amount_str.replace(',', ''))
        if not match:
            return 0.0
        return round(float(match.group(1).replace(',', '')), 2)

    customers = {}
    for row in rows:
        name = row.get('customer').strip().lower()
        if not name:
            continue
        amount = parse_amount(row.get('amount'))
        if amount == 0.0:
            continue
        if name not in customers:
            customers[name] = amount
        else:
            customers[name] += amount

    sorted_customers = sorted(customers.items(), key=lambda x: (-x[1], x[0]))
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import re
from unicodedata import normalize

def slugify(title: str, max_len: int = 50) -> str:
    # Transliterate accented characters to ASCII and lowercase
    ascii_title = normalize('NFKD', title).encode('ASCII', 'ignore').decode().lower()
    
    # Replace every run of non-alphanumeric characters with a single "-"
    slug = re.sub(r'[^a-z0-9]+', '-', ascii_title)
    
    # Strip leading/trailing "-"
    slug = slug.strip('-')
    
    # If the slug is longer than max_len, shorten it without cutting a word
    if len(slug) > max_len:
        # Find the last space before max_len that allows us to keep whole words
        last_space = max((i for i in range(max_len) if slug[i] == '-'), default=max_len)
        slug = slug[:last_space]
        # Ensure the result does not end with "-"
        if slug.endswith('-'):
            slug = slug[:-1]
    
    return slug
```
Coding: Token bucket Didn't work
```python
from collections import deque
import time

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock=time.time):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.tokens = capacity
        self.clock = clock
        self.last_refill_time = clock()
        self.refill_queue = deque()

    def _refill(self):
        current_time = self.clock()
        if not self.refill_queue or current_time >= self.refill_queue[-1]:
            elapsed_time = max(current_time - self.last_refill_time, 0)
            tokens_to_add = min(elapsed_time * self.refill_per_sec, self.capacity - self.tokens)
            self.tokens = min(self.capacity, self.tokens + tokens_to_add)
            self.last_refill_time = current_time
            if tokens_to_add > 0:
                self.refill_queue.append(current_time + 1 / self.refill_per_sec)

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens

    @tokens.setter
    def tokens(self, value: float):
        self._tokens = min(value, self.capacity)
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "partial_refund": 1.0
  },
  "defective": {
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 0% right
{
  "severity": {
    "0": 0.1,
    "1": 0.2,
    "2": 0.3,
    "3": 0.4
  },
  "page": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.0,
    "false": 1.0
  },
  "action_needed": {
    "true": 0.2,
    "false": 0.8
  }
}
Decisions: Meeting slot (hard) 0% right
{
  "slot": {
    "A": 0.0,
    "B": 0.25,
    "C": 0.5,
    "D": 0.25
  },
  "raj_last": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 1,
    "false": 0
  },
  "support": {
    "true": 1,
    "false": 0
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0,
    "technical": 1,
    "account": 0,
    "shipping": 0,
    "sales": 0
  },
  "urgency": {
    "0": 0,
    "1": 0,
    "2": 0,
    "3": 1
  },
  "outage": {
    "true": 1,
    "false": 0
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 1.0
  },
  "refund": {
    "true": 1.0
  },
  "tone": {
    "calm": 1.0
  }
}
Decisions: Moderation doxxing 50% right
{
  "policy": {
    "none": 0.9,
    "harassment": 0.1
  },
  "personal_info": {
    "true": 1
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
Decisions: Doc invoice missing due 50% right
{
  "doc_type": {
    "invoice": 1.0
  },
  "missing_due_date": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Phishing paypal 100% right
{
  "phishing": {
    "true": 1,
    "false": 0
  },
  "risk": {
    "0": 0,
    "1": 0,
    "2": 0,
    "3": 1
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0,
    "contact": 0,
    "financial": 0.1,
    "government_id": 1,
    "health": 0
  },
  "sensitive": {
    "true": 1,
    "false": 0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "negative": 1
  },
  "defect": {
    "true": 1
  },
  "recommend": {
    "false": 1
  }
}
Documents: Saas escalator (hard) 30% right
{
  "year2_price_per_seat_month": 47.68,
  "year3_price_per_seat_month": 49.91,
  "year1_invoice": 54000.00,
  "year2_invoice": null,
  "addon_months_billed": 6,
  "addon_invoice": 23040.00,
  "year3_invoice": null,
  "year3_discount_percent": 15.00,
  "total_contract_value": 133040.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 75% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {"date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60},
    {"date": "2025-02-24", "category": "ground_transport", "amount_usd": 39.72},
    {"date": "2025-02-25", "category": "meals", "amount_usd": 230.44},
    {"date": "2025-02-26", "category": "lodging", "amount_usd": 469.08},
    {"date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.84}
  ],
  "rejected_item_count": 2,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 1883.84,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 67% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2254,
  "late_fee_from_2025_06_01": 112.7,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2289,
  "move_in_payment": 4500
}
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-13T17:00",
  "sla_due_utc": "2025-09-13T22:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 33% right
{
  "q3_total_usd": 1951000,
  "q2_total_usd": 1622000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": -15.6,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "West",
  "regions_declining_q2_to_q3": ["Central"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 418000
}

Size: 7.6B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.