Review · updated OCT 11

Qwen2.5 VL 72B Instruct review: not one we'd recommend right now

It scored 14 out of 100, #53 of 56. It solved 4 of 30 coding jobs and scored 14 on reading documents. Runs on a Mac with 64 GB.

The short version
  • Qwen2.5 VL 72B Instruct is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's not one we'd recommend right now: 14 out of 100, #53 of 56.
  • It solved 4 of 30 coding jobs and scored 14 on reading documents. On our hardest tasks it scored 8.
  • Runs on a Mac with 64 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen2.5 VL 72B Instruct got 4 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen2.5 VL 72B Instruct scored 14; the best model scored 100.

TestScorePublic questionsSecret questions
Coding131413
Reading documents141813
Decisions898989

On the 18 hardest tasks (included in the scores above) it scored 8. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen2.5 VL 72B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBYesQ4
Mac, 96 GBYesQ6
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Qwen2.5 VL 72B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen2.5 VL 72B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    def interpolate(value: str, section_name: str, sections: dict[str, dict[str, str]]) -> str:
        def replace(match):
            if ':' in match.group(1):
                sect, key = match.group(1).split(':', 1)
                if sect not in sections:
                    raise KeyError(f"Section '{sect}' not found")
                return sections[sect][key]
            else:
                key = match.group(1)
                if key in sections[section_name]:
                    return sections[section_name][key]
                elif 'DEFAULT' in sections and key in sections['DEFAULT']:
                    return sections['DEFAULT'][key]
                else:
                    raise KeyError(f"Key '{key}' not found in section '{section_name}'")
        
        while '${' in value:
            try:
                new_value = re.sub(r'\$\{([^}]+)\}', replace, value)
                if new_value == value:
                    raise ValueError("Interpolation cycle detected")
                value = new_value
            except KeyError as e:
                raise KeyError(str(e)) from None
        return value

    lines = text.splitlines()
    current_section = None
    sections = {}
    default_section = {}

    for line in lines:
        line = line.strip()

        if not line or line.startswith(';') or line.startswith('#'):
            continue

        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            current_section = section_name
            sections[current_section] = {}
            continue

        if '=' in line or ':' in line:
            if current_section is None:
                raise ValueError("Key outside of any section")

            delimiter = '=' if '=' in line else ':'
            key, value = line.split(delimiter, 1)
            key = key.strip().lower()
            value = value.strip()

            if value.startswith('"') and value.endswith('"'):
                value = value[1:-1]

            if ' ;' in value or ' #' in value:
                value = value[:value.find(' ;')].strip() if ' ;' in value else value[:value.find(' #')].strip()

            if line.startswith(' ') or line.startswith('\t'):
                if key not in sections[current_section]:
                    raise ValueError("Continuation line without preceding key")
                sections[current_section][key] += '\n' + value
            else:
                sections[current_section][key] = value

            if current_section != 'DEFAULT':
                if key in default_section:
                    sections[current_section][key] = default_section[key]
                sections[current_section][key] = interpolate(sections[current_section][key], current_section, sections)

        else:
            raise ValueError("Malformed line")

    if 'DEFAULT' in sections:
        default_section = sections.pop('DEFAULT')

    for section in sections.values():
        for key in list(section.keys()):
            if key in default_section and key not in section:
                section[key] = default_section[key]

    return sections
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
import time

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def _remove_expired(self):
        now = self.clock()
        keys_to_remove = [key for key, (_, expiry) in self.cache.items() if now >= expiry]
        for key in keys_to_remove:
            del self.cache[key]
            self.stats["expirations"] += 1

    def put(self, key, value):
        now = self.clock()
        if self.capacity == 0:
            return
        self._remove_expired()
        if key in self.cache:
            del self.cache[key]
        elif len(self.cache) >= self.capacity:
            self.cache.popitem(last=False)
            self.stats["evictions"] += 1
        self.cache[key] = (value, now + self.ttl)
        self.cache.move_to_end(key)

    def get(self, key, default=None):
        now = self.clock()
        self._remove_expired()
        if key not in self.cache or now >= self.cache[key][1]:
            self.stats["misses"] += 1
            return default
        value, _ = self.cache[key]
        self.cache.move_to_end(key)
        self.stats["hits"] += 1
        return value

    def __len__(self):
        self._remove_expired()
        return len(self.cache)

    def stats(self):
        return self.stats.copy()
```
Coding: Parse duration Didn't work
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    
    sorted_nums = sorted(set(nums))
    ranges = []
    start = end = sorted_nums[0]
    
    for num in sorted_nums[1:]:
        if num == end + 1:
            end = num
        else:
            if start == end:
                ranges.append(str(start))
            elif start + 1 == end:
                ranges.extend([str(start), str(end)])
            else:
                ranges.append(f"{start}..{end}")
            start = end = num
    
    if start == end:
        ranges.append(str(start))
    elif start + 1 == end:
        ranges.extend([str(start), str(end)])
    else:
        ranges.append(f"{start}..{end}")
    
    return ",".join(ranges)
```
Coding: Top customers Didn't work
```
Coding: Slugify Didn't work
```python
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    # Step 1: Transliterate accented characters to ASCII and drop other non-ASCII characters
    normalized_title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
    
    # Step 2: Lowercase
    lowercased_title = normalized_title.lower()
    
    # Step 3: Replace runs of non-alphanumeric characters with a single "-"
    import re
    slug = re.sub(r'[^a-z0-9]+', '-', lowercased_title).strip('-')
    
    # Step 4: Shorten if necessary
    words = slug.split('-')
    shortened_slug = ''
    for word in words:
        if len(shortened_slug) + len(word) + (len(shortened_slug) > 0) <= max_len:
            shortened_slug += ('-' if shortened_slug else '') + word
        else:
            break
    
    # Step 5: Ensure the result never ends with "-"
    return shortened_slug.rstrip('-')
```
Coding: Token bucket Didn't work
```
Decisions: Refund window (hard) 50% right
{
    "outcome": {
        "full_refund": 0,
        "partial_refund": 0.95,
        "store_credit": 0.05,
        "nothing": 0
    },
    "defective": {
        "true": 0,
        "false": 1
    }
}
Decisions: Incident matrix (hard) 100% right
{
    "severity": {
        "0": 0.0,
        "1": 0.7,
        "2": 0.3,
        "3": 0.0
    },
    "page": {
        "true": 0.3,
        "false": 0.7
    }
}
Decisions: Tool followup (hard) 100% right
{
    "tool": {
        "calendar": 1,
        "web_search": 0,
        "calculator": 0,
        "email": 0,
        "none": 0
    },
    "confirm": {
        "true": 1,
        "false": 0
    }
}
Decisions: Legit security alert (hard) 50% right
{
    "phishing": {
        "true": 0.05,
        "false": 0.95
    },
    "action_needed": {
        "true": 0.9,
        "false": 0.1
    }
}
Decisions: Meeting slot (hard) 50% right
{
    "slot": {
        "A": 0,
        "B": 0,
        "C": 1,
        "D": 0
    },
    "raj_last": {
        "true": 1,
        "false": 0
    }
}
Decisions: Review mixed (hard) 100% right
{
    "hardware": {
        "true": 1,
        "false": 0
    },
    "support": {
        "true": 1,
        "false": 0
    }
}
Decisions: Support checkout down 100% right
{
    "department": {
        "billing": 0.05,
        "technical": 0.8,
        "account": 0.05,
        "shipping": 0.05,
        "sales": 0.05
    },
    "urgency": {
        "0": 0.0,
        "1": 0.0,
        "2": 0.1,
        "3": 0.9
    },
    "outage": {
        "true": 0.95,
        "false": 0.05
    }
}
Decisions: Refund wrong plan 100% right
{
    "department": {
        "billing": 1,
        "technical": 0,
        "account": 0,
        "shipping": 0,
        "sales": 0
    },
    "refund": {
        "true": 1,
        "false": 0
    },
    "tone": {
        "frustrated": 0,
        "calm": 1
    }
}
Decisions: Moderation doxxing 100% right
{
    "policy": {
        "none": 0,
        "harassment": 0.95,
        "hate": 0.02,
        "spam": 0.02,
        "self_harm": 0.01
    },
    "personal_info": {
        "true": 0.98,
        "false": 0.02
    }
}
Decisions: Route calendar 100% right
{
    "tool": {
        "web_search": 0,
        "calculator": 0,
        "calendar": 0.8,
        "email": 0.2,
        "none": 0
    },
    "confirm": {
        "true": 1,
        "false": 0
    }
}
Decisions: Doc invoice missing due 100% right
{
    "doc_type": {
        "invoice": 1,
        "resume": 0,
        "contract": 0,
        "bank_statement": 0,
        "other": 0
    },
    "missing_due_date": {
        "true": 1,
        "false": 0
    }
}
Decisions: Phishing paypal 100% right
{
    "phishing": {
        "true": 0.95,
        "false": 0.05
    },
    "risk": {
        "0": 0.0,
        "1": 0.0,
        "2": 0.05,
        "3": 0.95
    }
}
Decisions: Pii ssn email 100% right
{
    "data_kind": {
        "none": 0,
        "contact": 0.2,
        "financial": 0,
        "government_id": 0.6,
        "health": 0.2
    },
    "sensitive": {
        "true": 0.9,
        "false": 0.1
    }
}
Decisions: Review mixed 100% right
{
    "sentiment": {
        "positive": 0.2,
        "neutral": 0.1,
        "negative": 0.7
    },
    "defect": {
        "true": 0.9,
        "false": 0.1
    },
    "recommend": {
        "true": 0.1,
        "false": 0.9
    }
}
Documents: Saas escalator (hard) 0% right
Documents: Expense thread 92% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 130.00,
  "total_reimbursable_usd": 2250.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 0% right
Documents: Ticket SLA 0% right
Documents: Sales footnotes 0% right

Size: 73B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.