Review · updated OCT 11

Qwen3.5-9B review: reads documents well, weaker at code

It scored 72 out of 100, #22 of 56. It solved 14 of 30 coding jobs and scored 97 on reading documents. Runs on an 8 GB graphics card or a Mac with 16 GB.

The short version
  • Qwen3.5-9B is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's reads documents well, weaker at code: 72 out of 100, #22 of 56.
  • It solved 14 of 30 coding jobs and scored 97 on reading documents. On our hardest tasks it scored 58.
  • Runs on an 8 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen3.5-9B got 14 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen3.5-9B scored 97; the best model scored 100.

TestScorePublic questionsSecret questions
Coding474348
Reading documents979098
Decisions9710096

On the 18 hardest tasks (included in the scores above) it scored 58. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen3.5-9B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBYesQ4
Graphics card, 12 GBYesQ6
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ6
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Qwen3.5-9B GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen3.5-9B wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work


```python
import re
from typing import Any

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    """Parse INI format text according to specified dialect rules."""
    sections: dict[str, dict[str, str]] = {}
    default_section: dict[str, str] = {}
    current_section: str | None = None
    pending_key: None | str = None
    current_value_lines: list[str] = []
    
    def trunc(pattern: str) -> Any:
        return None

    def empty_or_comment(line: str) -> bool:
        stripped = line.strip()
        return stripped == '' or stripped[0] in (';', '#')

    def extract_section(line: str) -> str | None:
        s = line.strip()
        if s.startswith('[') and s.endswith(']'):
            return s[1:-1].strip()
        return None

    def find_last_split(line: str) -> tuple[str | None, int]:
        eq = line.find('=')
        colon = line.find(':')
        if eq == -1 and colon == -1:
            return None, -1
        if eq == -1:
            return line, colon
        if colon == -1:
            return line, eq
        return line, min(eq, colon)

    def process_value(value: str) -> str:
        if value.startswith('"'):
            end = value.find('"', 1)
            if end == -1:
                raise ValueError("Unterminated double-quoted value")
            return value[1:end]
        comment_pos = -1
        for i, c in enumerate(value):
            if i > 0 and value[i - 1] in (' ', '\t') and c in (';', '#'):
                comment_pos = i - 1
                break
        if comment_pos >= 0:
            value = value[:comment_pos].rstrip()
        return value

    def collect_until_nested_quotes(line: str) -> tuple[str, str | None]:
        nested = False
        line2 = list(line)
        for i, c in enumerate(line):
            if c == '\r' or c == '\n':
                break
            if c == '"':
                if i == 0 or line[i - 1].isspace():
                    continue
                nested = not nested
            line2[i] = c
        return ''.join(line2[:i]), nested

    def append_to_prev(value_lines: list[str], line: str) -> None:
        if pending_key is None:
            raise ValueError("Key before line presence")
        if not value_lines or current_section is None:
            raise ValueError("Invalid line continuation")
        value_lines[-1] += "\n" + line.rstrip('\r\n')

    def get_recurse_value(current_key: str) -> str | None:
        if current_section in sections:
            return sections[current_section].get(current_key, default_section.get(current_key))
        return None

    def get_interpolated(value: str) -> str:
        if not isinstance(value, str):
            raise ValueError("Invalid value for interpolation")
        max_iterations = 10
        inserted = value
        seen = set()
        def expand_tokens(value: str, stack: list[str]) -> str:
            for start in range(len(value)):
                if value[start: start + 2] == '{':
                    end = value.find('}', start)
                    if end == -1:
                        break
                    token = value[start + 1: end]
                    if start + 3 > end:
                        break
                    if token.startswith('$'):
                        if ':' in token:
                            section, key = token.split(':', 1)
                            sec = section if section else current_section
                            if sec != current_section and sec not in sections:
                                raise KeyError(f"Section ${section} missing interpolation")
                            res = default_section.get(key) or (sections.get(sec, {}).get(key, None))
                            if res is None: raise KeyError(key)
                            if key in seen and rec_stack:
                                raise ValueError("Circular reference")
                            seen.add(key)
                            stack.append(key)
                            resolved = expand_tokens(str(res), stack)
                            stack.pop()
                            del seen[-1]
                            value = value[:start] + str(resolved) + value[end + 1:]
                            return expand_tokens(value, stack)
                        else:
                            key = token.lstrip('$')
                            res = default_section.get(key) or (sections.get(current_section, {}).get(key, None))
                            if res is None: raise KeyError(key)
                            value = value.replace(f'${{key}}', str(res))
        return expand_tokens(value, [])
    
    for line_idx, line in enumerate(text.splitlines()):
        if empty_or_comment(line) or not line.strip():
            continue

        if current_section is None:
            raise ValueError("Key before section or malformed line")

        sec = extract_section(line)
        if sec:
            if pending_key:
                sections.setdefault(sec, {})[pending_key] = '\n'.join(current_value_lines)
                pending_key = None
                current_value_lines = []
            current_section = sec
            if current_section not in sections:
                sections[current_section] = {}
            continue

        eq_pos = find_last_split(line)
        if eq_pos[0] is None:
            raise ValueError("Key before value or delimiter missing")

        key, delim = eq_pos[1], eq_pos[0]
        val = line[key + delim + 1:]
        value = process_value(val)

        if pending_key is None:
            pending_key = key
            current_value_lines = [value]
        else:
            current_value_lines.append(value)

        if line.strip().endswith('\\'):
            continue

        sections[current_section][pending_key] = '\n'.join(current_value_lines)
        pending_key = None
        current_value_lines = []

    def process_interpolated(sections_dict: dict) -> None:
        for sec, sec_dict in sections_dict.items():
            interpolate_recursive(sec, sec_dict, default_section)
    
    def interpolate_recursive(sec: str, subdict: dict[str, str], parent: dict[str, str],
                           parent_seen: set[str] | None = None) -> None:
        subdict = dict(subdict)
        parent_seen = parent_seen if parent_seen is not None else set()
        if sec in parent_seen:
            raise ValueError("Circular reference")
        parent_seen.add(sec)
        
        for k, v in subdict.items():
            def replace_token(tok: str) -> None:
                if not tok.startswith('$'):
                    return
                if tok.startswith('${'):
                    nonlocal subdict[k]
                    suffix = tok[2:-1]
                    val = None
                    section_suffix = None
                    if ':' in suffix:
                        section_idx = suffix.index(':')
                        section_name = suffix[:section_idx].encode().decode().encode().decode()
                        key_name = suffix[section_idx + 1:]
                        section_dict = sections.get(section_name, {})
                        val = section_dict.get(key_name, default_section.get(key_name))
                        if val is None: raise KeyError(f"{section_name}.{key_name}")
                    else:
                        key_name = suffix
                        section_dict = sections.get(section_name, {})
                        val = section_dict.get(key_name, default_section.get(key_name))
                        if val is None: raise KeyError(f"section {key_name}")
                    subdict[k] = subdict[k].replace("${" + suffix + "}", "§INTER§")
                raise KeyError(f"Invalid interpolation token: {tok}")
            
            for i, c in enumerate(v):
                if v[i] == '{':
                    token = v[i+1:i+3] if i+3 <= len(v) and v[i+1:i+3] == '{$' else v[i:i+2]
                    if token == '{$':
                        idx = i + 2
                        end = v.find('}', idx) if idx + 2 <= len(v) else len(v)
                        token_str = v[idx - 1:end + 1]
                        section_name = token_str.split(':')[0] if ':' in token_str else current_section
                        key_name = token_str.split(':')[1]
                        res = sections.get(section_name, {}).get(key_name)
                        res = res or default_section.get(key_name)
                        if res is None: raise KeyError
                        v = v[:idx - 1] + str(res) + v[end + 1:]
                        continue
                raise ValueError("Interpolation error")
            pass
        
        for k, v in subdict.items():
            replace_token(f'${k}')
            if '$' in v:
                raise ValueError("Unprocessed interpolation")
        for k, v in subdict.items():
            for i in range(10):
                new_v = v.replace("${", "${_interp")
                if new_v == v:
                    break
                v = new_v
            def eval_interp(v: str) -> str:
                def repl(token: str) -> str:
                    if token.startswith('${'):
                        sep = token.find(':')
                        sn = token[2] if sep > 0 else '' if sep != 2 else ''
                        kn = token[2:sep].encode().decode()[:2]
                        sec = sn if sn else current_section
                        val = sections.get(sec, {}).get(kn, default_section.get(kn))
                        if val is None: raise KeyError
                    return str(val)
                return v.replace("${", "{:interp".format(resourceName))
    
    for sec in sections:
        if sec != 'DEFAULT':
            interpolate_recursive(sec, sections[sec], default_section)
    return sections if 'DEFAULT' not in sections else {k: v for k, v in sections.items() if k != 'DEFAULT'}
```
Coding: TTL LRU (hard) Didn't work
Coding: Parse duration Worked
```python
import re

def parse_duration(s: str) -> int:
    # Check for empty or blank strings
    if not s.strip():
        raise ValueError("ValueError")
    
    # Define Regex Patterns
    # 1. Matches a valid part: digits followed by a unit (h/m/s)
    #    Units are case-insensitive (covered by character class [hmsHMS])
    #    Note: '1 h' (with space) will not be matched as a single part.
    #    Depending on strict interpretation, '1 h' is often invalid as a "number followed by unit".
    #    However, if valid, the regex \d+\s*[hmsHMS] would be needed. 
    #    Given standard formats (1h30m, 45s), we assume contiguity or parts are separated by whitespace.
    #    So valid parts must be 'digits' + 'unit' without intervening characters.
    part_pattern = r'\d+[hmsHMS]'
    
    # Extract all valid parts
    parts = re.findall(part_pattern, s)
    
    # Check for leftover content (garbage, decimals, numbers without units, etc.)
    # re.sub removes all matches of the part_pattern. If anything remains (other than whitespace),
    # the string was malformed (e.g., '1 2h' -> '1' remains; '1.5h' -> '1.5h' remains; 'abc' -> 'abc' remains).
    remaining = re.sub(part_pattern, '', s)
    if remaining.strip():
        raise ValueError("Invalid format: number without unit, decimal, or unknown unit")
    
    # Mapping units to seconds and a order map (h < m < s)
    units_map = {'h': 3600, 'm': 60, 's': 1}
    unit_order = {'h': 0, 'm': 1, 's': 2}
    
    seen_units = set()
    last_order_idx = -1
    total_seconds = 0
    
    for part in parts:
        # Extract unit (case-insensitive)
        unit_char = part[-1].lower()
        
        # Validate unit existence (redundant with regex but safe)
        if unit_char not in units_map:
            raise ValueError("Unknown unit")
        
        current_order = unit_order[unit_char]
        
        # Check for out of order
        if current_order < last_order_idx:
            raise ValueError("Units out of order")
        
        # Check for repetition
        if unit_char in seen_units:
            raise ValueError("Repeated unit")
        
        seen_units.add(unit_char)
        last_order_idx = current_order
        
        # Parse number and add to total
        # part[:-1] is the number string (regex ensures exactly this structure)
        try:
            count = int(part[:-1])
        except ValueError:
            # Should not be reached given remaining checks, but for safety
            raise ValueError("Invalid number format")
            
        total_seconds += count * units_map[unit_char]
        
    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    # Handle empty list
    if not nums:
        return ""

    # Sort the numbers and remove duplicates
    sorted_nums = sorted(set(nums))
    
    # If list is empty after processing (e.g., all duplicates or initially empty)
    if not sorted_nums:
        return ""

    parts = []
    current_group = [sorted_nums[0]]
    
    for num in sorted_nums[1:]:
        # Check if the current number is consecutive to the last number in the group
        if num == current_group[-1] + 1:
            current_group.append(num)
        else:
            # The consecutive run has ended
            start, end = current_group[0], current_group[-1]
            run_length = len(current_group)
            
            if run_length >= 3:
                # Run of 3 or more consecutive integers
                parts.append(f"{start}..{end}")
            else:
                # Run of 1 or 2 integers, list individually
                for n in current_group:
                    parts.append(str(n))
            
            # Start a new group with the current number
            current_group = [num]
    
    # Process the final group
    start, end = current_group[0], current_group[-1]
    run_length = len(current_group)
    if run_length >= 3:
        parts.append(f"{start}..{end}")
    else:
        for n in current_group:
            parts.append(str(n))
            
    return ",".join(parts)
```
Coding: Top customers Worked
```python
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    """
    Returns the top n customers based on transaction amounts.
    """
    # Dict to aggregate amounts per customer (using lowercase names as keys)
    sums = {}
    # Dict to store the canonical name for each lowercase key (from first valid row)
    names = {}
    
    for row in rows:
        # Get and trim customer name
        customer = row.get("customer", "")
        if not isinstance(customer, str):
            continue
        name = customer.strip()
        
        # Skip rows with empty customer name
        if not name:
            continue
            
        # Get amount string
        amount_str = row.get("amount")
        if not amount_str or not isinstance(amount_str, str):
            continue
            
        # Try to parse amount
        # Formats like "$1,234.50", "1234.5", "-$5.00", "$1,000"
        # We remove '$' and ',' (thousands separator)
        try:
            clean_amount = amount_str.replace("$", "").replace(",", "")
            # float() handles string like "-5.00" or "5.00" 
            # It handles leading/trailing whitespace implicitly, 
            # but internal spaces/commas would raise ValueError.
            amount = float(clean_amount)
        except (ValueError, TypeError):
            # Skip rows with unparseable amounts
            continue
            
        # Normalize name for grouping (case-insensitive)
        key = name.lower()
        
        # Store canonical name only for the first valid row
        if key not in names:
            names[key] = name
        
        # Aggregate amount
        if key not in sums:
            sums[key] = 0.0
        sums[key] += amount
        
    # Prepare the result list
    customers = []
    for key, total in sums.items():
        canonical_name = names[key]
        rounded_total = round(total, 2)
        customers.append((canonical_name, rounded_total))
        
    # Sort by total descending, then name ascending
    customers.sort(key=lambda x: (-x[1], x[0]))
    
    # Return top n
    return customers[:n]
```
Coding: Slugify Didn't work
Coding: Token bucket Didn't work
```python
from __future__ import annotations

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.tokens = capacity
        self.last_refill = clock()

    @property
    def tokens(self) -> float:
        now = self.clock()
        elapsed = now - self.last_refill
        self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
        self.last_refill = now
        return self.tokens

    def allow(self, cost: float = 1.0) -> bool:
        self.tokens = min(self.capacity, self.tokens + (self.clock() - self.last_refill) * self.refill_per_sec)
        self.last_refill = self.clock()
        return self.tokens >= cost
```
Decisions: Refund window (hard) 100% right


{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.0,
    "store_credit": 1.0,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0.0,
    "1": 1.0,
    "2": 0.0,
    "3": 0.0
  },
  "page": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Tool followup (hard) 100% right


{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 1.0,
    "email": 0.0,
    "none": 0.0
  },
  "confirm": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Legit security alert (hard) 100% right


{
  "phishing": {
    "true": 0.0,
    "false": 1.0
  },
  "action_needed": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Meeting slot (hard) 100% right


{
  "slot": {
    "A": 0.0,
    "B": 1.0,
    "C": 0.0,
    "D": 0.0
  },
  "raj_last": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed (hard) 100% right


{
  "hardware": {
    "true": 0.99,
    "false": 0.01
  },
  "support": {
    "true": 0.99,
    "false": 0.01
  }
}
Decisions: Support checkout down 100% right


{
  "department": {
    "billing": 0.05,
    "technical": 0.95,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.05,
    "3": 0.95
  },
  "outage": {
    "true": 0.98,
    "false": 0.02
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.9,
    "technical": 0.05,
    "account": 0.04,
    "shipping": 0.01,
    "sales": 0.0
  },
  "refund": {
    "true": 1.0,
    "false": 0.0
  },
  "tone": {
    "calm": 1.0,
    "frustrated": 0.0
  }
}
Decisions: Moderation doxxing 100% right


{
  "policy": {
    "none": 0.0,
    "harassment": 1.0,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.0
  },
  "personal_info": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Route calendar 100% right


{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 1.0,
    "email": 0.0,
    "none": 0.0
  },
  "confirm": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Doc invoice missing due 100% right


{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Phishing paypal 100% right


{
  "phishing": {
    "true": 0.99,
    "false": 0.01
  },
  "risk": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.02,
    "3": 0.98
  }
}
Decisions: Pii ssn email 100% right
{
    "data_kind": {
        "none": 0.0,
        "contact": 0.0,
        "financial": 0.0,
        "government_id": 1.0,
        "health": 0.0
    },
    "sensitive": {
        "true": 1.0,
        "false": 0.0
    }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.05,
    "neutral": 0.05,
    "negative": 0.90
  },
  "defect": {
    "true": 0.99,
    "false": 0.01
  },
  "recommend": {
    "true": 0.01,
    "false": 0.99
  }
}
Documents: Saas escalator (hard) 60% right


{
  "year2_price_per_seat_month": 47.25,
  "year3_price_per_seat_month": 47.25,
  "year1_invoice": 58320.0,
  "year2_invoice": 61074.0,
  "addon_months_billed": 4,
  "addon_invoice": 25704.0,
  "year3_invoice": 134946.0,
  "year3_discount_percent": 15,
  "total_contract_value": 280044.0,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 100% right


{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.6
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.0,
  "total_reimbursable_usd": 2159.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right


{
  "tenants": [
    "Marcus Lin",
    "Sofia Lin"
  ],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 4700.00
}
Documents: Ticket SLA 92% right


{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": [
    "billing_address",
    "invoice_pdf"
  ],
  "affected_orders": [
    "SO-99812",
    "SO-99820",
    "SO-99827"
  ],
  "priority": "P2",
  "sla_due_local": "2025-09-15T15:30",
  "sla_due_utc": "2025-09-15T20:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 100% right


{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": [
    "East"
  ],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 9.7B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.