Review · updated OCT 11

Qwen3 VL 8B Instruct review: not one we'd recommend right now

It scored 39 out of 100, #40 of 56. It solved 11 of 30 coding jobs and scored 41 on reading documents. Runs on an 8 GB graphics card or a Mac with 16 GB.

The short version
  • Qwen3 VL 8B Instruct is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's not one we'd recommend right now: 39 out of 100, #40 of 56.
  • It solved 11 of 30 coding jobs and scored 41 on reading documents. On our hardest tasks it scored 25.
  • Runs on an 8 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen3 VL 8B Instruct got 11 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen3 VL 8B Instruct scored 41; the best model scored 100.

TestScorePublic questionsSecret questions
Coding372939
Reading documents414740
Decisions798277

On the 18 hardest tasks (included in the scores above) it scored 25. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen3 VL 8B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBYesQ4
Graphics card, 12 GBYesQ8
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ6
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Qwen3 VL 8B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen3 VL 8B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Process each line
    for i, line in enumerate(lines):
        line = line.strip()
        if not line:
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
                current_section = 'DEFAULT'
                # We don't store DEFAULT as a section, but collect its keys
                default_section = {}
                continue
            else:
                current_section = section_name
                sections[section_name] = {}
                continue

        # If we're not in a section yet, and this is not a section header, it's invalid
        if current_section is None:
            raise ValueError("Section header expected before key-value line")

        # Parse key-value pair
        if '=' in line:
            split_idx = line.index('=')
        elif ':' in line:
            split_idx = line.index(':')
        else:
            raise ValueError("Invalid key-value line")

        key_part = line[:split_idx].strip()
        value_part = line[split_idx + 1:].strip()

        # Handle quoted value
        if value_part.startswith('"') and value_part.endswith('"'):
            value_part = value_part[1:-1]
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)
        else:
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)

        # Process key (strip and lowercase)
        key = key_part.lower()

        # Store the value
        if current_section == 'DEFAULT':
            default_section[key] = value_part
        else:
            sections[current_section][key] = value_part

    # Apply DEFAULT section to all other sections
    for section_name, section_data in sections.items():
        if section_name != 'DEFAULT':
            for key, value in default_section.items():
                if key not in section_data:
                    section_data[key] = value

    return sections
```

Wait — I need to implement interpolation, which is more complex. Let me rewrite the solution to handle interpolation properly, including recursive interpolation and error handling for missing keys or cycles.

```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Parse lines
    i = 0
    while i < len(lines):
        line = lines[i].strip()
        if not line:
            i += 1
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            i += 1
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
                current_section = 'DEFAULT'
                default_section = {}
                i += 1
                continue
            else:
                current_section = section_name
                sections[section_name] = {}
                i += 1
                continue

        # If we're not in a section yet, it's invalid
        if current_section is None:
            raise ValueError("Section header expected before key-value line")

        # Parse key-value pair
        if '=' in line:
            split_idx = line.index('=')
        elif ':' in line:
            split_idx = line.index(':')
        else:
            raise ValueError("Invalid key-value line")

        key_part = line[:split_idx].strip()
        value_part = line[split_idx + 1:].strip()

        # Handle quoted value
        if value_part.startswith('"') and value_part.endswith('"'):
            value_part = value_part[1:-1]
        else:
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)

        # Process key (strip and lowercase)
        key = key_part.lower()

        # Store the value
        if current_section == 'DEFAULT':
            default_section[key] = value_part
        else:
            sections[current_section][key] = value_part

        i += 1

    # Apply DEFAULT section to all other sections
    for section_name, section_data in sections.items():
        if section_name != 'DEFAULT':
            for key, value in default_section.items():
                if key not in section_data:
                    section_data[key] = value

    return sections
```

Wait — I still haven't implemented interpolation, which is required. Let me fix that.

```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Parse lines
    i = 0
    while i < len(lines):
        line = lines[i].strip()
        if not line:
            i += 1
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            i += 1
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
                current_section = 'DEFAULT'
                default_section = {}
                i += 1
                continue
            else:
                current_section = section_name
                sections[section_name] = {}
                i += 1
                continue

        # If we're not in a section yet, it's invalid
        if current_section is None:
            raise ValueError("Section header expected before key-value line")

        # Parse key-value pair
        if '=' in line:
            split_idx = line.index('=')
        elif ':' in line:
            split_idx = line.index(':')
        else:
            raise ValueError("Invalid key-value line")

        key_part = line[:split_idx].strip()
        value_part = line[split_idx + 1:].strip()

        # Handle quoted value
        if value_part.startswith('"') and value_part.endswith('"'):
            value_part = value_part[1:-1]
        else:
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)

        # Process key (strip and lowercase)
        key = key_part.lower()

        # Store the value
        if current_section == 'DEFAULT':
            default_section[key] = value_part
        else:
            sections[current_section][key] = value_part

        i += 1

    # Apply DEFAULT section to all other sections
    for section_name, section_data in sections.items():
        if section_name != 'DEFAULT':
            for key, value in default_section.items():
                if key not in section_data:
                    section_data[key] = value

    return sections
```

I realize I'm still missing interpolation. Let me implement it properly.

```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Parse lines
    i = 0
    while i < len(lines):
        line = lines[i].strip()
        if not line:
            i += 1
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            i += 1
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
                current_section = 'DEFAULT'
                default_section = {}
                i += 1
                continue
            else:
                current_section = section_name
                sections[section_name] = {}
                i += 1
                continue

        # If we're not in a section yet, it's invalid
        if current_section is None:
            raise ValueError("Section header expected before key-value line")

        # Parse key-value pair
        if '=' in line:
            split_idx = line.index('=')
        elif ':' in line:
            split_idx = line.index(':')
        else:
            raise ValueError("Invalid key-value line")

        key_part = line[:split_idx].strip()
        value_part = line[split_idx + 1:].strip()

        # Handle quoted value
        if value_part.startswith('"') and value_part.endswith('"'):
            value_part = value_part[1:-1]
        else:
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)

        # Process key (strip and lowercase)
        key = key_part.lower()

        # Store the value
        if current_section == 'DEFAULT':
            default_section[key] = value_part
        else:
            sections[current_section][key] = value_part

        i += 1

    # Apply DEFAULT section to all other sections
    for section_name, section_data in sections.items():
        if section_name != 'DEFAULT':
            for key, value in default_section.items():
                if key not in section_data:
                    section_data[key] = value

    return sections
```

I'm still not handling interpolation. Let me rewrite this properly.

```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Parse lines
    i = 0
    while i < len(lines):
        line = lines[i].strip()
        if not line:
            i += 1
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            i += 1
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
                current_section = 'DEFAULT'
                default_section = {}
                i += 1
                continue
            else:
                current_section = section_name
                sections[section_name] = {}
                i += 1
                continue

        # If we're not in a section yet, it's invalid
        if current_section is None:
            raise ValueError("Section header expected before key-value line")

        # Parse key-value pair
        if '=' in line:
            split_idx = line.index('=')
        elif ':' in line:
            split_idx = line.index(':')
        else:
            raise ValueError("Invalid key-value line")

        key_part = line[:split_idx].strip()
        value_part = line[split_idx + 1:].strip()

        # Handle quoted value
        if value_part.startswith('"') and value_part.endswith('"'):
            value_part = value_part[1:-1]
        else:
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)

        # Process key (strip and lowercase)
        key = key_part.lower()

        # Store the value
        if current_section == 'DEFAULT':
            default_section[key] = value_part
        else:
            sections[current_section][key] = value_part

        i += 1

    # Apply DEFAULT section to all other sections
    for section_name, section_data in sections.items():
        if section_name != 'DEFAULT':
            for key, value in default_section.items():
                if key not in section_data:
                    section_data[key] = value

    return sections
```

I'm realizing I'm not implementing interpolation at all. Let me start over with a proper implementation.

```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Parse lines
    i = 0
    while i < len(lines):
        line = lines[i].strip()
        if not line:
            i += 1
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            i += 1
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
                current_section = 'DEFAULT'
                default_section = {}
                i += 1
                continue
            else:
                current_section = section_name
                sections[section_name] = {}
                i += 1
                continue

        # If we're not in a section yet, it's invalid
        if current_section is None:
            raise ValueError("Section header expected before key-value line")

        # Parse key-value pair
        if '=' in line:
            split_idx = line.index('=')
        elif ':' in line:
            split_idx = line.index(':')
        else:
            raise ValueError("Invalid key-value line")

        key_part = line[:split_idx].strip()
        value_part = line[split_idx + 1:].strip()

        # Handle quoted value
        if value_part.startswith('"') and value_part.endswith('"'):
            value_part = value_part[1:-1]
        else:
            # Handle continuation lines
            continuation_lines = []
            j = i + 1
            while j < len(lines):
                next_line = lines[j].strip()
                if not next_line:
                    j += 1
                    continue
                if next_line[0] in ';#':
                    break
                if next_line.startswith(' '):
                    continuation_lines.append(next_line[1:].strip())
                else:
                    break
                j += 1
            if continuation_lines:
                value_part += '\n' + '\n'.join(continuation_lines)

        # Process key (strip and lowercase)
        key = key_part.lower()

        # Store the value
        if current_section == 'DEFAULT':
            default_section[key] = value_part
        else:
            sections[current_section][key] = value_part

        i += 1

    # Apply DEFAULT section to all other sections
    for section_name, section_data in sections.items():
        if section_name != 'DEFAULT':
            for key, value in default_section.items():
                if key not in section_data:
                    section_data[key] = value

    return sections
```

I'm still not implementing interpolation. Let me implement it properly.

```python
import re
from typing import Dict, List, Optional, Tuple, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.strip().splitlines()
    if not lines:
        return {}

    sections: Dict[str, Dict[str, str]] = {}
    current_section = None
    default_section = {}

    # Parse lines
    i = 0
    while i < len(lines):
        line = lines[i].strip()
        if not line:
            i += 1
            continue

        # Check for comment or empty line
        if line[0] in ';#':
            i += 1
            continue

        # Check for section header
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if not section_name:
                raise ValueError("Empty section name")
            if section_name.upper() == 'DEFAULT':
  
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from typing import Any, Callable, Optional

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()  # key -> (value, expiry_time)
        self.stats = {
            "hits": 0,
            "misses": 0,
            "evictions": 0,
            "expirations": 0
        }

    def put(self, key: Any, value: Any) -> None:
        now = self.clock()
        expiry = now + self.ttl

        # Remove expired entries first
        expired_keys = []
        for k, (v, exp) in self.cache.items():
            if now >= exp:
                expired_keys.append(k)
        for k in expired_keys:
            self.cache.pop(k)
            self.stats["expirations"] += 1

        # Check if we need to evict
        if len(self.cache) >= self.capacity and self.capacity > 0:
            # Evict the least recently used
            lru_key = next(iter(self.cache))
            self.cache.pop(lru_key)
            self.stats["evictions"] += 1

        # Insert or update
        self.cache[key] = (value, expiry)

    def get(self, key: Any, default: Any = None) -> Any:
        now = self.clock()
        if key not in self.cache:
            self.stats["misses"] += 1
            return default

        value, expiry = self.cache[key]
        if now >= expiry:
            # Entry is expired
            self.cache.pop(key)
            self.stats["expirations"] += 1
            self.stats["misses"] += 1
            return default

        # Mark as most recently used
        self.cache.move_to_end(key)
        self.stats["hits"] += 1
        return value

    def __len__(self) -> int:
        now = self.clock()
        count = 0
        for k, (v, exp) in self.cache.items():
            if now < exp:
                count += 1
        return count

    def stats(self) -> dict:
        return self.stats.copy()
```
Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    current = ""
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
        i += 1
    
    # Validate order and uniqueness
    seen = set()
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        if unit == "h":
            continue
        if unit == "m":
            continue
        if unit == "s":
            continue
        # Check order: h, then m, then s
        if unit == "m" and "h" not in seen:
            raise ValueError("Units out of order")
        if unit == "s" and "h" not in seen and "m" not in seen:
            raise ValueError("Units out of order")
    
    # Validate order: h, m, s
    units_seen = []
    for num, unit in parts:
        units_seen.append(unit)
    if units_seen != ["h", "m", "s"]:
        if "h" not in units_seen or "m" not in units_seen or "s" not in units_seen:
            raise ValueError("Units out of order")
        # Check order
        if units_seen[0] != "h" or units_seen[1] != "m" or units_seen[2] != "s":
            raise ValueError("Units out of order")
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

Wait, I need to fix the logic for checking order and uniqueness. Let me revise:

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    expected_order = ["h", "m", "s"]
    current_index = 0
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if unit != expected_order[current_index]:
            raise ValueError("Units out of order")
        current_index += 1
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

Actually, I need to reconsider the unit order validation. The requirement says "Units must appear in the order h, m, s and each at most once." This means the units must appear in that exact order, but not necessarily all three. So I need to check that if "h" appears, then "m" can only appear after it, and "s" can only appear after "m", etc.

Let me fix this:

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    expected_order = ["h", "m", "s"]
    current_index = 0
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if current_index < len(expected_order) and unit != expected_order[current_index]:
            raise ValueError("Units out of order")
        current_index += 1
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

Actually, I think I'm overcomplicating this. Let me think again. The requirement says "Units must appear in the order h, m, s and each at most once." This means the units must appear in that exact order, but not necessarily all three. So I need to check that the first unit is "h" (if present), then "m" (if present), then "s" (if present).

Let me fix this:

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    expected_order = ["h", "m", "s"]
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if unit not in expected_order:
            raise ValueError("Unknown unit")
        # Check if unit appears in correct order
        if unit == "h" and "h" not in seen:
            continue
        if unit == "m" and "h" not in seen:
            raise ValueError("Units out of order")
        if unit == "s" and "h" not in seen and "m" not in seen:
            raise ValueError("Units out of order")
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

Actually, I think I'm still not handling the order correctly. Let me simplify and make it correct:

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if unit == "h":
            continue
        if unit == "m":
            if "h" not in seen:
                raise ValueError("Units out of order")
        if unit == "s":
            if "h" not in seen or "m" not in seen:
                raise ValueError("Units out of order")
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

Actually, I think I'm overthinking this. Let me write a cleaner version:

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if unit == "h":
            continue
        if unit == "m":
            if "h" not in seen:
                raise ValueError("Units out of order")
        if unit == "s":
            if "h" not in seen or "m" not in seen:
                raise ValueError("Units out of order")
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

Actually, I think I'm still not handling the order correctly. Let me try a different approach:

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if unit == "h":
            continue
        if unit == "m":
            if "h" not in seen:
                raise ValueError("Units out of order")
        if unit == "s":
            if "h" not in seen or "m" not in seen:
                raise ValueError("Units out of order")
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```

I think this is correct now. Let me test with some examples mentally:
- "1h30m" -> h then m -> valid
- "2H" -> H -> h -> valid
- "1h 5m 10s" -> h, then m, then s -> valid
- "90m" -> m -> but "h" not seen -> invalid (but according to requirement, if "m" appears, "h" must have appeared before it, so this should be invalid)
- "1h 2m 3s" -> valid
- "1s 2m" -> s before m -> invalid

Yes, this seems correct.

```python
def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")
    
    s = s.strip()
    parts = []
    i = 0
    
    # Split by whitespace to handle multiple parts
    tokens = s.split()
    for token in tokens:
        if not token:
            continue
        # Check if token has a unit
        if len(token) == 0:
            continue
        # Check for unit at the end
        unit = token[-1].lower()
        if unit not in "hms":
            raise ValueError("Unknown unit")
        # Check if the rest is a valid number
        num_str = token[:-1]
        if not num_str or not num_str.isdigit():
            raise ValueError("Number without unit")
        num = int(num_str)
        parts.append((num, unit))
    
    # Validate order and uniqueness
    seen = set()
    for num, unit in parts:
        if unit in seen:
            raise ValueError("Repeated unit")
        seen.add(unit)
        # Check if unit is in expected order
        if unit == "h":
            continue
        if unit == "m":
            if "h" not in seen:
                raise ValueError("Units out of order")
        if unit == "s":
            if "h" not in seen or "m" not in seen:
                raise ValueError("Units out of order")
    
    # Calculate total seconds
    total_seconds = 0
    for num, unit in parts:
        if unit == "h":
            total_seconds += num * 3600
        elif unit == "m":
            total_seconds += num * 60
        elif unit == "s":
            total_seconds += num
    
    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    
    nums = sorted(set(nums))
    result = []
    i = 0
    
    while i < len(nums):
        start = nums[i]
        # Look for consecutive numbers
        while i + 1 < len(nums) and nums[i + 1] == nums[i] + 1:
            i += 1
        end = nums[i]
        
        # If run is 3 or more consecutive numbers, use ".."
        if end - start + 1 >= 3:
            result.append(f"{start}..{end}")
        else:
            # Otherwise list each number individually
            for j in range(start, end + 1):
                result.append(str(j))
        
        i += 1
    
    return ",".join(result)
```
Coding: Top customers Worked
```python
import re
from typing import List, Dict, Tuple

def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    customer_totals = {}
    customer_first_seen = {}
    
    for row in rows:
        customer = row.get("customer", "").strip()
        amount_str = row.get("amount", "")
        
        if not customer or not amount_str:
            continue
            
        # Remove currency symbol and commas
        amount_str = re.sub(r'[$,]', '', amount_str)
        
        try:
            amount = float(amount_str)
        except ValueError:
            continue
            
        # Case-insensitive matching for customer name
        customer_lower = customer.lower()
        
        # Use first valid row for each customer
        if customer_lower not in customer_first_seen:
            customer_first_seen[customer_lower] = customer
            customer_totals[customer_lower] = 0.0
            
        customer_totals[customer_lower] += amount
    
    # Sort by total descending, then name ascending
    sorted_customers = sorted(
        [(customer_first_seen[k], customer_totals[k]) for k in customer_totals.keys()],
        key=lambda x: (-x[1], x[0])
    )
    
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    # Step 1: Transliterate accented characters to ASCII and drop non-ASCII
    normalized = unicodedata.normalize('NFKD', title)
    ascii_only = ''.join(c for c in normalized if ord(c) < 128)
    
    # Step 2: Lowercase
    lowercase = ascii_only.lower()
    
    # Step 3: Replace runs of non-alphanumeric with single "-", strip leading/trailing "-"
    slug = ''.join('-' if not c.isalnum() else c for c in lowercase)
    slug = slug.strip('-')
    
    # Step 4: Shorten if needed without cutting words
    if not slug:
        return ""
    
    # Split by hyphens to get words
    words = slug.split('-')
    
    # If max_len is 0, return empty string
    if max_len <= 0:
        return ""
    
    # If the entire slug is within max_len, return it
    if len(slug) <= max_len:
        return slug
    
    # Try to build the shortest prefix of whole words that fits within max_len
    result = []
    current_length = 0
    
    for word in words:
        if current_length > 0:
            current_length += 1  # Add one for the hyphen separator
        if len(word) + current_length <= max_len:
            result.append(word)
            current_length += len(word)
        else:
            # If even the first word is too long, hard-cut it
            if not result:
                if len(word) <= max_len:
                    result.append(word[:max_len])
                else:
                    result.append(word[:max_len])
            break
    
    # Join the words back with hyphens
    if not result:
        return ""
    
    final_slug = '-'.join(result)
    
    # Ensure it doesn't end with "-"
    if final_slug and final_slug[-1] == '-':
        final_slug = final_slug[:-1]
    
    return final_slug
```
Coding: Token bucket Didn't work
```python
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.tokens = capacity
        self.last_refill_time = clock()

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self.tokens

    def _refill(self):
        current_time = self.clock()
        time_elapsed = current_time - self.last_refill_time
        tokens_to_add = self.refill_per_sec * time_elapsed
        self.tokens = min(self.capacity, self.tokens + tokens_to_add)
        self.last_refill_time = current_time
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "partial_refund": 1.0
  },
  "defective": {
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 0% right
{
  "severity": {
    "0": 0.0,
    "1": 0.2,
    "2": 0.7,
    "3": 0.1
  },
  "page": {
    "true": 0.7,
    "false": 0.3
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.05,
    "false": 0.95
  }
}
Decisions: Meeting slot (hard) 0% right
{
  "slot": {
    "A": 0.0,
    "B": 0.25,
    "C": 0.75,
    "D": 0.0
  },
  "raj_last": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 0.95,
    "false": 0.05
  },
  "support": {
    "true": 0.98,
    "false": 0.02
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "technical": 0.95,
    "billing": 0.03,
    "account": 0.01,
    "shipping": 0.005,
    "sales": 0.005
  },
  "urgency": {
    "3": 0.98,
    "2": 0.01,
    "1": 0.005,
    "0": 0.005
  },
  "outage": {
    "true": 0.99,
    "false": 0.01
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.95,
    "technical": 0.02,
    "account": 0.01,
    "shipping": 0.01,
    "sales": 0.01
  },
  "refund": {
    "true": 0.98,
    "false": 0.02
  },
  "tone": {
    "frustrated": 0.1,
    "calm": 0.9
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "harassment": 0.95,
    "none": 0.05
  },
  "personal_info": {
    "true": 0.98,
    "false": 0.02
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 0.95,
    "false": 0.05
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 0.99,
    "resume": 0.00,
    "contract": 0.00,
    "bank_statement": 0.00,
    "other": 0.01
  },
  "missing_due_date": {
    "true": 0.95,
    "false": 0.05
  }
}
Decisions: Phishing paypal 100% right
{
  "phishing": {"true": 0.99, "false": 0.01},
  "risk": {"0": 0.0, "1": 0.0, "2": 0.0, "3": 0.99}
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "government_id": 1.0
  },
  "sensitive": {
    "true": 1.0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "negative": 0.95,
    "neutral": 0.05,
    "positive": 0.0
  },
  "defect": {
    "true": 0.98,
    "false": 0.02
  },
  "recommend": {
    "true": 0.0,
    "false": 1.0
  }
}
Documents: Saas escalator (hard) 10% right
{
  "year2_price_per_seat_month": 46.80,
  "year3_price_per_seat_month": 46.32,
  "year1_invoice": 54000.00,
  "year2_invoice": 56160.00,
  "addon_months_billed": 12,
  "addon_invoice": 7128.00,
  "year3_invoice": 56160.00,
  "year3_discount_percent": 15,
  "total_contract_value": 173548.00,
  "contract_end_date": "2026-02-28"
}
Documents: Expense thread 67% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.6
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 33.33
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 196.32
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 400.0
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 37.0
    }
  ],
  "rejected_item_count": 2,
  "per_diem_days": 2,
  "per_diem_usd": 65,
  "total_reimbursable_usd": 1950.45,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 67% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2237.0,
  "late_fee_from_2025_06_01": 107.5,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2587.0,
  "move_in_payment": 4300
}
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-12T17:30",
  "sla_due_utc": "2025-09-12T22:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 10% right
{
  "q3_total_usd": 15846,
  "q2_total_usd": 14464,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 17.5,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["Central", "East"],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 4191
}

Size: 8.8B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.