Review · updated OCT 11

Qwen3 VL 30B A3B Thinking review: reads documents well, weaker at code

It scored 77 out of 100, #17 of 56. It solved 18 of 30 coding jobs and scored 93 on reading documents. Runs on a 24 GB graphics card or a Mac with 32 GB.

The short version
  • Qwen3 VL 30B A3B Thinking is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's reads documents well, weaker at code: 77 out of 100, #17 of 56.
  • It solved 18 of 30 coding jobs and scored 93 on reading documents. On our hardest tasks it scored 51.
  • Runs on a 24 GB graphics card or a Mac with 32 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen3 VL 30B A3B Thinking got 18 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen3 VL 30B A3B Thinking scored 93; the best model scored 100.

TestScorePublic questionsSecret questions
Coding605761
Reading documents939892
Decisions969696

On the 18 hardest tasks (included in the scores above) it scored 51. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen3 VL 30B A3B Thinking in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBYesQ5
Graphics card, 32 GBYesQ6
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBYesQ4
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Qwen3 VL 30B A3B Thinking GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen3 VL 30B A3B Thinking wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    sections = {}
    default_section = {}
    current_section = None
    current_key = None
    current_value = None

    def split_key_value(line):
        index_eq = line.find('=')
        index_colon = line.find(':')
        if index_eq == -1 and index_colon == -1:
            raise ValueError("Invalid key-value line: no '=' or ':' found")
        if index_eq == -1:
            split_index = index_colon
        elif index_colon == -1:
            split_index = index_eq
        else:
            split_index = min(index_eq, index_colon)
        key = line[:split_index].strip()
        value = line[split_index+1:].strip()
        return key, value

    def process_value(value):
        value = value.strip()
        if len(value) >= 2 and value[0] == '"' and value[-1] == '"':
            value = value[1:-1]
        else:
            match = re.search(r'\s+;|\s+#', value)
            if match:
                value = value[:match.start()].strip()
        return value

    def resolve_references(value, section, sections, visited):
        pattern = r'\$\{([^}]+)\}'
        matches = re.findall(pattern, value)
        if not matches:
            return value
        def replace_match(match):
            ref = match.group(1)
            if ':' in ref:
                sec_name, key_name = ref.split(':', 1)
            else:
                sec_name = section
                key_name = ref
            if sec_name not in sections:
                raise KeyError(f"Section {sec_name} not found")
            if key_name not in sections[sec_name]:
                raise KeyError(f"Key {key_name} not found in section {sec_name}")
            if (sec_name, key_name) in visited:
                raise ValueError("Cyclic reference")
            visited.add((sec_name, key_name))
            resolved_val = resolve_references(sections[sec_name][key_name], sec_name, sections, visited)
            visited.remove((sec_name, key_name))
            return resolved_val
        return re.sub(pattern, replace_match, value)

    lines = text.splitlines()
    for line in lines:
        line = line.strip()
        if not line:
            continue
        if line[0] in ';#':
            continue
        if line[0] == '[' and line[-1] == ']':
            section_name = line[1:-1].strip()
            current_section = section_name
            if current_section == 'DEFAULT':
                default_section = {}
            else:
                sections[current_section] = {}
        elif line[0] in ' \t':
            if current_key is None:
                raise ValueError("Continuation line without key")
            current_value += '\n' + line.strip()
        else:
            if current_section is None:
                raise ValueError("Key line before section header")
            key, value = split_key_value(line)
            key = key.strip().lower()
            value = process_value(value)
            if current_section == 'DEFAULT':
                default_section[key] = value
            else:
                sections[current_section][key] = value
            current_key = key
            current_value = value

    sections['DEFAULT'] = default_section
    for section in sections:
        for key in list(sections[section].keys()):
            sections[section][key] = resolve_references(sections[section][key], section, sections, set())

    result = {}
    for section in sections:
        if section == 'DEFAULT':
            continue
        merged = sections['DEFAULT'].copy()
        merged.update(sections[section])
        result[section] = merged

    return result
```
Coding: TTL LRU (hard) Didn't work
To solve this problem, we need to implement a TTL (Time-To-Live) cache that combines the Least Recently Used (LRU) eviction policy with expiration of entries based on a specified time-to-live value. The cache must efficiently handle insertions, lookups, and evictions while maintaining accurate statistics on operations.

### Approach
1. **Data Structures**:
   - **Dictionary (`cache`)**: Maps keys to tuples containing the value, expiry time, and a node reference for the LRU list.
   - **Doubly Linked List**: Maintains the order of entries based on recency (most recently used at the head, least recently used at the tail). This allows O(1) operations for moving nodes to the head (for updates) and removing the tail (for eviction).
   - **Node Class**: Represents elements in the doubly linked list, storing the key and pointers to the previous and next nodes.

2. **Key Operations**:
   - **`put(key, value)`**:
     - If the key exists, update its value, expiry time (current time + TTL), and move it to the head of the LRU list.
     - If the key is new and the cache is full (non-expired entries ≥ capacity), first remove all expired entries (counting as expirations), then evict the least recently used entry (counting as eviction) until there is space.
     - Add the new entry to the head of the LRU list.
   - **`get(key, default=None)`**:
     - If the key exists and is not expired, return the value and move it to the head of the LRU list (counting as a hit).
     - If the key exists but is expired, remove it (counting as expiration and miss).
     - If the key does not exist, return the default (counting as a miss).
   - **`__len__()`**: Returns the count of non-expired entries (size of the cache).
   - **`stats()`**: Returns a dictionary with counts of hits, misses, evictions, and expirations.

3. **Handling Expired Entries**:
   - During `put` operations when the cache is full, expired entries are removed proactively (scanning the entire cache) to free space.
   - During `get` operations, expired entries are removed immediately upon discovery.
   - Each expiration (whether during `put` or `get`) increments the expiration count.

4. **Special Case for Capacity 0**:
   - If capacity is 0, all operations (put, get) are ignored, and `get` always returns the default (counting as a miss).

### Solution Code
```python
class Node:
    __slots__ = ('key', 'prev', 'next')
    def __init__(self, key):
        self.key = key
        self.prev = None
        self.next = None

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock: callable):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = {}
        self.head = None
        self.tail = None
        self.size = 0
        self.stats = {
            'hits': 0,
            'misses': 0,
            'evictions': 0,
            'expirations': 0
        }
    
    def _remove_node(self, node):
        if node == self.head:
            self.head = node.next
        if node == self.tail:
            self.tail = node.prev
        if node.prev:
            node.prev.next = node.next
        if node.next:
            node.next.prev = node.prev
        node.prev = None
        node.next = None
    
    def _add_node(self, node):
        if self.head is None:
            self.head = node
            self.tail = node
        else:
            node.next = self.head
            self.head.prev = node
            self.head = node

    def put(self, key, value):
        if self.capacity == 0:
            return
            
        if key in self.cache:
            current_time = self.clock()
            new_expiry = current_time + self.ttl
            node = self.cache[key][2]
            self._remove_node(node)
            self.cache[key] = (value, new_expiry, node)
            self._add_node(node)
            return
            
        if self.size >= self.capacity:
            expired_entries = []
            for k, (v, exp, node) in list(self.cache.items()):
                if self.clock() >= exp:
                    expired_entries.append((k, node))
            for k, node in expired_entries:
                del self.cache[k]
                self._remove_node(node)
                self.size -= 1
                self.stats['expirations'] += 1
            while self.size >= self.capacity:
                key_to_evict = self.tail.key
                node_to_evict = self.cache[key_to_evict][2]
                del self.cache[key_to_evict]
                self._remove_node(node_to_evict)
                self.size -= 1
                self.stats['evictions'] += 1
                
        current_time = self.clock()
        new_expiry = current_time + self.ttl
        node = Node(key)
        self.cache[key] = (value, new_expiry, node)
        self._add_node(node)
        self.size += 1

    def get(self, key, default=None):
        if self.capacity == 0:
            self.stats['misses'] += 1
            return default
            
        if key in self.cache:
            value, expiry, node = self.cache[key]
            current_time = self.clock()
            if current_time < expiry:
                self.stats['hits'] += 1
                self._remove_node(node)
                self._add_node(node)
                return value
            else:
                del self.cache[key]
                self._remove_node(node)
                self.size -= 1
                self.stats['expirations'] += 1
                self.stats['misses'] += 1
                return default
        else:
            self.stats['misses'] += 1
            return default

    def __len__(self):
        return self.size

    def stats(self):
        return self.stats
```

### Explanation
- **Initialization**: The cache initializes with a given capacity, TTL, and a clock function to get the current time. It sets up the cache dictionary, linked list (head and tail pointers), size counter, and stats dictionary.
- **`put` Method**:
  - For existing keys, updates the value and expiry time, then moves the node to the head of the LRU list.
  - For new keys, if the cache is full, it first removes all expired entries (scanning the entire cache) and then evicts the least recently used entries until space is available. It then adds the new entry to the head of the list.
- **`get` Method**:
  - If the key exists and is not expired, it moves the node to the head and returns the value (counting as a hit).
  - If the key exists but is expired, it removes the entry (counting as expiration and miss).
  - If the key is not found, it returns the default (counting as a miss).
- **`__len__` and `stats`**: Provide the current size of the cache and the operation statistics without side effects.
- **Special Case Handling**: When capacity is 0, all operations are ignored, and `get` operations count as misses.

This implementation efficiently manages the TTL cache with LRU eviction, handling both expiration and recency constraints while accurately tracking operational statistics.
Coding: Parse duration Worked
```python
def parse_duration(s: str) -> int:
    s_clean = s.strip().replace(' ', '')
    if not s_clean:
        raise ValueError("Empty string")
    
    tokens = []
    current_num = ''
    for char in s_clean:
        if char.isdigit():
            current_num += char
        else:
            if current_num == '':
                raise ValueError("Non-digit character without preceding number")
            unit = char.lower()
            if unit not in ['h', 'm', 's']:
                raise ValueError(f"Unknown unit: {char}")
            tokens.append((current_num, unit))
            current_num = ''
    
    if current_num != '':
        raise ValueError("Number without unit at the end")
    
    units = [unit for _, unit in tokens]
    if len(units) != len(set(units)):
        raise ValueError("Repeated unit")
    
    index_map = {'h': 0, 'm': 1, 's': 2}
    indices = [index_map[unit] for unit in units]
    if indices != sorted(indices):
        raise ValueError("Units out of order")
    
    total_seconds = 0
    for num_str, unit in tokens:
        num = int(num_str)
        if unit == 'h':
            total_seconds += num * 3600
        elif unit == 'm':
            total_seconds += num * 60
        else:
            total_seconds += num
    
    return total_seconds
```
Coding: Summarize ranges Worked
To solve this problem, we need to process a list of integers by sorting them, removing duplicates, and then collapsing consecutive sequences of three or more numbers into a range representation (e.g., "a..b"). Shorter sequences (1 or 2 numbers) should be listed individually. The final result should be a string where all elements are joined by commas without any spaces.

### Approach
1. **Handle Empty List**: If the input list is empty, return an empty string.
2. **Sort and Remove Duplicates**: Convert the list to a set to remove duplicates, then sort the resulting list in ascending order.
3. **Traverse the Sorted List**: Iterate through the sorted list to identify consecutive sequences:
   - For each starting index, extend the sequence as long as the next number is consecutive (i.e., current number + 1 equals the next number).
   - If the sequence length is 3 or more, represent it as "start..end".
   - If the sequence length is 1 or 2, add each number individually to the result list.
4. **Join Results**: Combine all parts (either ranges or individual numbers) into a single string separated by commas.

### Solution Code
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    sorted_nums = sorted(set(nums))
    n = len(sorted_nums)
    parts = []
    i = 0
    while i < n:
        j = i
        while j < n - 1 and sorted_nums[j] + 1 == sorted_nums[j + 1]:
            j += 1
        length = j - i + 1
        if length >= 3:
            parts.append(f"{sorted_nums[i]}..{sorted_nums[j]}")
        else:
            for k in range(i, j + 1):
                parts.append(str(sorted_nums[k]))
        i = j + 1
    return ",".join(parts)
```

### Explanation
1. **Initial Check**: The function first checks if the input list is empty and returns an empty string immediately if true.
2. **Sorting and Deduplication**: The list is converted to a set to remove duplicates, then sorted in ascending order.
3. **Sequence Identification**: Using a while loop, the algorithm processes each number in the sorted list:
   - For each starting index `i`, it finds the longest consecutive sequence by checking if the next number is exactly one more than the current number.
   - If the sequence length is 3 or more, it appends the range "start..end" to the result list.
   - If the sequence length is 1 or 2, it appends each number in the sequence as a string to the result list.
4. **Result Construction**: The individual parts (either ranges or single numbers) are joined into a single string separated by commas, which is then returned.

This approach efficiently processes the input list in linear time relative to the number of unique elements, ensuring optimal performance while meeting the problem requirements. The solution handles edge cases such as empty lists and single-element lists gracefully.
Coding: Top customers Worked
To solve this problem, we need to process a list of customer transactions, aggregate the amounts per customer (case-insensitively), round the totals to two decimal places, and return the top `n` customers sorted by their total amounts in descending order. If two customers have the same total, they should be sorted by their names in ascending order.

### Approach
1. **Initialization**: We use two dictionaries:
   - `first_occurrence` to store the first valid customer name (trimmed and case-insensitive) for each customer.
   - `total_dict` to accumulate the total amounts for each customer.
2. **Processing Rows**: For each row in the input list:
   - Skip rows with empty customer names.
   - Parse the amount string by removing commas and dollar signs, then converting it to a float. Skip rows with invalid amounts.
   - For valid entries, use the customer name in lowercase as a key to track the first occurrence of the customer and accumulate the amount in `total_dict`.
3. **Aggregation**: After processing all rows, convert each customer's total in `total_dict` to a rounded value (to two decimal places).
4. **Sorting and Selection**: Create a list of tuples (customer name, rounded total), sort this list primarily by the total in descending order and secondarily by the customer name in ascending order, then return the top `n` entries.

### Solution Code
```python
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    first_occurrence = {}
    total_dict = {}
    
    for row in rows:
        customer = row['customer'].strip()
        if customer == '':
            continue
        amount_str = row['amount'].strip()
        try:
            clean = amount_str.replace(',', '').replace('$', '')
            amount = float(clean)
        except:
            continue
        
        key = customer.lower()
        if key not in first_occurrence:
            first_occurrence[key] = customer
        total_dict[key] = total_dict.get(key, 0.0) + amount
    
    results = []
    for key, total in total_dict.items():
        rounded_total = round(total, 2)
        results.append((first_occurrence[key], rounded_total))
    
    results.sort(key=lambda x: (-x[1], x[0]))
    return results[:n]
```

### Explanation
1. **Initialization**: The dictionaries `first_occurrence` and `total_dict` are initialized to keep track of the first valid customer name (for reporting) and the cumulative amounts per customer, respectively.
2. **Row Processing**:
   - **Customer Name Check**: The customer name is stripped of leading/trailing whitespace. Rows with empty names are skipped.
   - **Amount Parsing**: The amount string is cleaned by removing commas and dollar signs. If conversion to a float fails, the row is skipped.
   - **Tracking First Occurrence**: The lowercase version of the customer name is used as a key. The first valid occurrence of this key is stored in `first_occurrence`.
   - **Accumulating Amounts**: The parsed amount is added to `total_dict` for the corresponding customer key.
3. **Result Preparation**:
   - **Rounding Totals**: Each customer's total from `total_dict` is rounded to two decimal places.
   - **Building Results**: A list of tuples (customer name, rounded total) is created using the first occurrence names from `first_occurrence`.
4. **Sorting and Selection**:
   - **Sorting**: The list is sorted by total in descending order (achieved by sorting on the negative of the total) and by customer name in ascending order.
   - **Slicing**: The top `n` entries from the sorted list are returned.

This approach efficiently processes the input data, aggregates customer amounts accurately, and delivers the required sorted results with minimal computational overhead.
Coding: Slugify Worked
```python
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    normalized = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
    normalized = normalized.lower()
    s = ''.join(['-' if not c.isalnum() else c for c in normalized])
    while '--' in s:
        s = s.replace('--', '-')
    s = s.strip('-')
    if len(s) <= max_len:
        return s
    words = s.split('-')
    if len(words[0]) > max_len:
        return words[0][:max_len]
    current_length = 0
    selected = []
    for word in words:
        if current_length == 0:
            selected.append(word)
            current_length = len(word)
        else:
            if current_length + 1 + len(word) <= max_len:
                selected.append(word)
                current_length += 1 + len(word)
            else:
                break
    return '-'.join(selected)
```
Coding: Token bucket Didn't work
```python
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.tokens = capacity
        self.last_refill_time = clock()

    def _refill(self):
        now = self.clock()
        elapsed = now - self.last_refill_time
        self.tokens += elapsed * self.refill_per_sec
        if self.tokens > self.capacity:
            self.tokens = self.capacity
        self.last_refill_time = now

    @property
    def tokens(self):
        self._refill()
        return self.tokens

    def allow(self, cost: float = 1.0) -> bool:
        self._refill()
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False
```
Decisions: Refund window (hard) 100% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.0,
    "store_credit": 1.0,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0.0,
    "1": 0.9,
    "2": 0.1,
    "3": 0.0
  },
  "page": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Tool followup (hard) 50% right
{"tool": {"web_search": 0.0, "calculator": 0.0, "calendar": 1.0, "email": 0.0, "none": 0.0}, "confirm": {"true": 0.0, "false": 1.0}}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.02,
    "false": 0.98
  }
}
Decisions: Meeting slot (hard) 100% right
{
  "slot": {
    "A": 0.0,
    "B": 1.0,
    "C": 0.0,
    "D": 0.0
  },
  "raj_last": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 0.99,
    "false": 0.01
  },
  "support": {
    "true": 0.99,
    "false": 0.01
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0.1,
    "technical": 0.9,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  },
  "outage": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 1.0,
    "technical": 0.0,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "refund": {
    "true": 1.0,
    "false": 0.0
  },
  "tone": {
    "frustrated": 0.0,
    "calm": 1.0
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.0,
    "harassment": 1.0,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.0
  },
  "personal_info": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0.05,
    "calculator": 0.02,
    "calendar": 0.9,
    "email": 0.01,
    "none": 0.02
  },
  "confirm": {
    "true": 0.95,
    "false": 0.05
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Phishing paypal 100% right
{"phishing": {"true": 1.0, "false": 0.0}, "risk": {"0": 0.0, "1": 0.0, "2": 0.0, "3": 1.0}}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.0,
    "financial": 0.0,
    "government_id": 1.0,
    "health": 0.0
  },
  "sensitive": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.1,
    "neutral": 0.0,
    "negative": 0.9
  },
  "defect": {
    "true": 1.0,
    "false": 0.0
  },
  "recommend": {
    "true": 0.0,
    "false": 1.0
  }
}
Documents: Saas escalator (hard) 100% right
{
  "year2_price_per_seat_month": 47.25,
  "year3_price_per_seat_month": 47.25,
  "year1_invoice": 58320.00,
  "year2_invoice": 61236.00,
  "addon_months_billed": 6,
  "addon_invoice": 38556.00,
  "year3_invoice": 134946.00,
  "year3_discount_percent": 15.00,
  "total_contract_value": 293058.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 100% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 2159.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 4700.00
}
Documents: Ticket SLA 92% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address", "invoice_pdf"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T15:30",
  "sla_due_utc": "2025-09-15T20:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 100% right
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 31B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.