Review · updated OCT 11

Qwen3.5-35B-A3B review: a solid all-rounder

It scored 88 out of 100, #10 of 56. It solved 24 of 30 coding jobs and scored 95 on reading documents. Runs on a 24 GB graphics card or a Mac with 48 GB.

The short version
  • Qwen3.5-35B-A3B is a free model from Alibaba's Qwen team that you can run on your own computer. In our tests it's a solid all-rounder: 88 out of 100, #10 of 56.
  • It solved 24 of 30 coding jobs and scored 95 on reading documents. On our hardest tasks it scored 73.
  • Runs on a 24 GB graphics card or a Mac with 48 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Qwen3.5-35B-A3B got 24 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Qwen3.5-35B-A3B scored 95; the best model scored 100.

TestScorePublic questionsSecret questions
Coding808678
Reading documents959895
Decisions9910099

On the 18 hardest tasks (included in the scores above) it scored 73. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Qwen3.5-35B-A3B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBYesQ4
Graphics card, 32 GBYesQ6
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBYesQ6
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Qwen3.5-35B-A3B GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Qwen3.5-35B-A3B wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    """
    Parses a string into a dictionary of dictionaries based on INI dialect rules.
    """
    # Regex to parse section headers [name]
    section_re = re.compile(r"^\s*\[\s*([^\]]+)\s*\]\s*$")
    # Regex to find interpolation patterns ${...}
    interp_re = re.compile(r"\$\{([^}]+)\}")
    
    # Dictionary to store raw parsed sections
    raw_sections: Dict[str, Dict[str, str]] = {}
    raw_sections['DEFAULT'] = {}
    
    current_section: str | None = None
    last_key: str | None = None
    current_val: str | None = None
    
    lines = text.splitlines()
    
    for line in lines:
        # Check if line is empty or a comment (Rule 3)
        # A comment line starts with ; or # (after stripping whitespace? No, "first non-space character")
        stripped = line.strip()
        if not stripped or stripped.startswith(';') or stripped.startswith('#'):
            # Rule 3: Ignored.
            # Does this break continuation?
            # Rule 6: "A line that starts with whitespace and follows a key line is a continuation".
            # If we ignore the line, it doesn't "exist" as a continuation target?
            # But Rule 3 says "Lines that are ... are ignored."
            # Usually ignored lines do not break the flow or are skipped entirely.
            # If skipped, the previous key is still active?
            # Let's assume ignored lines are skipped completely, and last_key is NOT reset.
            # (If it were a continuation line, it would be appended. If it's a comment, it's ignored).
            # But wait, if it starts with whitespace, it looks like a continuation line.
            # If it's ignored, we just skip it.
            # So `last_key` stays.
            continue
            
        # Check if line starts with whitespace (Rule 6 - Continuation)
        if line[0] in (' ', '\t'):
            if last_key is not None:
                # It is a continuation line
                # Append stripped text
                current_val = current_val + "\n" + stripped
            else:
                # Line starts with whitespace but no key was defined?
                # Rule 6 implies it follows a key line.
                # If it's the first non-empty line after a section header, is it allowed?
                # Rule 6 says "follows a key line".
                # If no key line was followed, it's invalid.
                raise ValueError("Line starting with whitespace does not follow a key line.")
            continue
        
        # If not continuation, check for section header (Rule 1)
        # Note: A section header resets continuation context
        match = section_re.match(line)
        if match:
            section_name = match.group(1)
            # If section name is DEFAULT, it is special
            # If not DEFAULT, check for duplicates? 
            # Rule 7 says later key wins. No rule about section names.
            # Assume valid.
            current_section = section_name
            if current_section not in raw_sections:
                raw_sections[current_section] = {}
            # Reset continuation
            last_key = None
            current_val = None
            continue
        
        # If we are here, it's not a section header or continuation.
        # It must be a key-value pair.
        if current_section is None:
            raise ValueError("Key-value line before any section header.")
        
        # Parse key-value (Rule 2)
        # Split on first `=` or `:`
        eq_pos = line.find('=')
        colon_pos = line.find(':')
        
        # Find the earliest separator
        if eq_pos != -1 and colon_pos != -1:
            if eq_pos < colon_pos:
                split_pos = eq_pos
            else:
                split_pos = colon_pos
        elif eq_pos != -1:
            split_pos = eq_pos
        else:
            split_pos = colon_pos
            
        if split_pos == -1:
            raise ValueError(f"Malformed line: {line}")
            
        raw_key = line[:split_pos].strip()
        raw_value = line[split_pos + 1:].strip()
        
        key = raw_key.lower()
        
        # Rule 2: Split on the FIRST. We did that.
        # Rule 7: If key appears twice, later wins.
        
        # Process value: Strip and handle quotes/comments
        value = _process_value(raw_value)
        
        raw_sections[current_section][key] = value
        last_key = key
        current_val = value # Keep track for potential continuation? 
        # Actually, `current_val` in the loop logic above was for appending.
        # But here `value` is the initial value.
        # If continuation happens, `current_val` in the loop logic handles the string concatenation.
        # Wait, I used `current_val` in the `if line[0] in ...` block.
        # So I need to update `current_val` here?
        # No, in that block, I do `current_val = current_val + ...`.
        # If this is a new key, `last_key` is set to `key`.
        # `current_val` variable in the loop logic is not strictly needed for storage 
        # because we store into `raw_sections[current_section][key]`.
        # However, `current_val` was used in the continuation logic as `current_val = current_val + ...`.
        # That logic assumes `current_val` holds the accumulated string.
        # So I should track the *string* to append to it.
        # But since I already stored it in `raw_sections`, maybe just store the current string in a local var?
        # Yes, `accumulated_value` is safer.
        
        # Let's correct the logic flow slightly in the implementation.
        # We need a variable for the current value string being built.
        # Let's call it `current_val_str`.
        # When a new key is found, set `current_val_str = value`.
        # When a continuation is found, `current_val_str += ...`.
        # And finally store `current_val_str` at the end?
        # No, `current_val_str` is the value of `last_key`.
        # If I encounter a new key, `current_val_str` for `last_key` is finished.
        # But wait, if I have `key = val\n  cont`.
        # `val` is set. `current_val_str` = `val`.
        # `cont` -> `current_val_str` += ...
        # So yes, I need to track `current_val_str`.
        # Let's refactor the loop logic.
        
        last_key = key
        current_val_str = value # This is the value for this key.
        # Wait, if I have `key = val`.
        # `raw_sections[current_section][key] = value`.
        # This is the raw value.
        # But if `cont` comes. I append to the *current* value string.
        # If I overwrite `raw_sections` now, I need to retrieve it later?
        # Or I just update `raw_sections` after the line loop ends?
        # No, that's messy.
        # Let's just store the accumulated string in `current_val_str` and write it to `raw_sections`
        # *only when* the key is completed?
        # Or just keep `current_val_str` updated and write to `raw_sections`?
        # If I have `key = val`. `raw_sections[key] = val`.
        # `cont` -> `current_val_str` (local var) += `cont`.
        # `cont` ends.
        # `raw_sections[key] = current_val_str`.
        # This works.
        # So I should NOT write to `raw_sections` immediately when new key is found.
        # I should write it either when new key found OR at end of loop?
        # If I have `key = val\n  cont\n  key2 = val2`.
        # When `key2` is found:
        # `val` is complete. `current_val_str` holds `val\ncont`.
        # I write `val\ncont` to `raw_sections[key]`.
        # Then `key` becomes `key2`.
        # Correct.
        
        # But what about the case where there is NO continuation?
        # `key = val`.
        # It is written immediately.
        # So logic:
        # 1. If continuation -> update `current_val_str`.
        # 2. Else -> if `last_key` exists -> save it. Then start new `last_key`.
        # Wait, I'm using `raw_sections` for storage.
        # So I need to write to `raw_sections[current_section][last_key]` before overwriting `last_key`.
        # And for the very last line, save it.
        
        # Revised flow inside loop:
        # ...
        # if not continuation and not header:
        #    if last_key is not None:
        #        raw_sections[current_section][last_key] = current_val_str
        #    last_key = key
        #    current_val_str = value
        #    raw_sections[current_section][key] = value # Wait, this is premature if continuation comes.
        #    # Better: Don't save yet.
        #    pass
        # ...
        # After loop: if last_key: raw_sections[...] = ...
        
        # Wait, the rule says "later value wins".
        # This applies to raw keys.
        # If `key = val` and `key = val2`.
        # `val` is overwritten by `val2`.
        # So if I don't save `val` yet.
        # And I overwrite `current_val_str` with `val2`.
        # Then if `key` has continuation after `val2`.
        # `current_val_str` becomes `val2\ncont`.
        # This is correct (val2 wins).
        # But what about `val`? It is lost.
        # Correct.
        
        # So I just track `current_val_str`.
        # And I don't write to `raw_sections` until I'm sure (new key or end).
        # Wait, if I have `key = val`.
        # And then `[Section]`.
        # Then I must save `val` to `raw_sections`.
        # So logic:
        #   When `key` line is found (and not continuation/header):
        #     If `last_key` is not None:
        #        save `current_val_str` to `raw_sections[current_section][last_key]`
        #     Start new `last_key = key`.
        #     `current_val_str` = `value`.
        #
        #   When `[Section]` is found:
        #     If `last_key` is not None:
        #        save.
        #     `last_key` = None.
        #
        #   At end of loop:
        #     If `last_key` is not None: save.
        
        # Wait, `last_key` logic inside the loop:
        # I already checked `line[0] in (' ', '\t')`.
        # If it is a continuation, I append to `current_val_str`?
        # Wait, I haven't defined `current_val_str` yet.
        # I need to initialize it when `last_key` is set.
        
        # Let's restructure.
        pass # Placeholder for logic
    
    # ... (I will write the full loop logic in the final code block) ...
    
    # Interpolation and Merging
    # We need to resolve values.
    # `resolved_data`
    
    # ... (Implementation details)
    pass

# Let's assemble the final logic.
# The structure:
# 1. Parse lines into raw_sections (with DEFAULT).
#    Track `last_key` and `current_val_str`.
#    Handle continuation.
#    Handle comments.
#    Handle section headers.
# 2. Process Interpolation.
#    `resolve(section, key, visited)` helper.
#    Iterate all sections (including DEFAULT if needed for lookup).
#    But output only non-DEFAULT sections.
#    Merge DEFAULT into sections?
#    Actually, we can just use `lookup` to get values.
#    Wait, output must be a dict.
#    So we must construct the final dicts.
#    For `Sec`: `res_Sec = {}`.
#    For `key` in `Sec`: `res_Sec[key] = resolve(key, Sec, set())`.
#    Wait, `Sec` has keys.
#    Does `Sec` have to inherit `DEFAULT`?
#    If `Sec` has `a`. `DEFAULT` has `a`.
#    `Sec` uses `a`.
#    So `resolve` logic handles lookup priority.
#    But the returned dict keys?
#    "The section DEFAULT is special: it is NOT returned as a section".
#    So keys from `DEFAULT` are included in `Sec` *if* `Sec` doesn't define them.
#    So `res_Sec` should contain `Sec`'s keys AND `DEFAULT`'s keys (if not overridden).
#    Wait, `DEFAULT` keys are "inherited".
#    So the output dict for `Sec` should effectively be `raw[Sec].update(raw[DEFAULT])`.
#    (With priority to `raw[Sec]`).
#    And then interpolate.
#    Wait, if `DEFAULT` has `x=${y}`.
#    Does `x` use `y` from `DEFAULT`?
#    Yes.
#    So `x` in `DEFAULT` is resolved first.
#    Then `Sec` copies resolved `x` (or inherits logic).
#    But `DEFAULT` values are not returned.
#    But their *resolved* values are what `Sec` inherits.
#    So:
#    1. Resolve `DEFAULT` keys first (if they are to be used).
#       Wait, `DEFAULT` keys might be used by `DEFAULT` keys?
#       Yes.
#       So we must resolve `DEFAULT` keys completely.
#       Wait, `DEFAULT` is not returned.
#       So we can just store resolved `DEFAULT` values in `raw_sections['DEFAULT']`.
#    2. For every other section `S`:
#       - Merge `raw['DEFAULT']` into `S` (overwriting `S` keys is not right, `S` has priority).
#         Wait, rule 8: "inherited ... unless that section defines the key itself".
#         This means if `S` has `key`, use `S`. Else use `DEFAULT`.
#         So `merged_keys = {**raw_sections['DEFAULT'], **raw_sections[S]}`.
#         Wait, merging raw keys.
#         Then resolve values in `merged_keys`?
#         No, resolve logic is key-dependent.
#         It seems cleaner to resolve on the fly.
#         But we need to return the dictionary.
#         The output requires the *resolved* values.
#         And the merged keys.
#         So `result[S]` contains all keys of `S` (raw and merged defaults).
#         Wait, "DEFAULT ... is NOT returned as a section".
#         So `DEFAULT`'s keys are merged into others.
#         So `result[S]` keys = `raw[S].keys() | raw['DEFAULT'].keys()`.
#         Values are resolved.
#         When resolving `key` in `S`:
#            - If `key` in `raw[S]`, source is `S`.
#            - Else source is `DEFAULT`.
#            - If source is `DEFAULT`, does it inherit from `DEFAULT`? Yes.
#            - If source is `S`, does it inherit from `S`? Yes (S's own keys).
#            - Note: `S` does not inherit from `DEFAULT` for *its own* values, only *lookup* fallback.
#              But wait, "inherited" usually means merged.
#              If `DEFAULT` has `a`. `S` has `a`. `S.a` = `S's raw`.
#              If `S` has `b=${a}`.
#              `S` looks up `a`. Finds `S`'s `a`.
#              So `b` resolves `S`'s `a`.
#              If `S` does not have `a`. `S` looks up `DEFAULT`'s `a`.
#              So `b` resolves `DEFAULT`'s `a`.
#              So yes, lookup logic handles this.
#         So `result[S]` keys = `raw[S].keys() | raw['DEFAULT'].keys()`.
#         For each key in this union:
#             `val = lookup_raw(key, S)`.
#             `resolved_val = resolve(val, context, visited)`.
#             `context`?
#             If `key` was in `raw[S]`, context is `S`.
#             If `key` was in `raw['DEFAULT']` (and not `S`), context is `DEFAULT`.
#             Wait, `DEFAULT` keys are not returned.
#             But `DEFAULT` keys are referenced.
#             And `DEFAULT` keys might be resolved using `DEFAULT` context.
#             So context should be the source section.
#         This seems consistent.
    
    # Wait, what about `${section:key}` in `S` referring to `DEFAULT:key`?
    # It resolves `DEFAULT:key`.
    # `DEFAULT:key` is in `DEFAULT`.
    # So context is `DEFAULT`.
    
    # One detail: "DEFAULT ... is NOT returned".
    # But `DEFAULT` keys are merged into output sections.
    # So `result` dict only has keys other than `DEFAULT`.
    
    # Interpolation recursion:
    # If `S`
Coding: TTL LRU (hard) Worked
```python
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self._cache = OrderedDict()
        self._stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
    
    def _now(self) -> float:
        return self.clock()
    
    def _is_expired(self, expiry_time: float) -> bool:
        return self._now() >= expiry_time
    
    def _remove_expired_keys(self):
        """Remove all expired keys and update expirations count."""
        expired_keys = []
        for key in self._cache:
            if self._is_expired(self._cache[key][1]):
                expired_keys.append(key)
        for key in expired_keys:
            del self._cache[key]
            self._stats["expirations"] += 1
    
    def _evict_lru(self):
        """Evict least recently used entry and increment eviction counter."""
        if self._cache:
            self._cache.popitem(last=False)
            self._stats["evictions"] += 1
    
    def put(self, key, value):
        if self.capacity == 0:
            return
        
        now = self._now()
        self._remove_expired_keys()
        
        if key in self._cache:
            _, _ = self._cache[key]
            self._cache[key] = (value, now + self.ttl)
            self._cache.move_to_end(key)
        else:
            if len(self._cache) >= self.capacity:
                self._evict_lru()
            self._cache[key] = (value, now + self.ttl)
    
    def get(self, key, default=None):
        if key not in self._cache:
            self._stats["misses"] += 1
            return default
        
        value, expiry_time = self._cache[key]
        
        if self._is_expired(expiry_time):
            del self._cache[key]
            self._stats["expirations"] += 1
            self._stats["misses"] += 1
            return default
        
        self._cache.move_to_end(key)
        self._stats["hits"] += 1
        return value
    
    def __len__(self) -> int:
        now = self._now()
        return sum(1 for _, (v, e) in self._cache.items() if not self._is_expired(e))
    
    def stats(self):
        return {
            "hits": self._stats["hits"],
            "misses": self._stats["misses"],
            "evictions": self._stats["evictions"],
            "expirations": self._stats["expirations"]
        }
```
Coding: Parse duration Worked
```python
import re

def parse_duration(s: str) -> int:
    if not s or not s.strip():
        raise ValueError("Empty or blank string")

    # Remove all whitespace to simplify parsing
    s_normalized = s.strip().replace(' ', '')

    # If no digits found, ensure no stray characters remain
    if not re.search(r'\d', s_normalized):
        raise ValueError("Invalid format")

    # Extract all tokens matching pattern: digits followed by single h/m/s (case-insensitive)
    tokens = re.findall(r'\d+[hmsHMS]', s_normalized)

    if not tokens:
        # Check if digits exist but no unit follows
        if re.search(r'\d', s_normalized):
            raise ValueError("Number without unit")
        raise ValueError("Empty or blank string")

    # Validate that entire normalized string is fully consumed by tokens
    remaining = re.sub(r'\d+[hmsHMS]', '', s_normalized)
    if remaining:
        raise ValueError("Invalid format")

    total_seconds = 0
    seen_units: set[str] = set()
    last_unit_index: int | None = None
    order = {'h': 0, 'm': 1, 's': 2}

    for token in tokens:
        unit_char = token[-1].lower()
        number = int(token[:-1])

        if unit_char not in order:
            raise ValueError("Unknown unit")

        if unit_char in seen_units:
            raise ValueError("Repeated unit")
        seen_units.add(unit_char)

        current_index = order[unit_char]

        if last_unit_index is not None and current_index <= last_unit_index:
            raise ValueError("Units out of order")

        if unit_char == 'h':
            total_seconds += number * 3600
        elif unit_char == 'm':
            total_seconds += number * 60
        elif unit_char == 's':
            total_seconds += number

        last_unit_index = current_index

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    sorted_nums = sorted(set(nums))
    result = []
    i = 0
    n = len(sorted_nums)
    
    while i < n:
        start = sorted_nums[i]
        j = i + 1
        
        while j < n and sorted_nums[j] == sorted_nums[j - 1] + 1:
            j += 1
            
        end = sorted_nums[j - 1]
        length = j - i
        
        if length >= 3:
            result.append(f"{start}..{end}")
        else:
            for k in range(i, j):
                result.append(str(sorted_nums[k]))
        
        i = j
        
    return ",".join(result)
```
Coding: Top customers Worked
```python
from collections import defaultdict

def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    customer_totals = defaultdict(float)
    first_valid_names = {}

    for row in rows:
        # Extract and validate customer name
        raw_name = row.get("customer", "")
        if not isinstance(raw_name, str):
            continue
        
        customer = raw_name.strip()
        if not customer:
            continue

        # Extract and validate amount
        raw_amount = row.get("amount", "")
        if not isinstance(raw_amount, str):
            continue
            
        # Parse amount
        try:
            # Remove currency symbols ('$') and thousands separators (',')
            # e.g., "$1,234.50" becomes "1234.50"
            clean_amount = raw_amount.replace("$", "").replace(",", "")
            amount = float(clean_amount)
        except (ValueError, TypeError):
            continue
        
        # Normalize customer name for aggregation (case-insensitive)
        normalized_name = customer.lower()
        
        # Preserve the trimmed spelling of the first valid row
        if normalized_name not in first_valid_names:
            first_valid_names[normalized_name] = customer
            
        customer_totals[normalized_name] += amount

    # Prepare list of tuples with rounded totals
    results = []
    for name_norm, total in customer_totals.items():
        display_name = first_valid_names[name_norm]
        rounded_total = round(total, 2)
        results.append((display_name, rounded_total))

    # Sort by total descending, then name ascending
    results.sort(key=lambda x: (-x[1], x[0]))

    # Return top n
    return results[:n]
```
Coding: Slugify Worked
```python
import re
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    # 1. Transliterate accented characters to ASCII and drop other non-ASCII
    # NFKD normalization decomposes characters (e.g., é -> e + combining acute)
    # Encoding to ASCII with 'ignore' drops the non-ASCII combining marks
    text = unicodedata.normalize('NFKD', title)
    text = text.encode('ascii', 'ignore').decode('ascii')
    
    # 2. Lowercase
    text = text.lower()
    
    # 3. Replace non-alphanumeric runs with single '-', strip leading/trailing
    text = re.sub(r'[^a-z0-9]+', '-', text)
    text = text.strip('-')
    
    # 4. Truncate to max_len respecting word boundaries
    if len(text) <= max_len:
        return text
    
    words = text.split('-')
    result_parts = []
    current_len = 0
    
    for word in words:
        # If this is the first word and it exceeds max_len, hard cut it
        if not result_parts and len(word) > max_len:
            return word[:max_len]
        
        # Calculate length if this word is added (including separator if not first)
        separator = 1 if result_parts else 0
        potential_len = current_len + len(word) + separator
        
        if potential_len <= max_len:
            result_parts.append(word)
            current_len = potential_len
        else:
            # If it doesn't fit, stop (unless it was the first word > max_len handled above)
            break
            
    slug = '-'.join(result_parts)
    
    # 5. Ensure result never ends with "-" (join logic handles this, but strip is safe fallback)
    return slug.strip('-')
```
Coding: Token bucket Worked
```python
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self._tokens = capacity
        self._last_refill = clock()

    def _refill(self) -> None:
        now = self.clock()
        elapsed = now - self._last_refill
        
        # Only refill if time has moved forward
        if elapsed > 0:
            tokens_to_add = elapsed * self.refill_per_sec
            self._tokens = min(self.capacity, self._tokens + tokens_to_add)
            self._last_refill = now

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False
```
Decisions: Refund window (hard) 100% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.0,
    "store_credit": 1.0,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0.0,
    "1": 0.95,
    "2": 0.05,
    "3": 0.0
  },
  "page": {
    "true": 0.05,
    "false": 0.95
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 1.0,
    "email": 0.0,
    "none": 0.0
  },
  "confirm": {
    "true": 0.85,
    "false": 0.15
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.0,
    "false": 1.0
  },
  "action_needed": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Meeting slot (hard) 100% right
{
  "slot": {
    "A": 0.0,
    "B": 1.0,
    "C": 0.0,
    "D": 0.0
  },
  "raj_last": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 0.99,
    "false": 0.01
  },
  "support": {
    "true": 0.98,
    "false": 0.02
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0.04,
    "technical": 0.96,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.01,
    "3": 0.99
  },
  "outage": {
    "true": 0.98,
    "false": 0.02
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.99,
    "technical": 0.0,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.01
  },
  "refund": {
    "true": 1.0,
    "false": 0.0
  },
  "tone": {
    "frustrated": 0.0,
    "calm": 1.0
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.05,
    "harassment": 0.95,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.0
  },
  "personal_info": {
    "true": 0.99,
    "false": 0.01
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 0.97,
    "email": 0.03,
    "none": 0.0
  },
  "confirm": {
    "true": 0.98,
    "false": 0.02
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Phishing paypal 100% right
{
  "phishing": {
    "true": 0.99,
    "false": 0.01
  },
  "risk": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.0,
    "financial": 0.02,
    "government_id": 0.98,
    "health": 0.0
  },
  "sensitive": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.05,
    "neutral": 0.01,
    "negative": 0.94
  },
  "defect": {
    "true": 0.99,
    "false": 0.01
  },
  "recommend": {
    "true": 0.01,
    "false": 0.99
  }
}
Documents: Saas escalator (hard) 100% right
{
  "year2_price_per_seat_month": 47.25,
  "year3_price_per_seat_month": 47.25,
  "year1_invoice": 58320.00,
  "year2_invoice": 61236.00,
  "addon_months_billed": 6,
  "addon_invoice": 38556.00,
  "year3_invoice": 134946.00,
  "year3_discount_percent": 15,
  "total_contract_value": 293058.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 100% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 2159.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right
{
  "tenants": [
    "Marcus Lin",
    "Sofia Lin"
  ],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2236,
  "late_fee_from_2025_06_01": 111.8,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2306,
  "move_in_payment": 4700
}
Documents: Ticket SLA 92% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": [
    "billing_address",
    "invoice_pdf"
  ],
  "affected_orders": [
    "SO-99812",
    "SO-99820",
    "SO-99827"
  ],
  "priority": "P2",
  "sla_due_local": "2025-09-15T15:30",
  "sla_due_utc": "2025-09-15T20:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 100% right
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": [
    "East"
  ],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 36B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.