Review · updated OCT 11

gpt-oss-120b review: one of the best local models we've tested

It scored 96 out of 100, #2 of 56. It solved 28 of 30 coding jobs and scored 99 on reading documents. Runs on a Mac with 96 GB.

The short version
  • gpt-oss-120b is a free model from OpenAI that you can run on your own computer. In our tests it's one of the best local models we've tested: 96 out of 100, #2 of 56.
  • It solved 28 of 30 coding jobs and scored 99 on reading documents. On our hardest tasks it scored 94.
  • Runs on a Mac with 96 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. gpt-oss-120b got 28 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. gpt-oss-120b scored 99; the best model scored 100.

TestScorePublic questionsSecret questions
Coding9310091
Reading documents9996100
Decisions9910099

On the 18 hardest tasks (included in the scores above) it scored 94. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

gpt-oss-120b comes in one standard size, which needs about 63.3 GB of memory.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBNo–
Mac, 96 GBYesstandard
Mac, 128 GBYesstandard

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run gpt-oss:120b

Its actual answers

What gpt-oss-120b wrote for our public questions, word for word.

Coding: INI parser (hard) Worked
```python
import re
from typing import Dict, Set


def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    """
    Parse an INI file according to the dialect described in the prompt.

    Returns a mapping ``section -> {key -> value}`` where *section* does not
    include the special ``DEFAULT`` section (its keys are inherited by all
    other sections).  All keys are lower‑cased.
    """
    # --------------------------------------------------------------------- #
    # 1.  Lexical parsing – sections, keys, continuations, comments, quoting #
    # --------------------------------------------------------------------- #
    section_re = re.compile(r'^\s*\[(.+?)\]\s*$')
    var_re = re.compile(r'\$\{([^}]+)\}')

    default: Dict[str, str] = {}
    sections: Dict[str, Dict[str, str]] = {}
    current: str | None = None               # name of the current section
    last_key: Dict[str, str | None] = {}     # last key seen per section

    for lineno, raw_line in enumerate(text.splitlines(), start=1):
        line = raw_line.rstrip("\r\n")

        # ignore empty lines and full‑line comments
        stripped = line.lstrip()
        if not stripped or stripped[0] in (";", "#"):
            continue

        # ----------------------------------------------------------------- #
        # Section header
        # ----------------------------------------------------------------- #
        m = section_re.match(line)
        if m:
            name = m.group(1)
            current = name
            if name != "DEFAULT":
                sections.setdefault(name, {})
            last_key[current] = None
            continue

        # ----------------------------------------------------------------- #
        # Continuation line (must follow a key line)
        # ----------------------------------------------------------------- #
        if line and line[0].isspace():
            if current is None:
                raise ValueError(f"Continuation line before any section (line {lineno})")
            key = last_key.get(current)
            if not key:
                raise ValueError(f"Continuation line without preceding key (line {lineno})")
            cont = line.strip()
            target = default if current == "DEFAULT" else sections[current]
            target[key] = f"{target[key]}\n{cont}"
            continue

        # ----------------------------------------------------------------- #
        # Key/value line
        # ----------------------------------------------------------------- #
        if current is None:
            raise ValueError(f"Key/value line before any section (line {lineno})")

        # locate first ':' or '='
        eq = line.find("=")
        colon = line.find(":")
        if eq == -1 and colon == -1:
            raise ValueError(f"Malformed line (no ':' or '=') (line {lineno})")
        if eq != -1 and colon != -1:
            pos = min(eq, colon)
        else:
            pos = eq if eq != -1 else colon

        raw_key = line[:pos]
        raw_val = line[pos + 1 :]

        key = raw_key.strip().lower()
        # ------------------------------------------------------------- #
        # Value processing – quoting, stripping, inline comments
        # ------------------------------------------------------------- #
        val = raw_val.lstrip()
        if val.startswith('"'):
            # quoted value – keep everything inside the quotes verbatim
            end = val.find('"', 1)
            if end == -1:
                raise ValueError(f"Unclosed double quote (line {lineno})")
            processed = val[1:end]  # no stripping
        else:
            # inline comment detection (space + ; or #)
            comment_pos = None
            for marker in (" ;", " #"):
                idx = val.find(marker)
                if idx != -1:
                    comment_pos = idx if comment_pos is None else min(comment_pos, idx)
            processed = val[:comment_pos] if comment_pos is not None else val
            processed = processed.strip()

        # store the raw (still possibly containing ${...}) value
        target = default if current == "DEFAULT" else sections[current]
        target[key] = processed
        last_key[current] = key

    # --------------------------------------------------------------------- #
    # 2.  Interpolation – recursive substitution with cycle detection
    # --------------------------------------------------------------------- #
    def get_raw(sec: str, k: str) -> str:
        """Return the raw (un‑interpolated) value for *k* in *sec*,
        falling back to DEFAULT if needed."""
        if sec != "DEFAULT" and k in sections.get(sec, {}):
            return sections[sec][k]
        if k in default:
            return default[k]
        raise KeyError(f"Key '{k}' not found in section '{sec}'")

    def resolve(sec: str, k: str, seen: Set[tuple[str, str]]) -> str:
        """Resolve *k* in *sec* with interpolation.  *seen* tracks the
        current resolution chain to detect cycles."""
        if (sec, k) in seen:
            raise ValueError(f"Interpolation cycle involving {sec}:{k}")
        seen.add((sec, k))
        raw = get_raw(sec, k)

        def repl(m: re.Match) -> str:
            inner = m.group(1)
            if ":" in inner:
                other_sec, other_key = inner.split(":", 1)
            else:
                other_sec, other_key = sec, inner
            return resolve(other_sec, other_key, seen)

        # Perform substitution; repeat until no more variables appear
        result = raw
        while True:
            if not var_re.search(result):
                break
            result = var_re.sub(repl, result)
        seen.remove((sec, k))
        return result

    # --------------------------------------------------------------------- #
    # 3.  Build final mapping – inherit DEFAULT, apply interpolation
    # --------------------------------------------------------------------- #
    final: Dict[str, Dict[str, str]] = {}
    for sec, kv in sections.items():
        if sec == "DEFAULT":
            continue
        merged_keys = set(default) | set(kv)
        resolved_section: Dict[str, str] = {}
        for k in merged_keys:
            resolved_section[k] = resolve(sec, k, set())
        final[sec] = resolved_section

    return final
```
Coding: TTL LRU (hard) Worked
```python
from __future__ import annotations

import collections
from collections import OrderedDict
from typing import Callable, Dict, Generic, Hashable, Iterable, Tuple, TypeVar

K = TypeVar("K", bound=Hashable)
V = TypeVar("V")
TClock = Callable[[], float]


class TTLCache(Generic[K, V]):
    """Least‑Recently‑Used cache with per‑entry TTL expiration.

    Parameters
    ----------
    capacity: int
        Maximum number of entries that can be stored. ``0`` disables storage.
    ttl: float
        Time‑to‑live for each entry in seconds.
    clock: Callable[[], float]
        A zero‑argument function returning the current time in seconds.
    """

    __slots__ = (
        "_capacity",
        "_ttl",
        "_clock",
        "_store",
        "_hits",
        "_misses",
        "_evictions",
        "_expirations",
    )

    def __init__(self, capacity: int, ttl: float, clock: TClock) -> None:
        if capacity < 0:
            raise ValueError("capacity must be non‑negative")
        if ttl < 0:
            raise ValueError("ttl must be non‑negative")
        self._capacity: int = capacity
        self._ttl: float = ttl
        self._clock: TClock = clock
        self._store: "OrderedDict[K, Tuple[V, float]]" = OrderedDict()
        self._hits = 0
        self._misses = 0
        self._evictions = 0
        self._expirations = 0

    # ------------------------------------------------------------------ helpers
    def _now(self) -> float:
        return self._clock()

    def _remove_expired(self) -> None:
        """Delete all expired entries, updating the expiration counter."""
        now = self._now()
        # list() copy to avoid mutation during iteration
        for key, (_, expiry) in list(self._store.items()):
            if now >= expiry:
                self._store.pop(key, None)
                self._expirations += 1

    # -------------------------------------------------------------------- public
    def put(self, key: K, value: V) -> None:
        """Insert or replace an entry.

        Replacements refresh the expiry time and mark the entry as most‑recently
        used. When a new key would overflow the capacity, expired entries are
        removed first; if the cache is still full, the least‑recently‑used entry
        is evicted.
        """
        if self._capacity == 0:
            return

        now = self._now()

        if key in self._store:
            # Replace existing entry.
            self._store[key] = (value, now + self._ttl)
            self._store.move_to_end(key)  # mark as most recent
            return

        # New key – possibly need to free space.
        self._remove_expired()

        if len(self._store) >= self._capacity:
            # Evict least‑recently‑used entry.
            evicted_key, _ = self._store.popitem(last=False)
            self._evictions += 1
            # (evicted_key is discarded)

        self._store[key] = (value, now + self._ttl)
        # New insert is automatically at the end (most recent).

    def get(self, key: K, default: V | None = None) -> V | None:
        """Retrieve a value if present and not expired.

        A successful lookup marks the entry as most‑recently‑used.
        Expired entries are removed, counted as expirations, and result in a miss.
        """
        if self._capacity == 0:
            self._misses += 1
            return default

        now = self._now()
        entry = self._store.get(key)

        if entry is None:
            self._misses += 1
            return default

        value, expiry = entry
        if now >= expiry:
            # Expired – remove and count as miss+expiration.
            self._store.pop(key, None)
            self._expirations += 1
            self._misses += 1
            return default

        # Hit – update recency.
        self._hits += 1
        self._store.move_to_end(key)
        return value

    def __len__(self) -> int:
        """Number of unexpired entries at the current time (does not mutate)."""
        now = self._now()
        return sum(1 for _, (_, expiry) in self._store.items() if now < expiry)

    def stats(self) -> Dict[str, int]:
        """Return a dictionary with cache statistics."""
        return {
            "hits": self._hits,
            "misses": self._misses,
            "evictions": self._evictions,
            "expirations": self._expirations,
        }
```
Coding: Parse duration Worked
```python
import re

def parse_duration(s: str) -> int:
    """
    Convert a duration string to a total number of seconds.

    Supported units (case‑insensitive):
        h – hours
        m – minutes
        s – seconds

    Rules
    -----
    * Each part is a non‑negative integer followed by a unit, e.g. "1h30m".
    * Whitespace may appear anywhere and is ignored.
    * Units must appear in the order: hours, minutes, seconds; each at most once.
    * Invalid input raises ValueError.

    Parameters
    ----------
    s: str
        The duration string.

    Returns
    -------
    int
        Total duration expressed in seconds.
    """
    if not isinstance(s, str):
        raise TypeError("Input must be a string")

    # Remove all whitespace characters
    s_clean = ''.join(s.split())
    if not s_clean:
        raise ValueError("Duration string is empty or blank")

    i = 0
    last_order = 0          # 0 = none, 1 = h, 2 = m, 3 = s
    total_seconds = 0
    unit_to_order = {'h': 1, 'm': 2, 's': 3}
    unit_to_seconds = {'h': 3600, 'm': 60, 's': 1}

    while i < len(s_clean):
        # --- parse the integer part ---
        start = i
        while i < len(s_clean) and s_clean[i].isdigit():
            i += 1
        if start == i:                     # no digits found
            raise ValueError(f"Expected number at position {start}")

        number = int(s_clean[start:i])

        # --- parse the unit character ---
        if i >= len(s_clean):
            raise ValueError(f"Missing unit after number at position {start}")
        unit_char = s_clean[i].lower()
        i += 1

        if unit_char not in unit_to_order:
            raise ValueError(f"Unknown unit '{unit_char}' at position {i-1}")

        order = unit_to_order[unit_char]
        if order <= last_order:            # repeated or out‑of‑order unit
            raise ValueError("Units must appear once and in the order h, m, s")

        total_seconds += number * unit_to_seconds[unit_char]
        last_order = order

    return total_seconds
```
Coding: Summarize ranges Worked
```python
from __future__ import annotations

def summarize_ranges(nums: list[int]) -> str:
    """Return a comma‑separated string that summarizes consecutive integer runs.

    * The input list is sorted and deduplicated.
    * Runs of three or more consecutive numbers become “a..b”.
    * Runs of one or two numbers are emitted individually.
    * An empty input returns an empty string.
    """
    if not nums:
        return ""

    # Sort and remove duplicates.
    sorted_unique = sorted(set(nums))

    parts: list[str] = []
    i = 0
    n = len(sorted_unique)

    while i < n:
        start = sorted_unique[i]
        j = i
        # Extend the run while the next number is exactly one larger.
        while j + 1 < n and sorted_unique[j + 1] == sorted_unique[j] + 1:
            j += 1

        run_len = j - i + 1
        if run_len >= 3:
            # Collapse long runs into "a..b".
            parts.append(f"{start}..{sorted_unique[j]}")
        else:
            # Emit each element of short runs individually.
            for k in range(i, j + 1):
                parts.append(str(sorted_unique[k]))

        i = j + 1

    return ",".join(parts)
```
Coding: Top customers Worked
```python
import re
from typing import List, Dict, Tuple, Optional


def _parse_amount(value: str) -> Optional[float]:
    """
    Convert a monetary string to a float.
    Accepted formats include "$1,234.50", "1234.5", "-$5.00", "$1,000", etc.
    Returns ``None`` if the string cannot be parsed.
    """
    if not isinstance(value, str):
        return None
    # Remove surrounding whitespace
    s = value.strip()
    # Quick reject empty strings
    if not s:
        return None

    # Remove dollar signs and commas
    s = s.replace("$", "").replace(",", "")
    # After cleanup, the string should represent a valid float
    try:
        return float(s)
    except ValueError:
        return None


def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    """
    Compute the top *n* customers by total transaction amount.

    Parameters
    ----------
    rows : list[dict]
        Each dict must contain at least the keys ``"customer"`` and ``"amount"``.
        ``"customer"`` is a string name (may contain surrounding whitespace).
        ``"amount"`` is a string that may include a leading ``$``, commas, and
        an optional leading ``-`` for refunds.
    n : int
        Number of top customers to return.

    Returns
    -------
    list[tuple[str, float]]
        A list of ``(customer_name, total)`` tuples, where *customer_name*
        preserves the original spelling from its first **valid** row,
        ``total`` is rounded to two decimal places, and the list is sorted
        by descending total then ascending name.
    """
    # Mapping: normalized name -> (original spelling, accumulated total)
    aggregates: Dict[str, Tuple[str, float]] = {}

    for row in rows:
        # Extract and clean customer name
        raw_name = row.get("customer", "")
        if not isinstance(raw_name, str):
            continue
        name = raw_name.strip()
        if not name:
            continue  # skip empty names

        # Parse amount
        raw_amount = row.get("amount", "")
        amount = _parse_amount(raw_amount)
        if amount is None:
            continue  # skip rows with unparsable amount

        norm = name.lower()
        if norm in aggregates:
            orig_name, current_total = aggregates[norm]
            aggregates[norm] = (orig_name, current_total + amount)
        else:
            aggregates[norm] = (name, amount)

    # Build list of (name, rounded_total)
    result = [
        (orig_name, round(total, 2))
        for orig_name, total in (v for v in aggregates.values())
    ]

    # Sort by total descending, then name ascending
    result.sort(key=lambda item: (-item[1], item[0]))

    # Return top n entries
    return result[:n]
```
Coding: Slugify Worked
```python
import re
import unicodedata


def slugify(title: str, max_len: int = 50) -> str:
    """
    Convert a title to a URL slug.

    Steps:
    1. Transliterate accented characters to ASCII, dropping any other non‑ASCII.
    2. Lower‑case the result.
    3. Replace runs of non‑alphanumeric characters with a single hyphen.
    4. Trim leading/trailing hyphens.
    5. Enforce ``max_len`` without breaking words (unless a single word exceeds the limit).
    """
    # 1. Transliterate / drop non‑ASCII
    normalized = unicodedata.normalize("NFKD", title)
    ascii_str = normalized.encode("ascii", "ignore").decode("ascii")

    # 2. Lowercase
    lower = ascii_str.lower()

    # 3. Replace runs of non‑alphanumerics with hyphen, then strip hyphens
    slug = re.sub(r"[^a-z0-9]+", "-", lower).strip("-")

    # 4. Enforce max_len without cutting words
    if len(slug) > max_len:
        words = slug.split("-")
        kept_words = []
        cur_len = 0
        for w in words:
            # length if we add this word (including a hyphen if not the first word)
            add_len = len(w) + (1 if kept_words else 0)
            if cur_len + add_len <= max_len:
                kept_words.append(w)
                cur_len += add_len
            else:
                break

        if not kept_words:
            # First word itself is longer than max_len → hard cut
            slug = slug[:max_len]
        else:
            slug = "-".join(kept_words)

    # 5. Ensure no trailing hyphen (should already be the case)
    return slug.rstrip("-")
```
Coding: Token bucket Worked
```python
from __future__ import annotations

from typing import Callable


class TokenBucket:
    """
    A simple token bucket rate limiter.

    Parameters
    ----------
    capacity : float
        Maximum number of tokens the bucket can hold.
    refill_per_sec : float
        Rate at which tokens are added to the bucket (tokens per second).
    clock : Callable[[], float]
        Zero‑argument function returning the current time in seconds.
        The implementation must never call ``time.time()`` directly.
    """

    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]) -> None:
        if capacity <= 0:
            raise ValueError("capacity must be positive")
        if refill_per_sec < 0:
            raise ValueError("refill_per_sec cannot be negative")
        self._capacity = float(capacity)
        self._refill_per_sec = float(refill_per_sec)
        self._clock = clock

        # The bucket starts full.
        self._tokens: float = self._capacity
        self._last_timestamp: float = self._clock()

    def _refill(self) -> None:
        """Refill the bucket based on elapsed time."""
        now = self._clock()
        elapsed = now - self._last_timestamp
        if elapsed > 0:
            added = elapsed * self._refill_per_sec
            self._tokens = min(self._capacity, self._tokens + added)
            self._last_timestamp = now
        else:
            # No time passed; keep the timestamp as‑is.
            # This also covers the case where the clock may go backwards.
            self._last_timestamp = now

    def allow(self, cost: float = 1) -> bool:
        """
        Attempt to consume ``cost`` tokens.

        Returns
        -------
        bool
            ``True`` if the bucket had enough tokens (they are deducted),
            ``False`` otherwise (no tokens are deducted).
        """
        if cost < 0:
            raise ValueError("cost cannot be negative")
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        """Current token count after refilling (read‑only)."""
        self._refill()
        return self._tokens
```
Decisions: Refund window (hard) 100% right
{"outcome":{"full_refund":0.0,"partial_refund":0.0,"store_credit":0.99,"nothing":0.01},"defective":{"true":0.0,"false":1.0}}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0.01,
    "1": 0.98,
    "2": 0.01,
    "3": 0.0
  },
  "page": {
    "true": 0.02,
    "false": 0.98
  }
}
Decisions: Tool followup (hard) 100% right
{"tool":{"web_search":0.01,"calculator":0.01,"calendar":0.96,"email":0.01,"none":0.01},"confirm":{"true":0.95,"false":0.05}}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {"true": 0.02, "false": 0.98},
  "action_needed": {"true": 0.03, "false": 0.97}
}
Decisions: Meeting slot (hard) 100% right
{"slot":{"A":0,"B":1,"C":0,"D":0},"raj_last":{"true":1,"false":0}}
Decisions: Review mixed (hard) 100% right
{"hardware":{"true":0.99,"false":0.01},"support":{"true":0.99,"false":0.01}}
Decisions: Support checkout down 100% right
{"department":{"billing":0.1,"technical":0.9,"account":0.0,"shipping":0.0,"sales":0.0},"urgency":{"0":0.0,"1":0.0,"2":0.05,"3":0.95},"outage":{"true":0.98,"false":0.02}}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.94,
    "technical": 0.02,
    "account": 0.02,
    "shipping": 0.01,
    "sales": 0.01
  },
  "refund": {
    "true": 0.99,
    "false": 0.01
  },
  "tone": {
    "calm": 0.98,
    "frustrated": 0.02
  }
}
Decisions: Moderation doxxing 100% right
{"policy":{"none":0.0,"harassment":1.0,"hate":0.0,"spam":0.0,"self_harm":0.0},"personal_info":{"true":1.0,"false":0.0}}
Decisions: Route calendar 100% right
{"tool":{"web_search":0.01,"calculator":0.01,"calendar":0.96,"email":0.01,"none":0.01},"confirm":{"true":0.98,"false":0.02}}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 0.98,
    "resume": 0.005,
    "contract": 0.005,
    "bank_statement": 0.005,
    "other": 0.005
  },
  "missing_due_date": {
    "true": 0.96,
    "false": 0.04
  }
}
Decisions: Phishing paypal 100% right
{"phishing":{"true":0.99,"false":0.01},"risk":{"0":0.01,"1":0.01,"2":0.01,"3":0.97}}
Decisions: Pii ssn email 100% right
{"data_kind":{"none":0.01,"contact":0.04,"financial":0.05,"government_id":0.88,"health":0.02},"sensitive":{"true":0.97,"false":0.03}}
Decisions: Review mixed 100% right
{"sentiment":{"positive":0.02,"neutral":0.08,"negative":0.90},"defect":{"true":0.98,"false":0.02},"recommend":{"true":0.03,"false":0.97}}
Documents: Saas escalator (hard) 100% right
{
  "year2_price_per_seat_month": 47.25,
  "year3_price_per_seat_month": 47.25,
  "year1_invoice": 58320.00,
  "year2_invoice": 61236.00,
  "addon_months_billed": 6,
  "addon_invoice": 38556.00,
  "year3_invoice": 134946.00,
  "year3_discount_percent": 15,
  "total_contract_value": 293058.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 100% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 2159.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 4700.00
}
Documents: Ticket SLA 92% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address", "invoice_pdf"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T15:30",
  "sla_due_utc": "2025-09-15T20:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 89% right
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "West",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 117B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.