Review · updated OCT 11

gpt-oss-20b review: a solid all-rounder

It scored 88 out of 100, #7 of 56. It solved 27 of 30 coding jobs and scored 87 on reading documents. Runs on a 16 GB graphics card or a Mac with 24 GB.

The short version
  • gpt-oss-20b is a free model from OpenAI that you can run on your own computer. In our tests it's a solid all-rounder: 88 out of 100, #7 of 56.
  • It solved 27 of 30 coding jobs and scored 87 on reading documents. On our hardest tasks it scored 69.
  • Runs on a 16 GB graphics card or a Mac with 24 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. gpt-oss-20b got 27 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. gpt-oss-20b scored 87; the best model scored 100.

TestScorePublic questionsSecret questions
Coding9010087
Reading documents879784
Decisions9510093

On the 18 hardest tasks (included in the scores above) it scored 69. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

gpt-oss-20b comes in one standard size, which needs about 12.1 GB of memory.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBYesstandard
Graphics card, 24 GBYesstandard
Graphics card, 32 GBYesstandard
Mac, 16 GBNo–
Mac, 24 GBYesstandard
Mac, 32 GBYesstandard
Mac, 48 GBYesstandard
Mac, 64 GBYesstandard
Mac, 96 GBYesstandard
Mac, 128 GBYesstandard

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run gpt-oss:20b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What gpt-oss-20b wrote for our public questions, word for word.

Coding: INI parser (hard) Worked
```python
import re
import string
from typing import Dict, Set, Tuple

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    """
    Parse a string in the described INI dialect and return a dictionary
    mapping section names to dictionaries of key/value pairs.
    """
    sections: Dict[str, Dict[str, str]] = {}
    current_section: str | None = None
    prev_key: str | None = None

    # regex for section headings
    section_regex = re.compile(r'^\s*\[([^\]]+)\]\s*$')

    def _add_section(name: str) -> Dict[str, str]:
        if name not in sections:
            sections[name] = {}
        return sections[name]

    # split a key/value line into key and value
    def _split_kv(line: str, lineno: int) -> Tuple[str, str]:
        pos_eq = line.find('=')
        pos_col = line.find(':')
        if pos_eq == -1 or (pos_col != -1 and pos_col < pos_eq):
            pos = pos_col
        else:
            pos = pos_eq
        if pos == -1:
            raise ValueError(f"Malformed line {lineno+1}: missing separator")
        key = line[:pos].strip()
        val_part = line[pos + 1 :].lstrip()

        # remove inline comment unless inside quotes
        in_quotes = False
        i = 0
        comment_start = None
        while i < len(val_part):
            ch = val_part[i]
            if ch == '"':
                in_quotes = not in_quotes
                i += 1
                continue
            if not in_quotes and ch in string.whitespace:
                if i + 1 < len(val_part) and val_part[i + 1] in ('#', ';'):
                    comment_start = i
                    break
            i += 1
        if comment_start is not None:
            val_part = val_part[:comment_start].rstrip()

        # quoted values keep inner text verbatim
        if val_part.startswith('"'):
            closing = None
            j = 1
            while j < len(val_part):
                if val_part[j] == '"':
                    closing = j
                    break
                j += 1
            if closing is None:
                raise ValueError(f"Unterminated quoted value at line {lineno+1}")
            value = val_part[1:closing]
        else:
            value = val_part
        return key, value

    for lineno, raw_line in enumerate(text.splitlines()):
        line = raw_line.rstrip('\n')
        stripped = line.strip()

        # skip empty or full-line comment lines
        if not stripped or stripped[0] in (';', '#'):
            continue

        # section header?
        m = section_regex.match(line)
        if m:
            current_section = m.group(1)
            _add_section(current_section)
            prev_key = None
            continue

        # continuation line (must follow a key line)
        if line[0].isspace():
            if prev_key is None:
                raise ValueError(f"Continuation line at line {lineno+1} without preceding key")
            prev_value = sections[current_section][prev_key]
            sections[current_section][prev_key] = prev_value + "\n" + line.strip()
            continue

        # must be a key/value line
        if current_section is None:
            raise ValueError(f"Key/value line outside any section at line {lineno+1}")
        key_raw, value = _split_kv(line, lineno)
        key = key_raw.lower()
        sections[current_section][key] = value
        prev_key = key

    # Separate DEFAULT section (which is not returned)
    defaults = sections.pop('DEFAULT', {})

    # Pre‑compute merged dictionaries (DEFAULT keys inherited)
    merged_sections: Dict[str, Dict[str, str]] = {}
    for sec_name, sec_vals in sections.items():
        merged = {}
        merged.update(defaults)      # inherited defaults
        merged.update(sec_vals)      # section overrides
        merged_sections[sec_name] = merged

    # Interpolation utilities
    placeholder_pattern = re.compile(r'\$\{([^}]+)\}')

    def _interpolate(
        sec_name: str,
        key: str,
        data: Dict[str, str],
        stack: Set[Tuple[str, str]],
    ) -> str:
        raw = data[key]
        if not isinstance(raw, str):
            return raw

        def replacer(match: re.Match[str]) -> str:
            token = match.group(1)
            if ':' in token:
                ref_sec, ref_key = token.split(':', 1)
            else:
                ref_sec, ref_key = sec_name, token
            ref_key = ref_key.lower()

            if ref_sec == 'DEFAULT':
                ref_data = defaults
            else:
                ref_data = merged_sections.get(ref_sec)
                if ref_data is None:
                    raise KeyError(f"Section {ref_sec!r} not found")
            if ref_key not in ref_data:
                raise KeyError(f"{ref_sec}:{ref_key!r} not found")

            if (ref_sec, ref_key) in stack:
                raise ValueError(f"Interpolation cycle detected for {ref_sec}:{ref_key}")

            return _interpolate(ref_sec, ref_key, ref_data, stack | {(ref_sec, ref_key)})

        result = raw
        while True:
            new, count = placeholder_pattern.subn(replacer, result)
            if count == 0:
                break
            result = new
        return result

    # Resolve all values and build final result
    result: Dict[str, Dict[str, str]] = {}
    for sec_name, sec_data in merged_sections.items():
        final_section: Dict[str, str] = {}
        for key in sec_data:
            final_section[key] = _interpolate(sec_name, key, sec_data, {(sec_name, key)})
        result[sec_name] = final_section

    return result
```
Coding: TTL LRU (hard) Worked
```python
import collections
from typing import Any, Callable, Dict

class TTLCache:
    """An LRU cache with time‑to‑live (TTL) expiration for each entry."""

    def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]) -> None:
        """
        Parameters
        ----------
        capacity : int
            Maximum number of items that may be stored. A capacity of `0` disables the cache.
        ttl : float
            Seconds until an entry expires after it is inserted or refreshed.
        clock : Callable[[], float]
            Zero‑argument callable returning the current time in seconds.
        """
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock

        # OrderedDict to preserve recency order: key -> (value, expiry_time)
        self._cache: collections.OrderedDict[Any, tuple] = collections.OrderedDict()

        # Statistics counters
        self._hits = 0
        self._misses = 0
        self._evictions = 0
        self._expirations = 0

    # ------------------------------------------------------------------ Public API

    def put(self, key: Any, value: Any) -> None:
        """Insert or replace an entry, refreshing its expiry and recency."""
        if self.capacity == 0:
            return

        now = self.clock()
        expiry = now + self.ttl

        if key in self._cache:
            # Replace existing entry
            self._cache[key] = (value, expiry)
            self._cache.move_to_end(key, last=True)
            return

        # Before inserting, clean up expired entries
        self._remove_expired(now)

        # If still full, evict the LRU entry
        if len(self._cache) >= self.capacity:
            lru_key, _ = self._cache.popitem(last=False)
            self._evictions += 1

        # Insert new entry
        self._cache[key] = (value, expiry)

    def get(self, key: Any, default: Any = None) -> Any:
        """Retrieve a value by key, or return `default` if missing or expired."""
        if self.capacity == 0:
            self._misses += 1
            return default

        now = self.clock()
        entry = self._cache.get(key)
        if entry is None:
            self._misses += 1
            return default

        value, expiry = entry
        if now >= expiry:
            # Entry expired
            del self._cache[key]
            self._expirations += 1
            self._misses += 1
            return default

        # Entry valid: update recency and return value
        self._cache.move_to_end(key, last=True)
        self._hits += 1
        return value

    def __len__(self) -> int:
        """Number of non‑expired entries currently in the cache."""
        now = self.clock()
        count = 0
        for _, (_, expiry) in self._cache.items():
            if now < expiry:
                count += 1
        return count

    def stats(self) -> Dict[str, int]:
        """Return a snapshot of cache statistics."""
        return {
            "hits": self._hits,
            "misses": self._misses,
            "evictions": self._evictions,
            "expirations": self._expirations,
        }

    # ------------------------------------------------------------------ Internals

    def _remove_expired(self, now: float) -> None:
        """Remove all expired entries and update expiration statistics."""
        to_remove = []
        for key, (_, expiry) in self._cache.items():
            if now >= expiry:
                to_remove.append(key)

        for key in to_remove:
            del self._cache[key]
            self._expirations += 1
```
Coding: Parse duration Worked
```python
import re
from typing import Optional

def parse_duration(s: str) -> int:
    """
    Convert a duration string into a total number of seconds.

    Supported units are 'h' (hours), 'm' (minutes), and 's' (seconds).
    The parts must appear in the order h → m → s and each may appear at most once.
    Leading and trailing whitespace, as well as spaces between parts, are allowed.
    Raises ValueError for invalid inputs.
    """
    if not isinstance(s, str):
        raise ValueError("Input must be a string")

    # Strip the string to detect an entirely blank input.
    stripped = s.strip()
    if not stripped:
        raise ValueError("Duration string cannot be empty or blank")

    # Regular expression that matches the entire string, allowing optional
    # whitespace and each unit only once and in the correct order.
    pattern = re.compile(
        r'^\s*(?:(?P<h>\d+)h)?\s*'   # hours, optional
        r'(?:(?P<m>\d+)m)?\s*'       # minutes, optional
        r'(?:(?P<s>\d+)s)?\s*$',
        re.IGNORECASE
    )

    match = pattern.match(s)
    if not match:
        raise ValueError(f"Invalid duration string: '{s}'")

    hours: Optional[str] = match.group('h')
    minutes: Optional[str] = match.group('m')
    seconds: Optional[str] = match.group('s')

    total = 0
    if hours is not None:
        total += int(hours) * 3600
    if minutes is not None:
        total += int(minutes) * 60
    if seconds is not None:
        total += int(seconds)

    return total
```
Coding: Summarize ranges Worked
```python
from typing import List

def summarize_ranges(nums: List[int]) -> str:
    """Return a comma‑separated string summarizing consecutive runs."""
    if not nums:
        return ""

    sorted_nums = sorted(set(nums))
    parts = []

    start = sorted_nums[0]
    prev = sorted_nums[0]

    for num in sorted_nums[1:]:
        if num == prev + 1:
            prev = num
            continue

        # Close the current run
        run_len = prev - start + 1
        if run_len >= 3:
            parts.append(f"{start}..{prev}")
        else:
            parts.extend(str(n) for n in range(start, prev + 1))

        # Start a new run
        start = prev = num

    # Close the final run
    run_len = prev - start + 1
    if run_len >= 3:
        parts.append(f"{start}..{prev}")
    else:
        parts.extend(str(n) for n in range(start, prev + 1))

    return ",".join(parts)
```
Coding: Top customers Worked
```python
from decimal import Decimal, InvalidOperation
from collections import defaultdict
from typing import Dict, List, Tuple


def _parse_amount(raw: str) -> Decimal:
    """
    Convert an amount string into a Decimal value.

    Accepts formats such as:
        "$1,234.50", "1234.5", "-$5.00", "$1,000"
    """
    if raw is None:
        raise ValueError("No amount provided")
    text = raw.strip()
    if not text:
        raise ValueError("Empty amount")

    sign = 1
    if text and text[0] == '-':
        sign = -1
        text = text[1:].strip()

    if text.startswith('$'):
        text = text[1:]

    # Remove thousands separators
    text = text.replace(',', '')

    try:
        value = Decimal(text)
    except InvalidOperation as exc:
        raise ValueError(f"Invalid numeric value: {raw}") from exc

    return sign * value


def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    """
    Return the top `n` customers by total transaction amount.

    Parameters
    ----------
    rows : List[Dict]
        Each dict must contain a 'customer' and an 'amount' key.
    n : int
        Number of top customers to return.
    """
    totals: Dict[str, Decimal] = defaultdict(Decimal)
    first_name: Dict[str, str] = {}

    for row in rows:
        # Extract and validate customer name
        raw_name = row.get('customer')
        if raw_name is None:
            continue
        name = raw_name.strip()
        if not name:
            continue

        # Parse amount
        try:
            amount = _parse_amount(row.get('amount'))
        except (ValueError, TypeError):
            continue

        key = name.lower()
        if key not in first_name:
            first_name[key] = name

        totals[key] += amount

    # Build result list
    result = [
        (first_name[key], float(totals[key].quantize(Decimal('0.01'))))
        for key in totals
    ]

    # Sort by total descending, then name ascending
    result.sort(key=lambda x: (-x[1], x[0]))

    return result[:n]
```
Coding: Slugify Worked
```python
import re
import unicodedata
from typing import List


def slugify(title: str, max_len: int = 50) -> str:
    """
    Build a URL slug from the given title.

    Parameters
    ----------
    title : str
        The original title string.
    max_len : int, optional
        Maximum length of the resulting slug (default 50).

    Returns
    -------
    str
        The slugified string.
    """
    if max_len <= 0:
        return ""

    # 1. Transliterate accented characters to ASCII and drop non‑ASCII.
    #    Normalization decomposes characters; encoding to ASCII with 'ignore'
    #    removes diacritics and any remaining non‑ASCII.
    ascii_title = unicodedata.normalize("NFKD", title).encode("ascii", "ignore").decode("ascii")

    # 2. Lowercase.
    ascii_title = ascii_title.lower()

    # 3. Replace runs of non a‑z0‑9 with a single hyphen.
    slug = re.sub(r"[^a-z0-9]+", "-", ascii_title)

    # Strip leading/trailing hyphens.
    slug = slug.strip("-")

    # If slug is empty after cleaning, return it.
    if not slug:
        return ""

    # 4. Truncate without cutting a word.
    if len(slug) <= max_len:
        return slug

    words: List[str] = slug.split("-")
    kept_words: List[str] = []
    current_len = 0

    for word in words:
        # Determine length if this word is added.
        add_len = len(word) if not kept_words else 1 + len(word)  # 1 for hyphen
        if current_len + add_len <= max_len:
            kept_words.append(word)
            current_len += add_len
        else:
            break

    if kept_words:
        truncated = "-".join(kept_words)
    else:
        # 5. If even the first word is too long, hard‑cut to max_len.
        truncated = slug[:max_len]

    # Ensure the result never ends with a hyphen.
    return truncated.rstrip("-")
```
Coding: Token bucket Worked
```python
from typing import Callable


class TokenBucket:
    """
    A token bucket rate limiter.

    Parameters
    ----------
    capacity : float
        The maximum number of tokens the bucket can hold.
    refill_per_sec : float
        The rate at which tokens are added to the bucket per second.
    clock : Callable[[], float]
        A zero-argument callable that returns the current time in seconds.
        The implementation must never call time.time() directly.
    """

    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self.capacity = float(capacity)
        self.refill_per_sec = float(refill_per_sec)
        self.clock = clock

        # The bucket starts full.
        self._tokens = self.capacity
        # The last time we updated the token count.
        self._last = self.clock()

    def _refill(self) -> None:
        """Replenish tokens based on elapsed time."""
        now = self.clock()
        elapsed = now - self._last
        if elapsed > 0:
            added = elapsed * self.refill_per_sec
            if added > 0:
                self._tokens = min(self.capacity, self._tokens + added)
        self._last = now

    def allow(self, cost: float = 1) -> bool:
        """
        Attempt to consume `cost` tokens from the bucket.

        Returns
        -------
        bool
            True if the bucket had enough tokens and the cost was deducted;
            False otherwise.
        """
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        """
        The current number of tokens in the bucket.

        Returns
        -------
        float
            The current token count after refilling.
        """
        self._refill()
        return self._tokens
```
Decisions: Refund window (hard) 100% right
{"outcome":{"full_refund":0.0,"partial_refund":0.0,"store_credit":1.0,"nothing":0.0},"defective":{"true":0.0,"false":1.0}}
Decisions: Incident matrix (hard) 100% right
{"severity":{"0":0,"1":1,"2":0,"3":0},"page":{"true":0,"false":1}}
Decisions: Tool followup (hard) 100% right
{"tool":{"web_search":0,"calculator":0,"calendar":0.97,"email":0,"none":0.03},"confirm":{"true":0.99,"false":0.01}}
Decisions: Legit security alert (hard) 100% right
{"phishing":{"true":0.05,"false":0.95},"action_needed":{"true":0.02,"false":0.98}}
Decisions: Meeting slot (hard) 100% right
{"slot":{"A":0,"B":1,"C":0,"D":0},"raj_last":{"true":1,"false":0}}
Decisions: Review mixed (hard) 100% right
{"hardware":{"true":0.99,"false":0.01},"support":{"true":0.99,"false":0.01}}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0.0,
    "technical": 1.0,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  },
  "outage": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.98,
    "technical": 0.005,
    "account": 0.005,
    "shipping": 0.005,
    "sales": 0.005
  },
  "refund": {
    "true": 0.99,
    "false": 0.01
  },
  "tone": {
    "frustrated": 0.02,
    "calm": 0.98
  }
}
Decisions: Moderation doxxing 100% right
{"policy":{"none":0.05,"harassment":0.95,"hate":0.0,"spam":0.0,"self_harm":0.0},"personal_info":{"true":1.0,"false":0.0}}
Decisions: Route calendar 100% right
{"tool":{"calendar":0.99,"none":0.01,"web_search":0.0,"calculator":0.0,"email":0.0},"confirm":{"true":1.0,"false":0.0}}
Decisions: Doc invoice missing due 100% right
{"doc_type":{"invoice":1.0,"resume":0.0,"contract":0.0,"bank_statement":0.0,"other":0.0},"missing_due_date":{"true":1.0,"false":0.0}}
Decisions: Phishing paypal 100% right
{"phishing":{"true":0.99,"false":0.01},"risk":{"0":0.01,"1":0.01,"2":0.01,"3":0.97}}
Decisions: Pii ssn email 100% right
{"data_kind":{"none":0.0,"contact":0.05,"financial":0.1,"government_id":0.85,"health":0.0},"sensitive":{"true":0.99,"false":0.01}}
Decisions: Review mixed 100% right
{"sentiment":{"positive":0.05,"neutral":0.15,"negative":0.8},"defect":{"true":0.99,"false":0.01},"recommend":{"true":0.01,"false":0.99}}
Documents: Saas escalator (hard) 100% right
{"year2_price_per_seat_month":47.25,"year3_price_per_seat_month":47.25,"year1_invoice":58320.00,"year2_invoice":61236.00,"addon_months_billed":6,"addon_invoice":38556.00,"year3_invoice":134946.00,"year3_discount_percent":15,"total_contract_value":293058.00,"contract_end_date":"2027-02-28"}
Documents: Expense thread 92% right
{"employee_id":"EMP-20417","destination_city":"Lisbon","trip_start":"2025-02-24","trip_end":"2025-02-27","approved_items":[{"date":"2025-02-24","category":"airfare","amount_usd":1184.60},{"date":"2025-02-24","category":"ground_transport","amount_usd":38.88},{"date":"2025-02-25","category":"meals","amount_usd":229.43},{"date":"2025-02-26","category":"lodging","amount_usd":466.56},{"date":"2025-02-27","category":"ground_transport","amount_usd":44.82}],"rejected_item_count":1,"per_diem_days":3,"per_diem_usd":195.00,"total_reimbursable_usd":2159.29,"approver_email":"priya.raman@corvane.com"}
Documents: Lease amendment 100% right
{"tenants":["Marcus Lin","Sofia Lin"],"landlord":"Ridgeline Property Group LLC","zip":"97205","lease_end":"2025-11-30","original_monthly_rent":2150,"monthly_rent_from_2025_06_01":2236,"late_fee_from_2025_06_01":111.8,"security_deposit":2150,"total_pet_deposits":800,"total_monthly_payment_july_2025":2306,"move_in_payment":4700}
Documents: Ticket SLA 92% right
{"ticket_id":"48213","account_id":"ACC-7731","open_issue":"inventory_sync","resolved_issues":["billing_address","invoice_pdf"],"affected_orders":["SO-99812","SO-99820","SO-99827"],"priority":"P2","sla_due_local":"2025-09-15T15:30","sla_due_utc":"2025-09-15T20:30:00Z","reissued_invoice":"INV-2025-0812"}
Documents: Sales footnotes 100% right
{"q3_total_usd":15346000,"q2_total_usd":14464000,"q2_central_originally_reported_usd":3047000,"q2_to_q3_change_pct":6.1,"top_region_q3":"East","fastest_growing_region_q1_to_q3":"International","regions_declining_q2_to_q3":["East"],"international_q3_organic_usd":1731000,"west_excluding_mountain_q3_usd":4201000}

Size: 21B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.