- gpt-oss-20b is a free model from OpenAI that you can run on your own computer. In our tests it's a solid all-rounder: 88 out of 100, #7 of 56.
- It solved 27 of 30 coding jobs and scored 87 on reading documents. On our hardest tasks it scored 69.
- Runs on a 16 GB graphics card or a Mac with 24 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. gpt-oss-20b got 27 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. gpt-oss-20b scored 87; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 90 | 100 | 87 |
| Reading documents | 87 | 97 | 84 |
| Decisions | 95 | 100 | 93 |
On the 18 hardest tasks (included in the scores above) it scored 69. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
gpt-oss-20b comes in one standard size, which needs about 12.1 GB of memory.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | Yes | standard |
| Graphics card, 24 GB | Yes | standard |
| Graphics card, 32 GB | Yes | standard |
| Mac, 16 GB | No | – |
| Mac, 24 GB | Yes | standard |
| Mac, 32 GB | Yes | standard |
| Mac, 48 GB | Yes | standard |
| Mac, 64 GB | Yes | standard |
| Mac, 96 GB | Yes | standard |
| Mac, 128 GB | Yes | standard |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run gpt-oss:20b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What gpt-oss-20b wrote for our public questions, word for word.
Coding: INI parser (hard) Worked
```python
import re
import string
from typing import Dict, Set, Tuple
def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
"""
Parse a string in the described INI dialect and return a dictionary
mapping section names to dictionaries of key/value pairs.
"""
sections: Dict[str, Dict[str, str]] = {}
current_section: str | None = None
prev_key: str | None = None
# regex for section headings
section_regex = re.compile(r'^\s*\[([^\]]+)\]\s*$')
def _add_section(name: str) -> Dict[str, str]:
if name not in sections:
sections[name] = {}
return sections[name]
# split a key/value line into key and value
def _split_kv(line: str, lineno: int) -> Tuple[str, str]:
pos_eq = line.find('=')
pos_col = line.find(':')
if pos_eq == -1 or (pos_col != -1 and pos_col < pos_eq):
pos = pos_col
else:
pos = pos_eq
if pos == -1:
raise ValueError(f"Malformed line {lineno+1}: missing separator")
key = line[:pos].strip()
val_part = line[pos + 1 :].lstrip()
# remove inline comment unless inside quotes
in_quotes = False
i = 0
comment_start = None
while i < len(val_part):
ch = val_part[i]
if ch == '"':
in_quotes = not in_quotes
i += 1
continue
if not in_quotes and ch in string.whitespace:
if i + 1 < len(val_part) and val_part[i + 1] in ('#', ';'):
comment_start = i
break
i += 1
if comment_start is not None:
val_part = val_part[:comment_start].rstrip()
# quoted values keep inner text verbatim
if val_part.startswith('"'):
closing = None
j = 1
while j < len(val_part):
if val_part[j] == '"':
closing = j
break
j += 1
if closing is None:
raise ValueError(f"Unterminated quoted value at line {lineno+1}")
value = val_part[1:closing]
else:
value = val_part
return key, value
for lineno, raw_line in enumerate(text.splitlines()):
line = raw_line.rstrip('\n')
stripped = line.strip()
# skip empty or full-line comment lines
if not stripped or stripped[0] in (';', '#'):
continue
# section header?
m = section_regex.match(line)
if m:
current_section = m.group(1)
_add_section(current_section)
prev_key = None
continue
# continuation line (must follow a key line)
if line[0].isspace():
if prev_key is None:
raise ValueError(f"Continuation line at line {lineno+1} without preceding key")
prev_value = sections[current_section][prev_key]
sections[current_section][prev_key] = prev_value + "\n" + line.strip()
continue
# must be a key/value line
if current_section is None:
raise ValueError(f"Key/value line outside any section at line {lineno+1}")
key_raw, value = _split_kv(line, lineno)
key = key_raw.lower()
sections[current_section][key] = value
prev_key = key
# Separate DEFAULT section (which is not returned)
defaults = sections.pop('DEFAULT', {})
# Pre‑compute merged dictionaries (DEFAULT keys inherited)
merged_sections: Dict[str, Dict[str, str]] = {}
for sec_name, sec_vals in sections.items():
merged = {}
merged.update(defaults) # inherited defaults
merged.update(sec_vals) # section overrides
merged_sections[sec_name] = merged
# Interpolation utilities
placeholder_pattern = re.compile(r'\$\{([^}]+)\}')
def _interpolate(
sec_name: str,
key: str,
data: Dict[str, str],
stack: Set[Tuple[str, str]],
) -> str:
raw = data[key]
if not isinstance(raw, str):
return raw
def replacer(match: re.Match[str]) -> str:
token = match.group(1)
if ':' in token:
ref_sec, ref_key = token.split(':', 1)
else:
ref_sec, ref_key = sec_name, token
ref_key = ref_key.lower()
if ref_sec == 'DEFAULT':
ref_data = defaults
else:
ref_data = merged_sections.get(ref_sec)
if ref_data is None:
raise KeyError(f"Section {ref_sec!r} not found")
if ref_key not in ref_data:
raise KeyError(f"{ref_sec}:{ref_key!r} not found")
if (ref_sec, ref_key) in stack:
raise ValueError(f"Interpolation cycle detected for {ref_sec}:{ref_key}")
return _interpolate(ref_sec, ref_key, ref_data, stack | {(ref_sec, ref_key)})
result = raw
while True:
new, count = placeholder_pattern.subn(replacer, result)
if count == 0:
break
result = new
return result
# Resolve all values and build final result
result: Dict[str, Dict[str, str]] = {}
for sec_name, sec_data in merged_sections.items():
final_section: Dict[str, str] = {}
for key in sec_data:
final_section[key] = _interpolate(sec_name, key, sec_data, {(sec_name, key)})
result[sec_name] = final_section
return result
``` Coding: TTL LRU (hard) Worked
```python
import collections
from typing import Any, Callable, Dict
class TTLCache:
"""An LRU cache with time‑to‑live (TTL) expiration for each entry."""
def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]) -> None:
"""
Parameters
----------
capacity : int
Maximum number of items that may be stored. A capacity of `0` disables the cache.
ttl : float
Seconds until an entry expires after it is inserted or refreshed.
clock : Callable[[], float]
Zero‑argument callable returning the current time in seconds.
"""
self.capacity = capacity
self.ttl = ttl
self.clock = clock
# OrderedDict to preserve recency order: key -> (value, expiry_time)
self._cache: collections.OrderedDict[Any, tuple] = collections.OrderedDict()
# Statistics counters
self._hits = 0
self._misses = 0
self._evictions = 0
self._expirations = 0
# ------------------------------------------------------------------ Public API
def put(self, key: Any, value: Any) -> None:
"""Insert or replace an entry, refreshing its expiry and recency."""
if self.capacity == 0:
return
now = self.clock()
expiry = now + self.ttl
if key in self._cache:
# Replace existing entry
self._cache[key] = (value, expiry)
self._cache.move_to_end(key, last=True)
return
# Before inserting, clean up expired entries
self._remove_expired(now)
# If still full, evict the LRU entry
if len(self._cache) >= self.capacity:
lru_key, _ = self._cache.popitem(last=False)
self._evictions += 1
# Insert new entry
self._cache[key] = (value, expiry)
def get(self, key: Any, default: Any = None) -> Any:
"""Retrieve a value by key, or return `default` if missing or expired."""
if self.capacity == 0:
self._misses += 1
return default
now = self.clock()
entry = self._cache.get(key)
if entry is None:
self._misses += 1
return default
value, expiry = entry
if now >= expiry:
# Entry expired
del self._cache[key]
self._expirations += 1
self._misses += 1
return default
# Entry valid: update recency and return value
self._cache.move_to_end(key, last=True)
self._hits += 1
return value
def __len__(self) -> int:
"""Number of non‑expired entries currently in the cache."""
now = self.clock()
count = 0
for _, (_, expiry) in self._cache.items():
if now < expiry:
count += 1
return count
def stats(self) -> Dict[str, int]:
"""Return a snapshot of cache statistics."""
return {
"hits": self._hits,
"misses": self._misses,
"evictions": self._evictions,
"expirations": self._expirations,
}
# ------------------------------------------------------------------ Internals
def _remove_expired(self, now: float) -> None:
"""Remove all expired entries and update expiration statistics."""
to_remove = []
for key, (_, expiry) in self._cache.items():
if now >= expiry:
to_remove.append(key)
for key in to_remove:
del self._cache[key]
self._expirations += 1
``` Coding: Parse duration Worked
```python
import re
from typing import Optional
def parse_duration(s: str) -> int:
"""
Convert a duration string into a total number of seconds.
Supported units are 'h' (hours), 'm' (minutes), and 's' (seconds).
The parts must appear in the order h → m → s and each may appear at most once.
Leading and trailing whitespace, as well as spaces between parts, are allowed.
Raises ValueError for invalid inputs.
"""
if not isinstance(s, str):
raise ValueError("Input must be a string")
# Strip the string to detect an entirely blank input.
stripped = s.strip()
if not stripped:
raise ValueError("Duration string cannot be empty or blank")
# Regular expression that matches the entire string, allowing optional
# whitespace and each unit only once and in the correct order.
pattern = re.compile(
r'^\s*(?:(?P<h>\d+)h)?\s*' # hours, optional
r'(?:(?P<m>\d+)m)?\s*' # minutes, optional
r'(?:(?P<s>\d+)s)?\s*$',
re.IGNORECASE
)
match = pattern.match(s)
if not match:
raise ValueError(f"Invalid duration string: '{s}'")
hours: Optional[str] = match.group('h')
minutes: Optional[str] = match.group('m')
seconds: Optional[str] = match.group('s')
total = 0
if hours is not None:
total += int(hours) * 3600
if minutes is not None:
total += int(minutes) * 60
if seconds is not None:
total += int(seconds)
return total
``` Coding: Summarize ranges Worked
```python
from typing import List
def summarize_ranges(nums: List[int]) -> str:
"""Return a comma‑separated string summarizing consecutive runs."""
if not nums:
return ""
sorted_nums = sorted(set(nums))
parts = []
start = sorted_nums[0]
prev = sorted_nums[0]
for num in sorted_nums[1:]:
if num == prev + 1:
prev = num
continue
# Close the current run
run_len = prev - start + 1
if run_len >= 3:
parts.append(f"{start}..{prev}")
else:
parts.extend(str(n) for n in range(start, prev + 1))
# Start a new run
start = prev = num
# Close the final run
run_len = prev - start + 1
if run_len >= 3:
parts.append(f"{start}..{prev}")
else:
parts.extend(str(n) for n in range(start, prev + 1))
return ",".join(parts)
``` Coding: Top customers Worked
```python
from decimal import Decimal, InvalidOperation
from collections import defaultdict
from typing import Dict, List, Tuple
def _parse_amount(raw: str) -> Decimal:
"""
Convert an amount string into a Decimal value.
Accepts formats such as:
"$1,234.50", "1234.5", "-$5.00", "$1,000"
"""
if raw is None:
raise ValueError("No amount provided")
text = raw.strip()
if not text:
raise ValueError("Empty amount")
sign = 1
if text and text[0] == '-':
sign = -1
text = text[1:].strip()
if text.startswith('$'):
text = text[1:]
# Remove thousands separators
text = text.replace(',', '')
try:
value = Decimal(text)
except InvalidOperation as exc:
raise ValueError(f"Invalid numeric value: {raw}") from exc
return sign * value
def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
"""
Return the top `n` customers by total transaction amount.
Parameters
----------
rows : List[Dict]
Each dict must contain a 'customer' and an 'amount' key.
n : int
Number of top customers to return.
"""
totals: Dict[str, Decimal] = defaultdict(Decimal)
first_name: Dict[str, str] = {}
for row in rows:
# Extract and validate customer name
raw_name = row.get('customer')
if raw_name is None:
continue
name = raw_name.strip()
if not name:
continue
# Parse amount
try:
amount = _parse_amount(row.get('amount'))
except (ValueError, TypeError):
continue
key = name.lower()
if key not in first_name:
first_name[key] = name
totals[key] += amount
# Build result list
result = [
(first_name[key], float(totals[key].quantize(Decimal('0.01'))))
for key in totals
]
# Sort by total descending, then name ascending
result.sort(key=lambda x: (-x[1], x[0]))
return result[:n]
``` Coding: Slugify Worked
```python
import re
import unicodedata
from typing import List
def slugify(title: str, max_len: int = 50) -> str:
"""
Build a URL slug from the given title.
Parameters
----------
title : str
The original title string.
max_len : int, optional
Maximum length of the resulting slug (default 50).
Returns
-------
str
The slugified string.
"""
if max_len <= 0:
return ""
# 1. Transliterate accented characters to ASCII and drop non‑ASCII.
# Normalization decomposes characters; encoding to ASCII with 'ignore'
# removes diacritics and any remaining non‑ASCII.
ascii_title = unicodedata.normalize("NFKD", title).encode("ascii", "ignore").decode("ascii")
# 2. Lowercase.
ascii_title = ascii_title.lower()
# 3. Replace runs of non a‑z0‑9 with a single hyphen.
slug = re.sub(r"[^a-z0-9]+", "-", ascii_title)
# Strip leading/trailing hyphens.
slug = slug.strip("-")
# If slug is empty after cleaning, return it.
if not slug:
return ""
# 4. Truncate without cutting a word.
if len(slug) <= max_len:
return slug
words: List[str] = slug.split("-")
kept_words: List[str] = []
current_len = 0
for word in words:
# Determine length if this word is added.
add_len = len(word) if not kept_words else 1 + len(word) # 1 for hyphen
if current_len + add_len <= max_len:
kept_words.append(word)
current_len += add_len
else:
break
if kept_words:
truncated = "-".join(kept_words)
else:
# 5. If even the first word is too long, hard‑cut to max_len.
truncated = slug[:max_len]
# Ensure the result never ends with a hyphen.
return truncated.rstrip("-")
``` Coding: Token bucket Worked
```python
from typing import Callable
class TokenBucket:
"""
A token bucket rate limiter.
Parameters
----------
capacity : float
The maximum number of tokens the bucket can hold.
refill_per_sec : float
The rate at which tokens are added to the bucket per second.
clock : Callable[[], float]
A zero-argument callable that returns the current time in seconds.
The implementation must never call time.time() directly.
"""
def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
self.capacity = float(capacity)
self.refill_per_sec = float(refill_per_sec)
self.clock = clock
# The bucket starts full.
self._tokens = self.capacity
# The last time we updated the token count.
self._last = self.clock()
def _refill(self) -> None:
"""Replenish tokens based on elapsed time."""
now = self.clock()
elapsed = now - self._last
if elapsed > 0:
added = elapsed * self.refill_per_sec
if added > 0:
self._tokens = min(self.capacity, self._tokens + added)
self._last = now
def allow(self, cost: float = 1) -> bool:
"""
Attempt to consume `cost` tokens from the bucket.
Returns
-------
bool
True if the bucket had enough tokens and the cost was deducted;
False otherwise.
"""
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
return False
@property
def tokens(self) -> float:
"""
The current number of tokens in the bucket.
Returns
-------
float
The current token count after refilling.
"""
self._refill()
return self._tokens
``` Decisions: Refund window (hard) 100% right
{"outcome":{"full_refund":0.0,"partial_refund":0.0,"store_credit":1.0,"nothing":0.0},"defective":{"true":0.0,"false":1.0}} Decisions: Incident matrix (hard) 100% right
{"severity":{"0":0,"1":1,"2":0,"3":0},"page":{"true":0,"false":1}} Decisions: Tool followup (hard) 100% right
{"tool":{"web_search":0,"calculator":0,"calendar":0.97,"email":0,"none":0.03},"confirm":{"true":0.99,"false":0.01}} Decisions: Legit security alert (hard) 100% right
{"phishing":{"true":0.05,"false":0.95},"action_needed":{"true":0.02,"false":0.98}} Decisions: Meeting slot (hard) 100% right
{"slot":{"A":0,"B":1,"C":0,"D":0},"raj_last":{"true":1,"false":0}} Decisions: Review mixed (hard) 100% right
{"hardware":{"true":0.99,"false":0.01},"support":{"true":0.99,"false":0.01}} Decisions: Support checkout down 100% right
{
"department": {
"billing": 0.0,
"technical": 1.0,
"account": 0.0,
"shipping": 0.0,
"sales": 0.0
},
"urgency": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
},
"outage": {
"true": 1.0,
"false": 0.0
}
} Decisions: Refund wrong plan 100% right
{
"department": {
"billing": 0.98,
"technical": 0.005,
"account": 0.005,
"shipping": 0.005,
"sales": 0.005
},
"refund": {
"true": 0.99,
"false": 0.01
},
"tone": {
"frustrated": 0.02,
"calm": 0.98
}
} Decisions: Moderation doxxing 100% right
{"policy":{"none":0.05,"harassment":0.95,"hate":0.0,"spam":0.0,"self_harm":0.0},"personal_info":{"true":1.0,"false":0.0}} Decisions: Route calendar 100% right
{"tool":{"calendar":0.99,"none":0.01,"web_search":0.0,"calculator":0.0,"email":0.0},"confirm":{"true":1.0,"false":0.0}} Decisions: Doc invoice missing due 100% right
{"doc_type":{"invoice":1.0,"resume":0.0,"contract":0.0,"bank_statement":0.0,"other":0.0},"missing_due_date":{"true":1.0,"false":0.0}} Decisions: Phishing paypal 100% right
{"phishing":{"true":0.99,"false":0.01},"risk":{"0":0.01,"1":0.01,"2":0.01,"3":0.97}} Decisions: Pii ssn email 100% right
{"data_kind":{"none":0.0,"contact":0.05,"financial":0.1,"government_id":0.85,"health":0.0},"sensitive":{"true":0.99,"false":0.01}} Decisions: Review mixed 100% right
{"sentiment":{"positive":0.05,"neutral":0.15,"negative":0.8},"defect":{"true":0.99,"false":0.01},"recommend":{"true":0.01,"false":0.99}} Documents: Saas escalator (hard) 100% right
{"year2_price_per_seat_month":47.25,"year3_price_per_seat_month":47.25,"year1_invoice":58320.00,"year2_invoice":61236.00,"addon_months_billed":6,"addon_invoice":38556.00,"year3_invoice":134946.00,"year3_discount_percent":15,"total_contract_value":293058.00,"contract_end_date":"2027-02-28"} Documents: Expense thread 92% right
{"employee_id":"EMP-20417","destination_city":"Lisbon","trip_start":"2025-02-24","trip_end":"2025-02-27","approved_items":[{"date":"2025-02-24","category":"airfare","amount_usd":1184.60},{"date":"2025-02-24","category":"ground_transport","amount_usd":38.88},{"date":"2025-02-25","category":"meals","amount_usd":229.43},{"date":"2025-02-26","category":"lodging","amount_usd":466.56},{"date":"2025-02-27","category":"ground_transport","amount_usd":44.82}],"rejected_item_count":1,"per_diem_days":3,"per_diem_usd":195.00,"total_reimbursable_usd":2159.29,"approver_email":"priya.raman@corvane.com"} Documents: Lease amendment 100% right
{"tenants":["Marcus Lin","Sofia Lin"],"landlord":"Ridgeline Property Group LLC","zip":"97205","lease_end":"2025-11-30","original_monthly_rent":2150,"monthly_rent_from_2025_06_01":2236,"late_fee_from_2025_06_01":111.8,"security_deposit":2150,"total_pet_deposits":800,"total_monthly_payment_july_2025":2306,"move_in_payment":4700} Documents: Ticket SLA 92% right
{"ticket_id":"48213","account_id":"ACC-7731","open_issue":"inventory_sync","resolved_issues":["billing_address","invoice_pdf"],"affected_orders":["SO-99812","SO-99820","SO-99827"],"priority":"P2","sla_due_local":"2025-09-15T15:30","sla_due_utc":"2025-09-15T20:30:00Z","reissued_invoice":"INV-2025-0812"} Documents: Sales footnotes 100% right
{"q3_total_usd":15346000,"q2_total_usd":14464000,"q2_central_originally_reported_usd":3047000,"q2_to_q3_change_pct":6.1,"top_region_q3":"East","fastest_growing_region_q1_to_q3":"International","regions_declining_q2_to_q3":["East"],"international_q3_organic_usd":1731000,"west_excluding_mountain_q3_usd":4201000} Size: 21B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.