- gpt-oss-120b is a free model from OpenAI that you can run on your own computer. In our tests it's one of the best local models we've tested: 96 out of 100, #2 of 56.
- It solved 28 of 30 coding jobs and scored 99 on reading documents. On our hardest tasks it scored 94.
- Runs on a Mac with 96 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. gpt-oss-120b got 28 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. gpt-oss-120b scored 99; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 93 | 100 | 91 |
| Reading documents | 99 | 96 | 100 |
| Decisions | 99 | 100 | 99 |
On the 18 hardest tasks (included in the scores above) it scored 94. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
gpt-oss-120b comes in one standard size, which needs about 63.3 GB of memory.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | No | – |
| Graphics card, 32 GB | No | – |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | No | – |
| Mac, 48 GB | No | – |
| Mac, 64 GB | No | – |
| Mac, 96 GB | Yes | standard |
| Mac, 128 GB | Yes | standard |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
Its actual answers
What gpt-oss-120b wrote for our public questions, word for word.
Coding: INI parser (hard) Worked
```python
import re
from typing import Dict, Set
def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
"""
Parse an INI file according to the dialect described in the prompt.
Returns a mapping ``section -> {key -> value}`` where *section* does not
include the special ``DEFAULT`` section (its keys are inherited by all
other sections). All keys are lower‑cased.
"""
# --------------------------------------------------------------------- #
# 1. Lexical parsing – sections, keys, continuations, comments, quoting #
# --------------------------------------------------------------------- #
section_re = re.compile(r'^\s*\[(.+?)\]\s*$')
var_re = re.compile(r'\$\{([^}]+)\}')
default: Dict[str, str] = {}
sections: Dict[str, Dict[str, str]] = {}
current: str | None = None # name of the current section
last_key: Dict[str, str | None] = {} # last key seen per section
for lineno, raw_line in enumerate(text.splitlines(), start=1):
line = raw_line.rstrip("\r\n")
# ignore empty lines and full‑line comments
stripped = line.lstrip()
if not stripped or stripped[0] in (";", "#"):
continue
# ----------------------------------------------------------------- #
# Section header
# ----------------------------------------------------------------- #
m = section_re.match(line)
if m:
name = m.group(1)
current = name
if name != "DEFAULT":
sections.setdefault(name, {})
last_key[current] = None
continue
# ----------------------------------------------------------------- #
# Continuation line (must follow a key line)
# ----------------------------------------------------------------- #
if line and line[0].isspace():
if current is None:
raise ValueError(f"Continuation line before any section (line {lineno})")
key = last_key.get(current)
if not key:
raise ValueError(f"Continuation line without preceding key (line {lineno})")
cont = line.strip()
target = default if current == "DEFAULT" else sections[current]
target[key] = f"{target[key]}\n{cont}"
continue
# ----------------------------------------------------------------- #
# Key/value line
# ----------------------------------------------------------------- #
if current is None:
raise ValueError(f"Key/value line before any section (line {lineno})")
# locate first ':' or '='
eq = line.find("=")
colon = line.find(":")
if eq == -1 and colon == -1:
raise ValueError(f"Malformed line (no ':' or '=') (line {lineno})")
if eq != -1 and colon != -1:
pos = min(eq, colon)
else:
pos = eq if eq != -1 else colon
raw_key = line[:pos]
raw_val = line[pos + 1 :]
key = raw_key.strip().lower()
# ------------------------------------------------------------- #
# Value processing – quoting, stripping, inline comments
# ------------------------------------------------------------- #
val = raw_val.lstrip()
if val.startswith('"'):
# quoted value – keep everything inside the quotes verbatim
end = val.find('"', 1)
if end == -1:
raise ValueError(f"Unclosed double quote (line {lineno})")
processed = val[1:end] # no stripping
else:
# inline comment detection (space + ; or #)
comment_pos = None
for marker in (" ;", " #"):
idx = val.find(marker)
if idx != -1:
comment_pos = idx if comment_pos is None else min(comment_pos, idx)
processed = val[:comment_pos] if comment_pos is not None else val
processed = processed.strip()
# store the raw (still possibly containing ${...}) value
target = default if current == "DEFAULT" else sections[current]
target[key] = processed
last_key[current] = key
# --------------------------------------------------------------------- #
# 2. Interpolation – recursive substitution with cycle detection
# --------------------------------------------------------------------- #
def get_raw(sec: str, k: str) -> str:
"""Return the raw (un‑interpolated) value for *k* in *sec*,
falling back to DEFAULT if needed."""
if sec != "DEFAULT" and k in sections.get(sec, {}):
return sections[sec][k]
if k in default:
return default[k]
raise KeyError(f"Key '{k}' not found in section '{sec}'")
def resolve(sec: str, k: str, seen: Set[tuple[str, str]]) -> str:
"""Resolve *k* in *sec* with interpolation. *seen* tracks the
current resolution chain to detect cycles."""
if (sec, k) in seen:
raise ValueError(f"Interpolation cycle involving {sec}:{k}")
seen.add((sec, k))
raw = get_raw(sec, k)
def repl(m: re.Match) -> str:
inner = m.group(1)
if ":" in inner:
other_sec, other_key = inner.split(":", 1)
else:
other_sec, other_key = sec, inner
return resolve(other_sec, other_key, seen)
# Perform substitution; repeat until no more variables appear
result = raw
while True:
if not var_re.search(result):
break
result = var_re.sub(repl, result)
seen.remove((sec, k))
return result
# --------------------------------------------------------------------- #
# 3. Build final mapping – inherit DEFAULT, apply interpolation
# --------------------------------------------------------------------- #
final: Dict[str, Dict[str, str]] = {}
for sec, kv in sections.items():
if sec == "DEFAULT":
continue
merged_keys = set(default) | set(kv)
resolved_section: Dict[str, str] = {}
for k in merged_keys:
resolved_section[k] = resolve(sec, k, set())
final[sec] = resolved_section
return final
``` Coding: TTL LRU (hard) Worked
```python
from __future__ import annotations
import collections
from collections import OrderedDict
from typing import Callable, Dict, Generic, Hashable, Iterable, Tuple, TypeVar
K = TypeVar("K", bound=Hashable)
V = TypeVar("V")
TClock = Callable[[], float]
class TTLCache(Generic[K, V]):
"""Least‑Recently‑Used cache with per‑entry TTL expiration.
Parameters
----------
capacity: int
Maximum number of entries that can be stored. ``0`` disables storage.
ttl: float
Time‑to‑live for each entry in seconds.
clock: Callable[[], float]
A zero‑argument function returning the current time in seconds.
"""
__slots__ = (
"_capacity",
"_ttl",
"_clock",
"_store",
"_hits",
"_misses",
"_evictions",
"_expirations",
)
def __init__(self, capacity: int, ttl: float, clock: TClock) -> None:
if capacity < 0:
raise ValueError("capacity must be non‑negative")
if ttl < 0:
raise ValueError("ttl must be non‑negative")
self._capacity: int = capacity
self._ttl: float = ttl
self._clock: TClock = clock
self._store: "OrderedDict[K, Tuple[V, float]]" = OrderedDict()
self._hits = 0
self._misses = 0
self._evictions = 0
self._expirations = 0
# ------------------------------------------------------------------ helpers
def _now(self) -> float:
return self._clock()
def _remove_expired(self) -> None:
"""Delete all expired entries, updating the expiration counter."""
now = self._now()
# list() copy to avoid mutation during iteration
for key, (_, expiry) in list(self._store.items()):
if now >= expiry:
self._store.pop(key, None)
self._expirations += 1
# -------------------------------------------------------------------- public
def put(self, key: K, value: V) -> None:
"""Insert or replace an entry.
Replacements refresh the expiry time and mark the entry as most‑recently
used. When a new key would overflow the capacity, expired entries are
removed first; if the cache is still full, the least‑recently‑used entry
is evicted.
"""
if self._capacity == 0:
return
now = self._now()
if key in self._store:
# Replace existing entry.
self._store[key] = (value, now + self._ttl)
self._store.move_to_end(key) # mark as most recent
return
# New key – possibly need to free space.
self._remove_expired()
if len(self._store) >= self._capacity:
# Evict least‑recently‑used entry.
evicted_key, _ = self._store.popitem(last=False)
self._evictions += 1
# (evicted_key is discarded)
self._store[key] = (value, now + self._ttl)
# New insert is automatically at the end (most recent).
def get(self, key: K, default: V | None = None) -> V | None:
"""Retrieve a value if present and not expired.
A successful lookup marks the entry as most‑recently‑used.
Expired entries are removed, counted as expirations, and result in a miss.
"""
if self._capacity == 0:
self._misses += 1
return default
now = self._now()
entry = self._store.get(key)
if entry is None:
self._misses += 1
return default
value, expiry = entry
if now >= expiry:
# Expired – remove and count as miss+expiration.
self._store.pop(key, None)
self._expirations += 1
self._misses += 1
return default
# Hit – update recency.
self._hits += 1
self._store.move_to_end(key)
return value
def __len__(self) -> int:
"""Number of unexpired entries at the current time (does not mutate)."""
now = self._now()
return sum(1 for _, (_, expiry) in self._store.items() if now < expiry)
def stats(self) -> Dict[str, int]:
"""Return a dictionary with cache statistics."""
return {
"hits": self._hits,
"misses": self._misses,
"evictions": self._evictions,
"expirations": self._expirations,
}
``` Coding: Parse duration Worked
```python
import re
def parse_duration(s: str) -> int:
"""
Convert a duration string to a total number of seconds.
Supported units (case‑insensitive):
h – hours
m – minutes
s – seconds
Rules
-----
* Each part is a non‑negative integer followed by a unit, e.g. "1h30m".
* Whitespace may appear anywhere and is ignored.
* Units must appear in the order: hours, minutes, seconds; each at most once.
* Invalid input raises ValueError.
Parameters
----------
s: str
The duration string.
Returns
-------
int
Total duration expressed in seconds.
"""
if not isinstance(s, str):
raise TypeError("Input must be a string")
# Remove all whitespace characters
s_clean = ''.join(s.split())
if not s_clean:
raise ValueError("Duration string is empty or blank")
i = 0
last_order = 0 # 0 = none, 1 = h, 2 = m, 3 = s
total_seconds = 0
unit_to_order = {'h': 1, 'm': 2, 's': 3}
unit_to_seconds = {'h': 3600, 'm': 60, 's': 1}
while i < len(s_clean):
# --- parse the integer part ---
start = i
while i < len(s_clean) and s_clean[i].isdigit():
i += 1
if start == i: # no digits found
raise ValueError(f"Expected number at position {start}")
number = int(s_clean[start:i])
# --- parse the unit character ---
if i >= len(s_clean):
raise ValueError(f"Missing unit after number at position {start}")
unit_char = s_clean[i].lower()
i += 1
if unit_char not in unit_to_order:
raise ValueError(f"Unknown unit '{unit_char}' at position {i-1}")
order = unit_to_order[unit_char]
if order <= last_order: # repeated or out‑of‑order unit
raise ValueError("Units must appear once and in the order h, m, s")
total_seconds += number * unit_to_seconds[unit_char]
last_order = order
return total_seconds
``` Coding: Summarize ranges Worked
```python
from __future__ import annotations
def summarize_ranges(nums: list[int]) -> str:
"""Return a comma‑separated string that summarizes consecutive integer runs.
* The input list is sorted and deduplicated.
* Runs of three or more consecutive numbers become “a..b”.
* Runs of one or two numbers are emitted individually.
* An empty input returns an empty string.
"""
if not nums:
return ""
# Sort and remove duplicates.
sorted_unique = sorted(set(nums))
parts: list[str] = []
i = 0
n = len(sorted_unique)
while i < n:
start = sorted_unique[i]
j = i
# Extend the run while the next number is exactly one larger.
while j + 1 < n and sorted_unique[j + 1] == sorted_unique[j] + 1:
j += 1
run_len = j - i + 1
if run_len >= 3:
# Collapse long runs into "a..b".
parts.append(f"{start}..{sorted_unique[j]}")
else:
# Emit each element of short runs individually.
for k in range(i, j + 1):
parts.append(str(sorted_unique[k]))
i = j + 1
return ",".join(parts)
``` Coding: Top customers Worked
```python
import re
from typing import List, Dict, Tuple, Optional
def _parse_amount(value: str) -> Optional[float]:
"""
Convert a monetary string to a float.
Accepted formats include "$1,234.50", "1234.5", "-$5.00", "$1,000", etc.
Returns ``None`` if the string cannot be parsed.
"""
if not isinstance(value, str):
return None
# Remove surrounding whitespace
s = value.strip()
# Quick reject empty strings
if not s:
return None
# Remove dollar signs and commas
s = s.replace("$", "").replace(",", "")
# After cleanup, the string should represent a valid float
try:
return float(s)
except ValueError:
return None
def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
"""
Compute the top *n* customers by total transaction amount.
Parameters
----------
rows : list[dict]
Each dict must contain at least the keys ``"customer"`` and ``"amount"``.
``"customer"`` is a string name (may contain surrounding whitespace).
``"amount"`` is a string that may include a leading ``$``, commas, and
an optional leading ``-`` for refunds.
n : int
Number of top customers to return.
Returns
-------
list[tuple[str, float]]
A list of ``(customer_name, total)`` tuples, where *customer_name*
preserves the original spelling from its first **valid** row,
``total`` is rounded to two decimal places, and the list is sorted
by descending total then ascending name.
"""
# Mapping: normalized name -> (original spelling, accumulated total)
aggregates: Dict[str, Tuple[str, float]] = {}
for row in rows:
# Extract and clean customer name
raw_name = row.get("customer", "")
if not isinstance(raw_name, str):
continue
name = raw_name.strip()
if not name:
continue # skip empty names
# Parse amount
raw_amount = row.get("amount", "")
amount = _parse_amount(raw_amount)
if amount is None:
continue # skip rows with unparsable amount
norm = name.lower()
if norm in aggregates:
orig_name, current_total = aggregates[norm]
aggregates[norm] = (orig_name, current_total + amount)
else:
aggregates[norm] = (name, amount)
# Build list of (name, rounded_total)
result = [
(orig_name, round(total, 2))
for orig_name, total in (v for v in aggregates.values())
]
# Sort by total descending, then name ascending
result.sort(key=lambda item: (-item[1], item[0]))
# Return top n entries
return result[:n]
``` Coding: Slugify Worked
```python
import re
import unicodedata
def slugify(title: str, max_len: int = 50) -> str:
"""
Convert a title to a URL slug.
Steps:
1. Transliterate accented characters to ASCII, dropping any other non‑ASCII.
2. Lower‑case the result.
3. Replace runs of non‑alphanumeric characters with a single hyphen.
4. Trim leading/trailing hyphens.
5. Enforce ``max_len`` without breaking words (unless a single word exceeds the limit).
"""
# 1. Transliterate / drop non‑ASCII
normalized = unicodedata.normalize("NFKD", title)
ascii_str = normalized.encode("ascii", "ignore").decode("ascii")
# 2. Lowercase
lower = ascii_str.lower()
# 3. Replace runs of non‑alphanumerics with hyphen, then strip hyphens
slug = re.sub(r"[^a-z0-9]+", "-", lower).strip("-")
# 4. Enforce max_len without cutting words
if len(slug) > max_len:
words = slug.split("-")
kept_words = []
cur_len = 0
for w in words:
# length if we add this word (including a hyphen if not the first word)
add_len = len(w) + (1 if kept_words else 0)
if cur_len + add_len <= max_len:
kept_words.append(w)
cur_len += add_len
else:
break
if not kept_words:
# First word itself is longer than max_len → hard cut
slug = slug[:max_len]
else:
slug = "-".join(kept_words)
# 5. Ensure no trailing hyphen (should already be the case)
return slug.rstrip("-")
``` Coding: Token bucket Worked
```python
from __future__ import annotations
from typing import Callable
class TokenBucket:
"""
A simple token bucket rate limiter.
Parameters
----------
capacity : float
Maximum number of tokens the bucket can hold.
refill_per_sec : float
Rate at which tokens are added to the bucket (tokens per second).
clock : Callable[[], float]
Zero‑argument function returning the current time in seconds.
The implementation must never call ``time.time()`` directly.
"""
def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]) -> None:
if capacity <= 0:
raise ValueError("capacity must be positive")
if refill_per_sec < 0:
raise ValueError("refill_per_sec cannot be negative")
self._capacity = float(capacity)
self._refill_per_sec = float(refill_per_sec)
self._clock = clock
# The bucket starts full.
self._tokens: float = self._capacity
self._last_timestamp: float = self._clock()
def _refill(self) -> None:
"""Refill the bucket based on elapsed time."""
now = self._clock()
elapsed = now - self._last_timestamp
if elapsed > 0:
added = elapsed * self._refill_per_sec
self._tokens = min(self._capacity, self._tokens + added)
self._last_timestamp = now
else:
# No time passed; keep the timestamp as‑is.
# This also covers the case where the clock may go backwards.
self._last_timestamp = now
def allow(self, cost: float = 1) -> bool:
"""
Attempt to consume ``cost`` tokens.
Returns
-------
bool
``True`` if the bucket had enough tokens (they are deducted),
``False`` otherwise (no tokens are deducted).
"""
if cost < 0:
raise ValueError("cost cannot be negative")
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
return False
@property
def tokens(self) -> float:
"""Current token count after refilling (read‑only)."""
self._refill()
return self._tokens
``` Decisions: Refund window (hard) 100% right
{"outcome":{"full_refund":0.0,"partial_refund":0.0,"store_credit":0.99,"nothing":0.01},"defective":{"true":0.0,"false":1.0}} Decisions: Incident matrix (hard) 100% right
{
"severity": {
"0": 0.01,
"1": 0.98,
"2": 0.01,
"3": 0.0
},
"page": {
"true": 0.02,
"false": 0.98
}
} Decisions: Tool followup (hard) 100% right
{"tool":{"web_search":0.01,"calculator":0.01,"calendar":0.96,"email":0.01,"none":0.01},"confirm":{"true":0.95,"false":0.05}} Decisions: Legit security alert (hard) 100% right
{
"phishing": {"true": 0.02, "false": 0.98},
"action_needed": {"true": 0.03, "false": 0.97}
} Decisions: Meeting slot (hard) 100% right
{"slot":{"A":0,"B":1,"C":0,"D":0},"raj_last":{"true":1,"false":0}} Decisions: Review mixed (hard) 100% right
{"hardware":{"true":0.99,"false":0.01},"support":{"true":0.99,"false":0.01}} Decisions: Support checkout down 100% right
{"department":{"billing":0.1,"technical":0.9,"account":0.0,"shipping":0.0,"sales":0.0},"urgency":{"0":0.0,"1":0.0,"2":0.05,"3":0.95},"outage":{"true":0.98,"false":0.02}} Decisions: Refund wrong plan 100% right
{
"department": {
"billing": 0.94,
"technical": 0.02,
"account": 0.02,
"shipping": 0.01,
"sales": 0.01
},
"refund": {
"true": 0.99,
"false": 0.01
},
"tone": {
"calm": 0.98,
"frustrated": 0.02
}
} Decisions: Moderation doxxing 100% right
{"policy":{"none":0.0,"harassment":1.0,"hate":0.0,"spam":0.0,"self_harm":0.0},"personal_info":{"true":1.0,"false":0.0}} Decisions: Route calendar 100% right
{"tool":{"web_search":0.01,"calculator":0.01,"calendar":0.96,"email":0.01,"none":0.01},"confirm":{"true":0.98,"false":0.02}} Decisions: Doc invoice missing due 100% right
{
"doc_type": {
"invoice": 0.98,
"resume": 0.005,
"contract": 0.005,
"bank_statement": 0.005,
"other": 0.005
},
"missing_due_date": {
"true": 0.96,
"false": 0.04
}
} Decisions: Phishing paypal 100% right
{"phishing":{"true":0.99,"false":0.01},"risk":{"0":0.01,"1":0.01,"2":0.01,"3":0.97}} Decisions: Pii ssn email 100% right
{"data_kind":{"none":0.01,"contact":0.04,"financial":0.05,"government_id":0.88,"health":0.02},"sensitive":{"true":0.97,"false":0.03}} Decisions: Review mixed 100% right
{"sentiment":{"positive":0.02,"neutral":0.08,"negative":0.90},"defect":{"true":0.98,"false":0.02},"recommend":{"true":0.03,"false":0.97}} Documents: Saas escalator (hard) 100% right
{
"year2_price_per_seat_month": 47.25,
"year3_price_per_seat_month": 47.25,
"year1_invoice": 58320.00,
"year2_invoice": 61236.00,
"addon_months_billed": 6,
"addon_invoice": 38556.00,
"year3_invoice": 134946.00,
"year3_discount_percent": 15,
"total_contract_value": 293058.00,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 100% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{
"date": "2025-02-24",
"category": "airfare",
"amount_usd": 1184.60
},
{
"date": "2025-02-24",
"category": "ground_transport",
"amount_usd": 38.88
},
{
"date": "2025-02-25",
"category": "meals",
"amount_usd": 229.39
},
{
"date": "2025-02-26",
"category": "lodging",
"amount_usd": 466.56
},
{
"date": "2025-02-27",
"category": "ground_transport",
"amount_usd": 44.82
}
],
"rejected_item_count": 1,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 2159.25,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 100% right
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2236.00,
"late_fee_from_2025_06_01": 111.80,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2306.00,
"move_in_payment": 4700.00
} Documents: Ticket SLA 92% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address", "invoice_pdf"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-15T15:30",
"sla_due_utc": "2025-09-15T20:30:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 89% right
{
"q3_total_usd": 15346000,
"q2_total_usd": 14464000,
"q2_central_originally_reported_usd": 3047000,
"q2_to_q3_change_pct": 6.1,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "West",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731000,
"west_excluding_mountain_q3_usd": 4201000
} Size: 117B parameters. First tested OCT 11.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.