- Gemma 3 27B is a free model from Google that you can run on your own computer. In our tests it's not one we'd recommend right now: 40 out of 100, #38 of 56.
- It solved 10 of 30 coding jobs and scored 46 on reading documents. On our hardest tasks it scored 29.
- Runs on a 24 GB graphics card or a Mac with 32 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 3 27B got 10 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 3 27B scored 46; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 33 | 29 | 35 |
| Reading documents | 46 | 65 | 41 |
| Decisions | 88 | 93 | 87 |
On the 18 hardest tasks (included in the scores above) it scored 29. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Gemma 3 27B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | Yes | Q5 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | Yes | Q4 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run gemma3:27b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Gemma 3 27B wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
.
```python
import re
from typing import Dict
def parse_ini(text: str) -> dict[str, dict[str, str]]:
"""Parses an INI string according to the specified rules."""
result: Dict[str, Dict[str, str]] = {}
default_section: Dict[str, str] = {}
current_section: str = None
current_key: str = None
current_value: str = None
in_string: bool = False
interpolation_stack: list[tuple[str, str]] = []
for line in text.splitlines():
line = line.strip()
if not line or line.startswith((';', '#')):
continue
if line.startswith('[') and line.endswith(']'):
section_name = line[1:-1].strip()
if current_section is not None:
if current_key:
if current_value is None:
current_value = ''
result[current_section][current_key] = current_value
current_section = section_name
result[current_section] = {}
current_key = None
current_value = None
in_string = False
interpolation_stack = []
else:
if current_section is None:
raise ValueError("Key line before section header")
match = re.match(r'(\w+)\s*[:=]\s*(.+)', line)
if not match:
raise ValueError("Malformed line")
key, value = match.groups()
key = key.lower().strip()
if current_key is None:
current_key = key
in_string = False
# Extract value, handling quotes and comments
value_start = 0
while value_start < len(value):
if value[value_start] == '"':
value_start += 1
value_end = value.find('"', value_start)
if value_end == -1:
current_value = value[value_start-1:]
break
current_value = value[value_start : value_end]
value_start = value_end + 1
else:
comment_start = value.find(';', value_start)
hash_comment_start = value.find('#', value_start)
if comment_start != -1 and (hash_comment_start == -1 or comment_start < hash_comment_start):
current_value = value[value_start:comment_start].strip()
break
elif hash_comment_start != -1:
current_value = value[value_start:hash_comment_start].strip()
break
else:
current_value = value[value_start:].strip()
break
else:
# Continuation line
current_value += "\n" + value
if current_section is not None and current_key is not None:
if current_value is None:
current_value = ''
result[current_section][current_key] = current_value
#Apply DEFAULT section
for key, value in default_section.items():
if key not in result.get('DEFAULT', {}):
result['DEFAULT'] = result.get('DEFAULT',{})
result['DEFAULT'][key] = value
final_result: Dict[str, Dict[str, str]] = {}
for section_name, section_data in result.items():
if section_name != 'DEFAULT':
#Inherit from default
inherited_data = default_section.copy()
for key, value in section_data.items():
inherited_data[key] = value
final_result[section_name] = inherited_data
else:
final_result[section_name] = section_data
for section_name, section_data in final_result.items():
for key, value in section_data.items():
final_result[section_name][key] = interpolate(value, final_result)
return final_result
def interpolate(value: str, data: Dict[str, Dict[str, str]]) -> str:
"""Interpolates values in a string using the provided data."""
def replace_match(match):
full_match = match.group(0)
if full_match.startswith("${"):
ref = full_match[2:-1]
if ":" in ref:
section, key = ref.split(":", 1)
if section in data and key in data[section]:
return data[section][key]
else:
raise KeyError(f"Section or key not found: {section}:{key}")
else:
if "DEFAULT" in data and ref in data["DEFAULT"]:
return data["DEFAULT"][ref]
else:
raise KeyError(f"Key not found: {ref}")
else:
return match.group(0)
pattern = r"\${[^}]+}"
while True:
original_value = value
value = re.sub(pattern, replace_match, value)
if value == original_value:
break
return value
``` Coding: TTL LRU (hard) Worked
```python
from collections import OrderedDict
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.hits = 0
self.misses = 0
self.evictions = 0
self.expirations = 0
def put(self, key, value):
if self.capacity == 0:
return
now = self.clock()
expired = []
# Remove expired entries
for k, (v, expiry) in self.cache.items():
if now >= expiry:
expired.append(k)
for k in expired:
del self.cache[k]
self.expirations += 1
if key in self.cache:
self.cache.move_to_end(key)
self.cache[key] = (value, now + self.ttl)
return
if len(self.cache) >= self.capacity:
lru_key, _ = self.cache.popitem(last=False)
self.evictions += 1
self.cache[key] = (value, now + self.ttl)
self.cache.move_to_end(key)
def get(self, key, default=None):
if self.capacity == 0:
self.misses += 1
return default
now = self.clock()
if key in self.cache:
value, expiry = self.cache[key]
if now >= expiry:
del self.cache[key]
self.expirations += 1
self.misses += 1
return default
else:
self.hits += 1
self.cache.move_to_end(key)
self.cache[key] = (value, expiry)
return value
else:
self.misses += 1
return default
def __len__(self):
now = self.clock()
count = 0
for _, (value, expiry) in self.cache.items():
if now < expiry:
count += 1
return count
def stats(self):
return {
"hits": self.hits,
"misses": self.misses,
"evictions": self.evictions,
"expirations": self.expirations,
}
``` Coding: Parse duration Didn't work
```python
import re
def parse_duration(s: str) -> int:
"""Converts a duration string into a total number of seconds.
Supported units: `h` (hours), `m` (minutes), `s` (seconds). Units are case-insensitive.
Each part is a non-negative integer followed by a unit, e.g. "1h30m", "45s", "2H", "1h 5m 10s", "90m".
Whitespace is allowed before, after and between parts.
Units must appear in the order h, m, s and each at most once.
Raise ValueError for empty/blank strings, numbers without a unit, unknown units, decimals, repeated units, or units out of order.
"""
s = s.strip()
if not s:
raise ValueError("Empty duration string")
parts = re.findall(r"(\d+)[hms]", s, re.IGNORECASE)
units = re.findall(r"[hms]", s, re.IGNORECASE)
if not parts:
raise ValueError("Invalid duration format")
if len(parts) != len(units):
raise ValueError("Invalid duration format: numbers without unit")
hours = 0
minutes = 0
seconds = 0
hour_found = False
minute_found = False
second_found = False
for i in range(len(parts)):
try:
value = int(parts[i])
except ValueError:
raise ValueError("Invalid duration format: non-integer value")
unit = units[i].lower()
if unit == 'h':
if hour_found:
raise ValueError("Repeated unit: hour")
hours = value
hour_found = True
elif unit == 'm':
if hour_found and not minute_found:
pass
elif not hour_found and minute_found:
raise ValueError("Units out of order")
if minute_found:
raise ValueError("Repeated unit: minute")
minutes = value
minute_found = True
elif unit == 's':
if hour_found and minute_found and not second_found:
pass
elif hour_found and not minute_found and second_found:
raise ValueError("Units out of order")
elif not hour_found and minute_found and second_found:
raise ValueError("Units out of order")
if second_found:
raise ValueError("Repeated unit: second")
seconds = value
second_found = True
else:
raise ValueError("Unknown unit")
total_seconds = hours * 3600 + minutes * 60 + seconds
return total_seconds
``` Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
nums = sorted(list(set(nums)))
result = []
start = nums[0]
end = nums[0]
for i in range(1, len(nums)):
if nums[i] == end + 1:
end = nums[i]
else:
if start == end:
result.append(str(start))
else:
result.append(f"{start}..{end}")
start = nums[i]
end = nums[i]
if start == end:
result.append(str(start))
else:
result.append(f"{start}..{end}")
return ",".join(result)
``` Coding: Top customers Didn't work
```python
import re
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
"""
Finds the top n customers based on their total spending, considering refunds.
Args:
rows: A list of dictionaries, each representing a transaction with "customer" and "amount" keys.
n: The number of top customers to return.
Returns:
A list of (name, total) tuples, sorted by total descending, then name ascending.
"""
customer_totals = {}
for row in rows:
customer = row.get("customer")
amount_str = row.get("amount")
if not customer or not amount_str:
continue
customer = customer.strip()
try:
amount = float(re.sub(r'[^\d\.-]', '', amount_str))
except ValueError:
continue
if customer not in customer_totals:
customer_totals[customer] = 0.0
customer_totals[customer] += amount
result = [(name, round(total, 2)) for name, total in customer_totals.items()]
result.sort(key=lambda x: (-x[1], x[0]))
return result[:n]
``` Coding: Slugify Didn't work
.
```python
import unicodedata
def slugify(title: str, max_len: int = 50) -> str:
"""Builds a URL slug."""
title = ''.join(c for c in unicodedata.normalize('NFKD', title)
if unicodedata.category(c) != 'Mn')
title = ''.join(c for c in title if 0 <= ord(c) <= 0x10ffff)
title = title.lower()
title = ''.join(c if 'a' <= c <= 'z' or '0' <= c <= '9' else '-' for c in title)
title = title.strip('-')
if len(title) > max_len:
words = title.split('-')
current_len = 0
result = []
for word in words:
if current_len + len(word) + 1 <= max_len:
result.append(word)
current_len += len(word) + 1
else:
if len(word) > max_len:
result.append(word[:max_len])
break
title = '-'.join(result).strip('-')
return title
``` Coding: Token bucket Worked
```python
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock):
"""
Initializes a TokenBucket rate limiter.
Args:
capacity: The maximum number of tokens the bucket can hold.
refill_per_sec: The rate at which tokens are added per second.
clock: A callable that returns the current time in seconds.
"""
self._capacity = capacity
self._refill_per_sec = refill_per_sec
self._clock = clock
self._tokens = capacity
self._last_refill_time = 0.0
def allow(self, cost: float = 1) -> bool:
"""
Refills the bucket and attempts to subtract tokens.
Args:
cost: The number of tokens to subtract.
Returns:
True if the tokens were subtracted successfully, False otherwise.
"""
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
else:
return False
@property
def tokens(self) -> float:
"""
Returns the current number of tokens in the bucket.
"""
self._refill()
return self._tokens
def _refill(self):
"""
Refills the bucket with tokens based on elapsed time.
"""
now = self._clock()
time_elapsed = now - self._last_refill_time
refill_amount = time_elapsed * self._refill_per_sec
self._tokens = min(self._capacity, self._tokens + refill_amount)
self._last_refill_time = now
``` Decisions: Refund window (hard) 50% right
```json
{
"outcome": {
"full_refund": 0.05,
"partial_refund": 0.9,
"store_credit": 0.03,
"nothing": 0.02
},
"defective": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Incident matrix (hard) 100% right
```json
{
"severity": {
"0": 0.05,
"1": 0.75,
"2": 0.15,
"3": 0.05
},
"page": {
"true": 0.3,
"false": 0.7
}
}
``` Decisions: Tool followup (hard) 100% right
```json
{
"tool": {
"calendar": 0.95,
"web_search": 0.01,
"calculator": 0.01,
"email": 0.02,
"none": 0.01
},
"confirm": {
"true": 0.85,
"false": 0.15
}
}
``` Decisions: Legit security alert (hard) 100% right
```json
{
"phishing": {
"true": 0.05,
"false": 0.95
},
"action_needed": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Meeting slot (hard) 50% right
```json
{
"slot": {
"A": 0.05,
"B": 0.9,
"C": 0.03,
"D": 0.02
},
"raj_last": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Review mixed (hard) 100% right
```json
{
"hardware": {
"true": 0.95,
"false": 0.05
},
"support": {
"true": 0.85,
"false": 0.15
}
}
``` Decisions: Support checkout down 100% right
```json
{
"department": {
"billing": 0.1,
"technical": 0.85,
"account": 0.02,
"shipping": 0.01,
"sales": 0.02
},
"urgency": {
"0": 0.01,
"1": 0.05,
"2": 0.2,
"3": 0.74
},
"outage": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Refund wrong plan 100% right
```json
{
"department": {
"billing": 0.95,
"technical": 0.01,
"account": 0.02,
"shipping": 0.01,
"sales": 0.01
},
"refund": {
"true": 0.98,
"false": 0.02
},
"tone": {
"frustrated": 0.2,
"calm": 0.8
}
}
``` Decisions: Moderation doxxing 100% right
```json
{
"policy": {
"none": 0.05,
"harassment": 0.9,
"hate": 0.01,
"spam": 0.01,
"self_harm": 0.03
},
"personal_info": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Route calendar 100% right
```json
{
"tool": {
"calendar": 0.95,
"email": 0.03,
"web_search": 0.01,
"calculator": 0.005,
"none": 0.005
},
"confirm": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Doc invoice missing due 100% right
```json
{
"doc_type": {
"invoice": 0.99,
"resume": 0.001,
"contract": 0.005,
"bank_statement": 0.002,
"other": 0.002
},
"missing_due_date": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Phishing paypal 100% right
```json
{
"phishing": {
"true": 0.95,
"false": 0.05
},
"risk": {
"0": 0.01,
"1": 0.04,
"2": 0.15,
"3": 0.80
}
}
``` Decisions: Pii ssn email 100% right
```json
{
"data_kind": {
"none": 0.05,
"contact": 0.2,
"financial": 0.05,
"government_id": 0.65,
"health": 0.05
},
"sensitive": {
"true": 0.9,
"false": 0.1
}
}
``` Decisions: Review mixed 100% right
```json
{
"sentiment": {
"positive": 0.05,
"neutral": 0.1,
"negative": 0.85
},
"defect": {
"true": 0.95,
"false": 0.05
},
"recommend": {
"true": 0.05,
"false": 0.95
}
}
``` Documents: Saas escalator (hard) 10% right
```json
{
"year2_price_per_seat_month": 47.84,
"year3_price_per_seat_month": 47.40,
"year1_invoice": 5400.00,
"year2_invoice": 5731.20,
"addon_months_billed": 27,
"addon_invoice": 7200.00,
"year3_invoice": 5668.80,
"year3_discount_percent": 0.15,
"total_contract_value": 24000.00,
"contract_end_date": "2027-02-28"
}
``` Documents: Expense thread 75% right
```json
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{ "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
{ "date": "2025-02-24", "category": "ground_transport", "amount_usd": 41.04 },
{ "date": "2025-02-25", "category": "meals", "amount_usd": 229.39 },
{ "date": "2025-02-26", "category": "lodging", "amount_usd": 468.24 }
],
"rejected_item_count": 1,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 2118.27,
"approver_email": "priya.raman@corvane.com"
}
``` Documents: Lease amendment 67% right
```json
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2238.00,
"late_fee_from_2025_06_01": 111.90,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2238.00,
"move_in_payment": 2550.00
}
``` Documents: Ticket SLA 82% right
```json
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-12T17:00",
"sla_due_utc": "2025-09-12T21:00:00Z",
"reissued_invoice": "INV-2025-0812"
}
``` Documents: Sales footnotes 89% right
```json
{
"q3_total_usd": 15346000,
"q2_total_usd": 14464000,
"q2_central_originally_reported_usd": 3047000,
"q2_to_q3_change_pct": 5.9,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731000,
"west_excluding_mountain_q3_usd": 4201000
}
``` Size: 27B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.