- Gemma 2 27B is a free model from Google that you can run on your own computer. In our tests it's not one we'd recommend right now: 23 out of 100, #46 of 56.
- It solved 4 of 30 coding jobs and scored 34 on reading documents. On our hardest tasks it scored 12.
- Runs on a 24 GB graphics card or a Mac with 32 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 2 27B got 4 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 2 27B scored 34; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 13 | 14 | 13 |
| Reading documents | 34 | 38 | 32 |
| Decisions | 80 | 86 | 77 |
On the 18 hardest tasks (included in the scores above) it scored 12. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Gemma 2 27B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | Yes | Q5 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | Yes | Q4 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run gemma2:27b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Gemma 2 27B wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
def parse_ini(text: str) -> dict[str, dict[str, str]]:
sections = {}
current_section = None
default_section = {}
def add_key(key, value):
if current_section is not None:
if key in sections.get(current_section, {}):
sections[current_section][key] = value
else:
sections[current_section][key] = value
else:
default_section[key] = value
for line in text.splitlines():
line = line.strip()
if not line:
continue
if line.startswith('[') and line.endswith(']'):
current_section = line[1:-1].strip()
sections[current_section] = {}
elif line.startswith('#') or line.startswith(';'):
continue
else:
parts = line.split('=', 1)
if len(parts) == 1:
parts = parts[0].split(':', 1)
key = parts[0].strip().lower()
value =''.join(parts[1:]).strip()
if value.startswith(('"', "'") and value.endswith(('"', "'")):
value = value[1:-1]
if value:
add_key(key, value.strip())
for section, data in sections.items():
for key, value in data.items():
if '${' in value:
def resolve_key(key):
key = key.strip()
if key.startswith("${") and key.endswith("}"):
key = key[2:-1]
if key.count(':') == 1:
section_name, key = key.split(':')
try:
return sections.get(section_name, {})[key]
except KeyError:
raise KeyError(f"{{ {key} }}")
else:
if key in sections.get(section, {}):
return sections.get(section, {})[key]
elif key in default_section:
return default_section[key]
raise KeyError(f"{{ {key} }}")
while '${' in value:
value = value.replace(key, resolve_key(key))
add_key(key, value)
return sections
``` Coding: TTL LRU (hard) Didn't work
```python
import heapq
from collections import defaultdict
class TTLCache:
def __init__(self, capacity, ttl, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = {}
self.expiry_heap = []
self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
def put(self, key, value):
now = self.clock()
self._remove_expired()
if key in self.cache:
self._remove_entry(key)
if len(self.cache) == self.capacity:
_, evicted_key = heapq.heappop(self.expiry_heap)
del self.cache[evicted_key]
self.stats["evictions"] += 1
expiry = now + self.ttl
self.cache[key] = (value, expiry)
heapq.heappush(self.expiry_heap, (expiry, key))
def get(self, key, default=None):
now = self.clock()
if key in self.cache:
value, expiry = self.cache[key]
if now >= expiry:
self._remove_entry(key)
self.stats["expirations"] += 1
return default
self._remove_entry(key)
heapq.heappush(self.expiry_heap, (expiry, key))
self.stats["hits"] += 1
return value
self.stats["misses"] += 1
return default
def __len__(self):
self._remove_expired()
return len(self.cache)
def _remove_expired(self):
now = self.clock()
while self.expiry_heap and self.expiry_heap[0][0] <= now:
_, key = heapq.heappop(self.expiry_heap)
del self.cache[key]
def _remove_entry(self, key):
_, key = heapq.heappop(self.expiry_heap)
del self.cache[key]
``` Coding: Parse duration Didn't work
```python
def parse_duration(s: str) =
s = s.strip().lower()
if not s:
raise ValueError("Empty/blank string not allowed.")
parts = s.split()
if len(parts) == 0 or len(parts) > 3:
raise ValueError("Invalid duration format.")
total_seconds = 0
seen_units = set()
for part in parts:
if len(part) < 2:
raise ValueError("Invalid duration format (number without unit).")
number = part[:-1]
unit = part[-1]
if not number.isdigit():
raise ValueError("Invalid duration format (non-digit number).")
if unit not in ('h', 'm', 's'):
raise ValueError("Unknown unit")
if unit in seen_units:
raise ValueError("Repeated unit.")
seen_units.add(unit)
if unit == 'h':
total_seconds += int(number) * 3600
elif unit == 'm':
total_seconds += int(number) * 60
else:
total_seconds += int(number)
return total_seconds
``` Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
nums = sorted(set(nums))
if not nums:
return ""
ranges = []
start = nums[0]
for i in range(1, len(nums)):
if nums[i] != nums[i - 1] + 1:
if i - 1 >= 2 and nums[i - 1] - start >= 2:
ranges.append(f"{start}..{nums[i - 1]}")
else:
ranges.extend([str(x) for x in range(start, i)])
start = nums[i]
if i == len(nums) - 1:
if i - 1 >= 2 and nums[i] - start >= 2:
ranges.append(f"{start}..{nums[i]}")
else:
ranges.extend([str(x) for x in range(start, i + 1)])
return ",".join(ranges)
``` Coding: Top customers Didn't work
```python
import re
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
customer_sums = {}
for row in rows:
customer = row.get("customer", "").strip().lower()
amount_str = row.get("amount")
if customer and re.match(r"^\$\d+(\.\d+)?$", amount_str):
amount = float(re.sub(r"\$", "", amount_str))
if amount < 0:
amount *= -1
if customer not in customer_sums:
customer_sums[customer] = 0
customer_sums[customer] += amount
top_customers = sorted(customer_sums.items(), key=lambda x: (-x[1], x[0]))
return [(name, round(total, 2)) for name, total in top_customers[:n]]
```
Coding: Slugify Didn't work
```python
import unicodedata
def slugify(title: str, max_len: int = 50) -> str:
title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
title = title.lower()
title = ''.join('-'.join(title.split('-')).split() )
words = title.split('-')
for i in range(len(words)):
if len(words[i]) > max_len:
if i == 0:
return words[i][:max_len]
else:
return '-'.join(words[:i])
return '-'.join(words)
``` Coding: Token bucket Worked
```python
from typing import Callable, Optional
import time
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
self._capacity = capacity
self._refill_per_sec = refill_per_sec
self._clock = clock
self._tokens = self._capacity
self._last_refill = self._clock()
def allow(self, cost: float = 1) -> bool:
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
return False
@property
def tokens(self) -> float:
self._refill()
return self._tokens
def _refill(self):
now = self._clock()
elapsed = now - self._last_refill
self._tokens = min(self._capacity, self._tokens + self._refill_per_sec * elapsed)
self._last_refill = now
``` Decisions: Refund window (hard) 50% right
```json
{
"outcome": {
"full_refund": 0,
"partial_refund": 0.8,
"store_credit": 0.2,
"nothing": 0
},
"defective": {
"true": 0,
"false": 1
}
}
``` Decisions: Incident matrix (hard) 50% right
```json
{
"severity": {
"0": 0.1,
"1": 0.7,
"2": 0.2,
"3": 0
},
"page": {
"true": 0.7,
"false": 0.3
}
}```
Decisions: Tool followup (hard) 100% right
:```json
{
"tool": {
"calendar": 1.0
},
"confirm": {
"true": 1.0
}
}
``` Decisions: Legit security alert (hard) 100% right
```json
{
"phishing": {
"true": 0.02,
"false": 0.98
},
"action_needed": {
"true": 0.2,
"false": 0.8
}
}```
Decisions: Meeting slot (hard) 50% right
```json
{
"slot": {
"A": 0.0,
"B": 1.0,
"C": 0.0,
"D": 0.0
},
"raj_last": {
"true": 0.0,
"false": 1.0
}
}
```
Decisions: Review mixed (hard) 100% right
```json
{
"hardware": {"true": 1.0, "false": 0.0},
"support": {"true": 1.0, "false": 0.0}
}
``` Decisions: Support checkout down 100% right
```json
{
"department": {
"billing": 0.01,
"technical": 0.99
},
"urgency": {
"0": 0.0,
"1": 0.0,
"2": 0.1,
"3": 0.9
},
"outage": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Refund wrong plan 100% right
```json
{
"department": {
"billing": 1.0
},
"refund": {
"true": 1.0
},
"tone": {
"calm": 1.0
}
}
``` Decisions: Moderation doxxing 100% right
```json
{
"policy": {
"none": 0.0,
"harassment": 1.0,
"hate": 0.0,
"spam": 0.0,
"self_harm": 0.0
},
"personal_info": {
"true": 1.0,
"false": 0.0
}
}```
Decisions: Route calendar 100% right
```json
{
"tool": {
"web_search": 0.0,
"calculator": 0.0,
"calendar": 1.0,
"email": 1.0,
"none": 0.0
},
"confirm": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Doc invoice missing due 100% right
```json
{
"doc_type": {
"invoice": 0.99,
"resume": 0.001,
"contract": 0.005,
"bank_statement": 0.004,
"other": 0.00
},
"missing_due_date": {
"true": 0.95,
"false": 0.05
}
}``` Decisions: Phishing paypal 100% right
```json
{
"phishing": {
"true": 0.95,
"false": 0.05
},
"risk": {
"0": 0.0,
"1": 0.0,
"2": 0.4,
"3": 0.6
}
}
``` Decisions: Pii ssn email 50% right
```json
{
"data_kind": {
"none": 0.0,
"contact": 0.5,
"financial": 0.2,
"government_id": 0.3,
"health": 0.0
},
"sensitive": {
"true": 0.5,
"false": 0.5
}
}
```
Decisions: Review mixed 100% right
```json
{
"sentiment": {
"positive": 0.2,
"neutral": 0.1,
"negative": 0.7
},
"defect": {
"true": 0.8,
"false": 0.2
},
"recommend": {
"true": 0.2,
"false": 0.8
}
}
```
Documents: Saas escalator (hard) 20% right
```json
{
"year2_price_per_seat_month": 47.00,
"year3_price_per_seat_month": 50.00,
"year1_invoice": 6480.00,
"year2_invoice": 5616.00,
"addon_months_billed": 4,
"addon_invoice": 3200.00,
"year3_invoice": null,
"year3_discount_percent": 15.00,
"total_contract_value": null,
"contract_end_date": "2027-02-28"
}
``` Documents: Expense thread 0% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{
"date": "2025-02-24",
"category": "airfare",
"amount_usd": 1184.60
},
{
"date": "2025-02-24",
"category": "ground_transport",
"amount_usd": 36.00
},
{
"date": "2025-02-25",
"category": "meals",
"amount_usd": 212.40 * 1.08
},
{
"date": "2025-02-26",
"category": "lodging",
"amount_usd": 432.00 * 1.08
}
],
"rejected_item_count": 2,
"per_diem_days": 2,
"per_diem_usd": 65.00,
"total_reimbursable_usd": null,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 83% right
```json
{
"tenants": [
"Marcus Lin",
"Sofia Lin"
],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150,
"monthly_rent_from_2025_06_01": 2236,
"late_fee_from_2025_06_01": 111.8,
"security_deposit": 2150,
"total_pet_deposits": 800,
"total_monthly_payment_july_2025": 2271,
"move_in_payment": 4300
}
```
Documents: Ticket SLA 75% right
```json
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": [
"billing_address",
"invoice_pdf"
],
"affected_orders": [
"SO-99812",
"SO-99820",
"SO-99827"
],
"priority": "P2",
"sla_due_local": "2025-09-13T17:00",
"sla_due_utc": "2025-09-14T00:00:00Z",
"reissued_invoice": "INV-2025-0812"
}
```
Documents: Sales footnotes 11% right
```json
{
"q3_total_usd": 17041,
"q2_total_usd": 15464,
"q2_central_originally_reported_usd": 3047,
"q2_to_q3_change_pct": 9.7,
"top_region_q3": "West",
"fastest_growing_region_q1_to_q3": "West",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731,
"west_excluding_mountain_q3_usd": 4291
}``` Size: 27B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.