- Gemma 3 4B is a free model from Google that you can run on your own computer. In our tests it's not one we'd recommend right now: 14 out of 100, #52 of 56.
- It solved 1 of 30 coding jobs and scored 25 on reading documents. On our hardest tasks it scored 6.
- Runs on an 8 GB graphics card or a Mac with 16 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 3 4B got 1 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 3 4B scored 25; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 3 | 0 | 4 |
| Reading documents | 25 | 40 | 21 |
| Decisions | 68 | 64 | 69 |
On the 18 hardest tasks (included in the scores above) it scored 6. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
The Q4 download
Most people don't run the full-size model; they download a smaller Q4 version. We ran that version with Ollama on our Mac mini (M4, 16 GB) and gave it the same questions. It scored 13, about the same as the full-size model (14). You lose almost nothing by downloading it.
| Test | Full size (online) | Q4 (our Mac) |
|---|---|---|
| Coding | 3 | 0 |
| Reading documents | 25 | 27 |
| Hardest tasks | 6 | 8 |
| Overall | 14 | 13 |
Can your computer run it?
You can download Gemma 3 4B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | Yes | Q8 |
| Graphics card, 12 GB | Yes | Q8 |
| Graphics card, 16 GB | Yes | Q8 |
| Graphics card, 24 GB | Yes | Q8 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | Yes | Q8 |
| Mac, 24 GB | Yes | Q8 |
| Mac, 32 GB | Yes | Q8 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run gemma3:4b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Gemma 3 4B wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, Any
def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
"""Parses an INI file into a dictionary."""
result: Dict[str, Dict[str, str]] = {}
current_section: str = None
current_key: str = None
current_value: str = None
state = "scan" # "scan", "parse", "continue"
def _interpolate(value: str, section: str) -> str:
"""Recursively interpolates keys."""
value = value.replace('$', f'${section}:')
def _interpolate_recursive(val: str) -> str:
match = re.match(r'([^\s;#"]+|""[^"]*)"', val)
if match:
unquoted = match.group(1)
if unquoted.startswith('"') and unquoted.endswith('"'):
return unquoted
else:
return unquoted
else:
return val
return _interpolate_recursive(value)
for line in text.splitlines():
line = line.strip()
if not line:
continue
if line.startswith(";") or line.startswith("#"):
continue
if line.startswith("[") and line.endswith("]"):
section = line[1:-1]
if section == "DEFAULT":
continue
result[section] = {}
current_section = section
state = "parse"
continue
if current_section is None:
raise ValueError("Key line before any section header")
match = re.match(r"^\s*([a-zA-Z0-9_]+)\s*[=:]\s*(.*)", line)
if match:
if current_key is not None:
raise ValueError("Malformed line: multiple key lines")
key = match.group(1).lower()
value = match.group(2).strip()
if value.startswith('"') and value.endswith('"'):
value = value[1:-1]
result[current_section][key] = value
current_key = None
state = "parse"
continue
if state == "parse":
if current_key is None:
raise ValueError("Malformed line: value before key")
if line.startswith(" "):
value = line.strip()
if value.startswith('"') and value.endswith('"'):
value = value[1:-1]
if current_value is not None:
current_value += "\n" + value
else:
current_value = value
state = "continue"
continue
if current_value is not None:
result[current_section][current_key] = _interpolate(current_value, current_section)
current_value = None
current_key = None
state = "scan"
continue
else:
raise ValueError("Malformed line: key without value")
if current_key is not None:
raise ValueError("Malformed line: key without value")
return result
``` Coding: TTL LRU (hard) Didn't work
```python
import time
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = {}
self.lru_list = []
self.hits = 0
self.misses = 0
self.evictions = 0
self.expirations = 0
def put(self, key, value):
if key in self.cache:
self.cache[key] = value
self.lru_list.remove(key)
self.lru_list.append(key)
else:
if len(self.cache) >= self.capacity:
self.evict_lru()
self.cache[key] = value
self.lru_list.append(key)
now = self.clock()
self.cache[key]['expiry'] = now + self.ttl
def get(self, key, default=None):
if key in self.cache:
now = self.clock()
if now >= self.cache[key]['expiry']:
self.expirations += 1
self.lru_list.remove(key)
del self.cache[key]
return default
else:
self.hits += 1
self.lru_list.remove(key)
self.lru_list.append(key)
return self.cache[key]['value']
else:
self.misses += 1
return default
def __len__(self):
count = 0
for key in self.cache:
now = self.clock()
if now < self.cache[key]['expiry']:
count += 1
return count
def stats(self):
return {
"hits": self.hits,
"misses": self.misses,
"evictions": self.evictions,
"expirations": self.expirations,
}
def evict_lru(self):
if self.lru_list:
lru_key = self.lru_list.pop(0)
del self.cache[lru_key]
self.evictions += 1
``` Coding: Parse duration Didn't work
```python
import re
def parse_duration(s: str) -> int:
"""
Converts a duration string into a total number of seconds.
- Supported units: `h` (hours), `m` (minutes), `s` (seconds). Units are case-insensitive.
- Each part is a non-negative integer followed by a unit, e.g. "1h30m", "45s", "2H", "1h 5m 10s", "90m".
- Whitespace is allowed before, after and between parts.
- Units must appear in the order h, m, s and each at most once.
- Raise `ValueError` for empty/blank strings, numbers without a unit, unknown units, decimals, repeated units, or units out of order.
"""
if not s or s.strip() == "":
raise ValueError("Empty or blank string provided.")
parts = re.split(r"(\d+)(\b[hm]s\b)", s)
if len(parts) % 2 != 0:
raise ValueError("Invalid duration string format.")
hours = 0
minutes = 0
seconds = 0
for i in range(0, len(parts), 2):
try:
num = int(parts[i])
except ValueError:
raise ValueError("Invalid number format.")
unit = parts[i+1].lower()
if unit not in ['h', 'm', 's']:
raise ValueError(f"Unknown unit: {unit}")
if unit == 'h':
hours = num
elif unit == 'm':
minutes = num
else:
seconds = num
if hours > 24:
raise ValueError("Hours cannot exceed 24.")
total_seconds = (hours * 3600) + (minutes * 60) + seconds
return total_seconds
``` Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
"""Sort the numbers and remove duplicates. Then collapse runs of 3 or more consecutive integers into "a..b". Numbers in shorter runs (1 or 2 numbers) are listed individually. Join everything with "," (no spaces). An empty list returns "".
Examples: [1,2,3,5,7,8] -> "1..3,5,7,8"; [-3,-2,-1,1] -> "-3..-1,1".
"""
if not nums:
return ""
nums = sorted(list(set(nums)))
result = []
i = 0
while i < len(nums):
if i + 1 < len(nums) and nums[i] + 1 == nums[i + 1]:
j = i + 1
while j + 1 < len(nums) and nums[j] + 1 == nums[j + 1]:
j += 1
result.append(f"{nums[i]}..{nums[j]}")
i = j + 1
else:
result.append(str(nums[i]))
i += 1
return ",".join(result)
``` Coding: Top customers Didn't work
```python
from typing import List, Tuple
import re
def top_customers(rows: List[dict], n: int) -> List[Tuple[str, float]]:
"""
Calculates the total amount spent by each customer and returns the top n customers.
Args:
rows: A list of dictionaries, where each dictionary represents a row with a "customer" name and an "amount" string.
n: The number of top customers to return.
Returns:
A list of tuples, where each tuple contains the customer name and total amount spent, sorted by total descending,
then name ascending.
"""
customer_totals = {}
for row in rows:
customer_name = row.get("customer", "").strip()
amount_str = row.get("amount", "")
if not customer_name or not amount_str:
continue
try:
amount = float(amount_str.replace("$", "").replace(",", ""))
if amount < 0:
amount = -amount # Handle refunds
if customer_name not in customer_totals:
customer_totals[customer_name] = 0.0
customer_totals[customer_name] += amount
except ValueError:
continue
sorted_customers = sorted(customer_totals.items(), key=lambda item: (-item[1], item[0]))
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import re
def slugify(title: str, max_len: int = 50) -> str:
"""
Builds a URL slug.
1. Transliterate accented characters to ASCII (é -> e, ü -> u) and drop any other non-ASCII characters.
2. Lowercase.
3. Replace every run of characters other than a-z and 0-9 with a single "-", and strip leading/trailing "-".
4. If the slug is longer than `max_len`, shorten it without cutting a word: keep the longest prefix of whole words (words are separated by "-") whose length is <= max_len. If even the first word is longer than max_len, hard-cut it to max_len characters.
5. The result never ends with "-".
"""
title = title.lower()
title = re.sub(r'[^\w\s-]', '', title)
title = re.sub(r'[-\s]+', '-', title)
title = title.strip('-')
if len(title) > max_len:
parts = title.split('-')
parts = [part for part in parts if len(part) <= max_len]
title = '-'.join(parts)
if title and title[-1] == '-':
title = title[:-1]
return title
``` Coding: Token bucket Didn't work
```python
import time
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.clock = clock
self.tokens = capacity
self.last_refill = self.clock()
def refill(self):
now = self.clock()
elapsed = now - self.last_refill
if elapsed > 0:
refill_amount = elapsed * self.refill_per_sec
self.tokens = min(self.capacity, self.tokens + refill_amount)
self.last_refill = now
def allow(self, cost: float = 1) -> bool:
self.refill()
if self.tokens >= cost:
self.tokens -= cost
return True
else:
return False
@property
def tokens(self):
self.refill()
return self.tokens
``` Decisions: Refund window (hard) 100% right
```json
{
"outcome": {
"full_refund": 0.1,
"partial_refund": 0.3,
"store_credit": 0.6,
"nothing": 0.0
},
"defective": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Incident matrix (hard) 0% right
```json
{
"severity": "3",
"page": "true"
}
``` Decisions: Tool followup (hard) 100% right
```json
{
"tool": {
"calendar": 0.95,
"none": 0.05
},
"confirm": {
"true": 0.8,
"false": 0.2
}
}
``` Decisions: Legit security alert (hard) 0% right
```json
{
"phishing": {
"true": 0.8,
"false": 0.2
},
"action_needed": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Meeting slot (hard) 50% right
```json
{
"slot": {
"A": 0.1,
"B": 0.2,
"C": 0.6,
"D": 0.1
},
"raj_last": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Review mixed (hard) 100% right
```json
{
"hardware": {
"true": 0.95,
"false": 0.05
},
"support": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Support checkout down 67% right
```json
{
"department": {
"billing": 0.2,
"technical": 0.6,
"account": 0.1,
"shipping": 0.05,
"sales": 0.05
},
"urgency": {
"0": 0.05,
"1": 0.2,
"2": 0.5,
"3": 0.25
},
"outage": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Refund wrong plan 33% right
```json
{
"department": {
"billing": 0.85,
"technical": 0.05,
"account": 0.05,
"shipping": 0.0,
"sales": 0.0
},
"refund": "true",
"tone": "calm"
}
``` Decisions: Moderation doxxing 0% right
```json
{
"policy": {
"none": 0.7,
"harassment": 0.3
},
"personal_info": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Route calendar 100% right
```json
{
"tool": {
"calendar": 0.95,
"none": 0.05
},
"confirm": {
"true": 0.98,
"false": 0.02
}
}
``` Decisions: Doc invoice missing due 100% right
```json
{
"doc_type": {
"invoice": 0.99,
"resume": 0.01,
"contract": 0.01,
"bank_statement": 0.01,
"other": 0.0
},
"missing_due_date": {
"true": 0.8,
"false": 0.2
}
}
``` Decisions: Phishing paypal 50% right
```json
{
"phishing": {
"true": 0.95,
"false": 0.05
},
"risk": {
"0": 0.05,
"1": 0.15,
"2": 0.60,
"3": 0.20
}
}
``` Decisions: Pii ssn email 100% right
```json
{
"data_kind": {
"government_id": 0.95,
"contact": 0.05,
"financial": 0.0,
"health": 0.0
},
"sensitive": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Review mixed 100% right
```json
{
"sentiment": {
"positive": 0.1,
"neutral": 0.2,
"negative": 0.7
},
"defect": {
"true": 0.95,
"false": 0.05
},
"recommend": {
"true": 0.05,
"false": 0.95
}
}
``` Documents: Saas escalator (hard) 10% right
```json
{
"year2_price_per_seat_month": 45.00,
"year3_price_per_seat_month": 45.00,
"year1_invoice": 5400.00,
"year2_invoice": 5400.00,
"addon_months_billed": 160,
"addon_invoice": 7200.00,
"year3_invoice": 5400.00,
"year3_discount_percent": 15.00,
"total_contract_value": 21600.00,
"contract_end_date": "2026-02-28"
}
``` Documents: Expense thread 71% right
```json
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{
"date": "2025-02-24",
"category": "airfare",
"amount_usd": 1184.60
},
{
"date": "2025-02-24",
"category": "ground_transport",
"amount_usd": 36.00
},
{
"date": "2025-02-25",
"category": "meals",
"amount_usd": 212.40
},
{
"date": "2025-02-26",
"category": "lodging",
"amount_usd": 432.00
},
{
"date": "2025-02-27",
"category": "ground_transport",
"amount_usd": 41.50
}
],
"rejected_item_count": 2,
"per_diem_days": 3,
"per_diem_usd": 205.00,
"total_reimbursable_usd": 2002.50,
"approver_email": "priya.raman@corvane.com"
}
``` Documents: Lease amendment 67% right
```json
{
"tenants": [
"Marcus Lin",
"Sofia Lin"
],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2295.00,
"late_fee_from_2025_06_01": 111.75,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2406.75,
"move_in_payment": 6400.00
}
``` Documents: Ticket SLA 54% right
```json
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": [
"billing_address",
"invoice_pdf"
],
"affected_orders": [
"SO-99812",
"SO-99815",
"SO-99820",
"SO-99827"
],
"priority": "P2",
"sla_due_local": "2025-09-13T09:00",
"sla_due_utc": "2025-09-13T09:00:00Z",
"reissued_invoice": "INV-2025-0812"
}
``` Documents: Sales footnotes 0% right
```json
{
"q3_total_usd": 14,763,
"q2_total_usd": 11,962,
"q2_central_originally_reported_usd": 3047,
"q2_to_q3_change_pct": 23.7,
"top_region_q3": "West",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["Central"],
"international_q3_organic_usd": 1951,
"west_excluding_mountain_q3_usd": 4201
}
``` Size: 4.3B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.