Review · updated OCT 11

Gemma 3 4B review: not one we'd recommend right now

It scored 14 out of 100, #52 of 56. It solved 1 of 30 coding jobs and scored 25 on reading documents. Runs on an 8 GB graphics card or a Mac with 16 GB.

The short version
  • Gemma 3 4B is a free model from Google that you can run on your own computer. In our tests it's not one we'd recommend right now: 14 out of 100, #52 of 56.
  • It solved 1 of 30 coding jobs and scored 25 on reading documents. On our hardest tasks it scored 6.
  • Runs on an 8 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 3 4B got 1 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 3 4B scored 25; the best model scored 100.

TestScorePublic questionsSecret questions
Coding304
Reading documents254021
Decisions686469

On the 18 hardest tasks (included in the scores above) it scored 6. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

The Q4 download

Most people don't run the full-size model; they download a smaller Q4 version. We ran that version with Ollama on our Mac mini (M4, 16 GB) and gave it the same questions. It scored 13, about the same as the full-size model (14). You lose almost nothing by downloading it.

TestFull size (online)Q4 (our Mac)
Coding30
Reading documents2527
Hardest tasks68
Overall1413

Can your computer run it?

You can download Gemma 3 4B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBYesQ8
Graphics card, 12 GBYesQ8
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ8
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run gemma3:4b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Gemma 3 4B wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, Any

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    """Parses an INI file into a dictionary."""

    result: Dict[str, Dict[str, str]] = {}
    current_section: str = None
    current_key: str = None
    current_value: str = None
    state = "scan"  # "scan", "parse", "continue"
    
    def _interpolate(value: str, section: str) -> str:
        """Recursively interpolates keys."""
        value = value.replace('$', f'${section}:')
        
        def _interpolate_recursive(val: str) -> str:
            match = re.match(r'([^\s;#"]+|""[^"]*)"', val)
            if match:
                unquoted = match.group(1)
                if unquoted.startswith('"') and unquoted.endswith('"'):
                    return unquoted
                else:
                    return unquoted
            else:
                return val

        return _interpolate_recursive(value)
        
    for line in text.splitlines():
        line = line.strip()
        if not line:
            continue

        if line.startswith(";") or line.startswith("#"):
            continue

        if line.startswith("[") and line.endswith("]"):
            section = line[1:-1]
            if section == "DEFAULT":
                continue
            result[section] = {}
            current_section = section
            state = "parse"
            continue

        if current_section is None:
            raise ValueError("Key line before any section header")

        match = re.match(r"^\s*([a-zA-Z0-9_]+)\s*[=:]\s*(.*)", line)
        if match:
            if current_key is not None:
                raise ValueError("Malformed line: multiple key lines")
            
            key = match.group(1).lower()
            value = match.group(2).strip()
            
            if value.startswith('"') and value.endswith('"'):
                value = value[1:-1]
            
            result[current_section][key] = value
            current_key = None
            state = "parse"
            continue

        if state == "parse":
            if current_key is None:
                raise ValueError("Malformed line: value before key")
            
            if line.startswith(" "):
                value = line.strip()
                if value.startswith('"') and value.endswith('"'):
                    value = value[1:-1]
                
                if current_value is not None:
                    current_value += "\n" + value
                else:
                    current_value = value
                state = "continue"
                continue
            
            if current_value is not None:
                result[current_section][current_key] = _interpolate(current_value, current_section)
                current_value = None
                current_key = None
                state = "scan"
                continue
            else:
                raise ValueError("Malformed line: key without value")

    if current_key is not None:
        raise ValueError("Malformed line: key without value")

    return result
```
Coding: TTL LRU (hard) Didn't work
```python
import time

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = {}
        self.lru_list = []
        self.hits = 0
        self.misses = 0
        self.evictions = 0
        self.expirations = 0

    def put(self, key, value):
        if key in self.cache:
            self.cache[key] = value
            self.lru_list.remove(key)
            self.lru_list.append(key)
        else:
            if len(self.cache) >= self.capacity:
                self.evict_lru()
            self.cache[key] = value
            self.lru_list.append(key)
        now = self.clock()
        self.cache[key]['expiry'] = now + self.ttl

    def get(self, key, default=None):
        if key in self.cache:
            now = self.clock()
            if now >= self.cache[key]['expiry']:
                self.expirations += 1
                self.lru_list.remove(key)
                del self.cache[key]
                return default
            else:
                self.hits += 1
                self.lru_list.remove(key)
                self.lru_list.append(key)
                return self.cache[key]['value']
        else:
            self.misses += 1
            return default

    def __len__(self):
        count = 0
        for key in self.cache:
            now = self.clock()
            if now < self.cache[key]['expiry']:
                count += 1
        return count

    def stats(self):
        return {
            "hits": self.hits,
            "misses": self.misses,
            "evictions": self.evictions,
            "expirations": self.expirations,
        }

    def evict_lru(self):
        if self.lru_list:
            lru_key = self.lru_list.pop(0)
            del self.cache[lru_key]
            self.evictions += 1
```
Coding: Parse duration Didn't work
```python
import re

def parse_duration(s: str) -> int:
    """
    Converts a duration string into a total number of seconds.

    - Supported units: `h` (hours), `m` (minutes), `s` (seconds). Units are case-insensitive.
    - Each part is a non-negative integer followed by a unit, e.g. "1h30m", "45s", "2H", "1h 5m 10s", "90m".
    - Whitespace is allowed before, after and between parts.
    - Units must appear in the order h, m, s and each at most once.
    - Raise `ValueError` for empty/blank strings, numbers without a unit, unknown units, decimals, repeated units, or units out of order.
    """
    if not s or s.strip() == "":
        raise ValueError("Empty or blank string provided.")

    parts = re.split(r"(\d+)(\b[hm]s\b)", s)
    if len(parts) % 2 != 0:
        raise ValueError("Invalid duration string format.")

    hours = 0
    minutes = 0
    seconds = 0
    
    for i in range(0, len(parts), 2):
        try:
            num = int(parts[i])
        except ValueError:
            raise ValueError("Invalid number format.")
        
        unit = parts[i+1].lower()

        if unit not in ['h', 'm', 's']:
            raise ValueError(f"Unknown unit: {unit}")

        if unit == 'h':
            hours = num
        elif unit == 'm':
            minutes = num
        else:
            seconds = num
    
    if hours > 24:
        raise ValueError("Hours cannot exceed 24.")

    total_seconds = (hours * 3600) + (minutes * 60) + seconds

    return total_seconds
```
Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
    """Sort the numbers and remove duplicates. Then collapse runs of 3 or more consecutive integers into "a..b". Numbers in shorter runs (1 or 2 numbers) are listed individually. Join everything with "," (no spaces). An empty list returns "".

    Examples: [1,2,3,5,7,8] -> "1..3,5,7,8"; [-3,-2,-1,1] -> "-3..-1,1".
    """
    if not nums:
        return ""

    nums = sorted(list(set(nums)))
    result = []

    i = 0
    while i < len(nums):
        if i + 1 < len(nums) and nums[i] + 1 == nums[i + 1]:
            j = i + 1
            while j + 1 < len(nums) and nums[j] + 1 == nums[j + 1]:
                j += 1
            result.append(f"{nums[i]}..{nums[j]}")
            i = j + 1
        else:
            result.append(str(nums[i]))
            i += 1

    return ",".join(result)
```
Coding: Top customers Didn't work
```python
from typing import List, Tuple
import re

def top_customers(rows: List[dict], n: int) -> List[Tuple[str, float]]:
    """
    Calculates the total amount spent by each customer and returns the top n customers.

    Args:
        rows: A list of dictionaries, where each dictionary represents a row with a "customer" name and an "amount" string.
        n: The number of top customers to return.

    Returns:
        A list of tuples, where each tuple contains the customer name and total amount spent, sorted by total descending,
        then name ascending.
    """

    customer_totals = {}
    for row in rows:
        customer_name = row.get("customer", "").strip()
        amount_str = row.get("amount", "")

        if not customer_name or not amount_str:
            continue

        try:
            amount = float(amount_str.replace("$", "").replace(",", ""))
            if amount < 0:
                amount = -amount  # Handle refunds

            if customer_name not in customer_totals:
                customer_totals[customer_name] = 0.0
            customer_totals[customer_name] += amount
        except ValueError:
            continue

    sorted_customers = sorted(customer_totals.items(), key=lambda item: (-item[1], item[0]))
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import re

def slugify(title: str, max_len: int = 50) -> str:
    """
    Builds a URL slug.

    1. Transliterate accented characters to ASCII (é -> e, ü -> u) and drop any other non-ASCII characters.
    2. Lowercase.
    3. Replace every run of characters other than a-z and 0-9 with a single "-", and strip leading/trailing "-".
    4. If the slug is longer than `max_len`, shorten it without cutting a word: keep the longest prefix of whole words (words are separated by "-") whose length is <= max_len. If even the first word is longer than max_len, hard-cut it to max_len characters.
    5. The result never ends with "-".
    """
    title = title.lower()
    title = re.sub(r'[^\w\s-]', '', title)
    title = re.sub(r'[-\s]+', '-', title)
    title = title.strip('-')

    if len(title) > max_len:
        parts = title.split('-')
        parts = [part for part in parts if len(part) <= max_len]
        title = '-'.join(parts)
        if title and title[-1] == '-':
            title = title[:-1]

    return title
```
Coding: Token bucket Didn't work
```python
import time

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.tokens = capacity
        self.last_refill = self.clock()

    def refill(self):
        now = self.clock()
        elapsed = now - self.last_refill
        if elapsed > 0:
            refill_amount = elapsed * self.refill_per_sec
            self.tokens = min(self.capacity, self.tokens + refill_amount)
            self.last_refill = now

    def allow(self, cost: float = 1) -> bool:
        self.refill()
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        else:
            return False

    @property
    def tokens(self):
        self.refill()
        return self.tokens
```
Decisions: Refund window (hard) 100% right
```json
{
  "outcome": {
    "full_refund": 0.1,
    "partial_refund": 0.3,
    "store_credit": 0.6,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Incident matrix (hard) 0% right
```json
{
  "severity": "3",
  "page": "true"
}
```
Decisions: Tool followup (hard) 100% right
```json
{
  "tool": {
    "calendar": 0.95,
    "none": 0.05
  },
  "confirm": {
    "true": 0.8,
    "false": 0.2
  }
}
```
Decisions: Legit security alert (hard) 0% right
```json
{
  "phishing": {
    "true": 0.8,
    "false": 0.2
  },
  "action_needed": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Meeting slot (hard) 50% right
```json
{
  "slot": {
    "A": 0.1,
    "B": 0.2,
    "C": 0.6,
    "D": 0.1
  },
  "raj_last": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Review mixed (hard) 100% right
```json
{
  "hardware": {
    "true": 0.95,
    "false": 0.05
  },
  "support": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Support checkout down 67% right
```json
{
  "department": {
    "billing": 0.2,
    "technical": 0.6,
    "account": 0.1,
    "shipping": 0.05,
    "sales": 0.05
  },
  "urgency": {
    "0": 0.05,
    "1": 0.2,
    "2": 0.5,
    "3": 0.25
  },
  "outage": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Refund wrong plan 33% right
```json
{
  "department": {
    "billing": 0.85,
    "technical": 0.05,
    "account": 0.05,
    "shipping": 0.0,
    "sales": 0.0
  },
  "refund": "true",
  "tone": "calm"
}
```
Decisions: Moderation doxxing 0% right
```json
{
  "policy": {
    "none": 0.7,
    "harassment": 0.3
  },
  "personal_info": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Route calendar 100% right
```json
{
  "tool": {
    "calendar": 0.95,
    "none": 0.05
  },
  "confirm": {
    "true": 0.98,
    "false": 0.02
  }
}
```
Decisions: Doc invoice missing due 100% right
```json
{
  "doc_type": {
    "invoice": 0.99,
    "resume": 0.01,
    "contract": 0.01,
    "bank_statement": 0.01,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 0.8,
    "false": 0.2
  }
}
```
Decisions: Phishing paypal 50% right
```json
{
  "phishing": {
    "true": 0.95,
    "false": 0.05
  },
  "risk": {
    "0": 0.05,
    "1": 0.15,
    "2": 0.60,
    "3": 0.20
  }
}
```
Decisions: Pii ssn email 100% right
```json
{
  "data_kind": {
    "government_id": 0.95,
    "contact": 0.05,
    "financial": 0.0,
    "health": 0.0
  },
  "sensitive": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Review mixed 100% right
```json
{
  "sentiment": {
    "positive": 0.1,
    "neutral": 0.2,
    "negative": 0.7
  },
  "defect": {
    "true": 0.95,
    "false": 0.05
  },
  "recommend": {
    "true": 0.05,
    "false": 0.95
  }
}
```
Documents: Saas escalator (hard) 10% right
```json
{
  "year2_price_per_seat_month": 45.00,
  "year3_price_per_seat_month": 45.00,
  "year1_invoice": 5400.00,
  "year2_invoice": 5400.00,
  "addon_months_billed": 160,
  "addon_invoice": 7200.00,
  "year3_invoice": 5400.00,
  "year3_discount_percent": 15.00,
  "total_contract_value": 21600.00,
  "contract_end_date": "2026-02-28"
}
```
Documents: Expense thread 71% right
```json
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 36.00
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 212.40
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 432.00
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 41.50
    }
  ],
  "rejected_item_count": 2,
  "per_diem_days": 3,
  "per_diem_usd": 205.00,
  "total_reimbursable_usd": 2002.50,
  "approver_email": "priya.raman@corvane.com"
}
```
Documents: Lease amendment 67% right
```json
{
  "tenants": [
    "Marcus Lin",
    "Sofia Lin"
  ],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2295.00,
  "late_fee_from_2025_06_01": 111.75,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2406.75,
  "move_in_payment": 6400.00
}
```
Documents: Ticket SLA 54% right
```json
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": [
    "billing_address",
    "invoice_pdf"
  ],
  "affected_orders": [
    "SO-99812",
    "SO-99815",
    "SO-99820",
    "SO-99827"
  ],
  "priority": "P2",
  "sla_due_local": "2025-09-13T09:00",
  "sla_due_utc": "2025-09-13T09:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
```
Documents: Sales footnotes 0% right
```json
{
  "q3_total_usd": 14,763,
  "q2_total_usd": 11,962,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 23.7,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["Central"],
  "international_q3_organic_usd": 1951,
  "west_excluding_mountain_q3_usd": 4201
}
```

Size: 4.3B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.