Review · updated OCT 11

Gemma 3 12B review: not one we'd recommend right now

It scored 35 out of 100, #43 of 56. It solved 9 of 30 coding jobs and scored 40 on reading documents. Runs on a 12 GB graphics card or a Mac with 16 GB.

The short version
  • Gemma 3 12B is a free model from Google that you can run on your own computer. In our tests it's not one we'd recommend right now: 35 out of 100, #43 of 56.
  • It solved 9 of 30 coding jobs and scored 40 on reading documents. On our hardest tasks it scored 24.
  • Runs on a 12 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 3 12B got 9 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 3 12B scored 40; the best model scored 100.

TestScorePublic questionsSecret questions
Coding302930
Reading documents405636
Decisions839679

On the 18 hardest tasks (included in the scores above) it scored 24. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Gemma 3 12B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBYesQ5
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ5
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run gemma3:12b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Gemma 3 12B wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    """Parses an INI dialect string and returns a dictionary."""

    result: dict[str, dict[str, str]] = {}
    default_section: dict[str, str] = {}
    current_section: str | None = None
    previous_value: str | None = None

    for line in text.splitlines():
        line = line.strip()
        if not line or line.startswith(';') or line.startswith('#'):
            if previous_value is not None:
                previous_value = None
            continue

        match = re.match(r"^\s*\[(\w+)\]\s*$", line)
        if match:
            if current_section is not None:
                result[current_section] = section
            current_section = match.group(1)
            if current_section == "DEFAULT":
                section = default_section
            else:
                section = {}
                result[current_section] = section
            previous_value = None
            continue

        match = re.match(r"^(\w+)\s*(=|:)\s*(.*)$", line)
        if match:
            key = match.group(1).strip().lower()
            sep = match.group(2)
            value = match.group(3).strip()

            if current_section is None:
                raise ValueError("Key line before section header")

            if value.startswith('"') and value.endswith('"'):
                value = value[1:-1]
            else:
                value = interpolate(value, result, current_section)

            section[key] = value
            previous_value = value
            continue

        if previous_value is not None:
            if line.startswith(" "):
                previous_value += "\n" + line.strip()
            else:
                previous_value = None
        else:
            raise ValueError("Malformed line")

    if current_section is not None:
        result[current_section] = section

    # Merge DEFAULT section into other sections
    for key, value in default_section.items():
        for section_name, section in result.items():
            if section_name != "DEFAULT" and key not in section:
                section[key] = value

    return result


def interpolate(value: str, result: dict[str, dict[str, str]], current_section: str) -> str:
    """Interpolates variables in a value."""

    while True:
        match = re.search(r"\${(\w+)}", value)
        if not match:
            break

        key = match.group(1)
        try:
            if current_section == "DEFAULT":
                replacement = result["DEFAULT"].get(key)
            else:
                replacement = result[current_section].get(key)
            if replacement is None:
                raise KeyError(f"Key '{key}' not found")
            value = value.replace("${" + key + "}", replacement, 1)
        except KeyError as e:
            raise e
        except ValueError as e:
            raise e


    match = re.search(r"\${(\w+):(\w+)}", value)
    if match:
        section = match.group(1)
        key = match.group(2)
        try:
            if section not in result:
                raise KeyError(f"Section '{section}' not found")
            replacement = result[section].get(key)
            if replacement is None:
                raise KeyError(f"Key '{key}' not found in section '{section}'")

            value = value.replace("${" + section + ":" + key + "}", replacement, 1)
        except KeyError as e:
            raise e
        except ValueError as e:
            raise e

    return value
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        if capacity < 0:
            raise ValueError("Capacity must be non-negative")
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.hits = 0
        self.misses = 0
        self.evictions = 0
        self.expirations = 0

    def put(self, key, value):
        now = self.clock()
        if self.capacity == 0:
            return

        if key in self.cache:
            self.cache.move_to_end(key)
            self.cache[key] = (value, now + self.ttl)
        else:
            while len(self.cache) >= self.capacity:
                self._remove_expired()
                if len(self.cache) >= self.capacity:
                    self.evictions += 1
                    self.cache.popitem(last=False)  # LRU eviction

            self.cache[key] = (value, now + self.ttl)

    def get(self, key, default=None):
        now = self.clock()
        if self.capacity == 0:
            self.misses += 1
            return default

        if key in self.cache:
            value, expiry = self.cache.pop(key)
            if now >= expiry:
                self.expirations += 1
                self.misses += 1
                return default
            else:
                self.cache[key] = (value, now + self.ttl)
                self.cache.move_to_end(key)
                self.hits += 1
                return value
        else:
            self.misses += 1
            return default

    def __len__(self):
        now = self.clock()
        count = 0
        temp_cache = OrderedDict()
        for key, (value, expiry) in self.cache.items():
            if now < expiry:
                count += 1
                temp_cache[key] = (value, expiry)
        self.cache = temp_cache
        return count

    def stats(self):
        return {
            "hits": self.hits,
            "misses": self.misses,
            "evictions": self.evictions,
            "expirations": self.expirations,
        }

    def _remove_expired(self):
        now = self.clock()
        keys_to_remove = []
        for key, (value, expiry) in self.cache.items():
            if now >= expiry:
                keys_to_remove.append(key)

        for key in keys_to_remove:
            self.cache.pop(key)
            self.expirations += 1
```
Coding: Parse duration Didn't work
```python
import re

def parse_duration(s: str) -> int:
    """Converts a duration string into a total number of seconds.

    - Supported units: `h` (hours), `m` (minutes), `s` (seconds). Units are case-insensitive.
    - Each part is a non-negative integer followed by a unit, e.g. "1h30m", "45s", "2H", "1h 5m 10s", "90m".
    - Whitespace is allowed before, after and between parts.
    - Units must appear in the order h, m, s and each at most once.
    - Raise `ValueError` for empty/blank strings, numbers without a unit, unknown units, decimals, repeated units, or units out of order.
    """
    s = s.strip()
    if not s:
        raise ValueError("Empty string")

    total_seconds = 0
    parts = re.findall(r"(\d+)\s*([hms])", s, re.IGNORECASE)

    if not parts:
        raise ValueError("No valid duration parts found")

    h_found, m_found, s_found = False, False, False

    for num, unit in parts:
        num = int(num)
        unit = unit.lower()

        if h_found or m_found or s_found:
            raise ValueError("Units out of order")

        if unit == 'h':
            total_seconds += num * 3600
            h_found = True
        elif unit == 'm':
            total_seconds += num * 60
            m_found = True
        elif unit == 's':
            total_seconds += num
            s_found = True
        else:
            raise ValueError(f"Unknown unit: {unit}")

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""

    nums.sort()
    unique_nums = []
    for num in nums:
        if not unique_nums or num != unique_nums[-1]:
            unique_nums.append(num)

    result = []
    i = 0
    while i < len(unique_nums):
        start = unique_nums[i]
        j = i + 1
        while j < len(unique_nums) and unique_nums[j] == unique_nums[j - 1] + 1:
            j += 1

        if j - i >= 3:
            result.append(f"{start}..{unique_nums[j - 1]}")
        elif j - i == 2:
            result.append(f"{start},{unique_nums[j - 1]}")
        else:
            result.append(str(start))

        i = j

    return ",".join(result)
```
Coding: Top customers Didn't work
```python
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    customer_totals = {}
    for row in rows:
        customer = row.get("customer")
        amount_str = row.get("amount")

        if not customer or not amount_str:
            continue

        customer = customer.strip()
        if not customer:
            continue

        try:
            amount = float(amount_str.replace("$", "").replace(",", ""))
        except ValueError:
            continue

        if customer not in customer_totals:
            customer_totals[customer] = 0.0

        customer_totals[customer] += amount

    sorted_customers = sorted(
        customer_totals.items(),
        key=lambda item: (-item[1], item[0]),
    )

    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    """Builds a URL slug."""

    title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
    title = title.lower()
    title = re.sub(r'[^\w\s-]', '', title).strip()
    title = re.sub(r'[-\s]+', '-', title).strip('-_')

    if len(title) > max_len:
        parts = title.split('-')
        while len(title) > max_len and parts:
            title = '-'.join(parts[:-1])
            parts = title.split('-')
        if len(title) > max_len:
            title = title[:max_len]

    if title.endswith('-'):
        title = title[:-1]

    return title

import re
```
Coding: Token bucket Worked
```python
import time

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._tokens = capacity
        self._last_refill_time = clock()

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        else:
            return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens

    def _refill(self):
        now = self._clock()
        elapsed_time = now - self._last_refill_time
        refill_amount = elapsed_time * self._refill_per_sec
        self._tokens = min(self._capacity, self._tokens + refill_amount)
        self._last_refill_time = now
```
Decisions: Refund window (hard) 100% right
```json
{
  "outcome": {
    "full_refund": 0.05,
    "partial_refund": 0.25,
    "store_credit": 0.60,
    "nothing": 0.10
  },
  "defective": {
    "true": 0.05,
    "false": 0.95
  }
}
```
Decisions: Incident matrix (hard) 100% right
```json
{
  "severity": {
    "0": 0.1,
    "1": 0.4,
    "2": 0.3,
    "3": 0.2
  },
  "page": {
    "true": 0.2,
    "false": 0.8
  }
}
```
Decisions: Tool followup (hard) 100% right
```json
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 0.95,
    "email": 0.05,
    "none": 0.0
  },
  "confirm": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Legit security alert (hard) 100% right
```json
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.3,
    "false": 0.7
  }
}
```
Decisions: Meeting slot (hard) 50% right
```json
{
  "slot": {
    "A": 0.05,
    "B": 0.15,
    "C": 0.65,
    "D": 0.15
  },
  "raj_last": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Review mixed (hard) 100% right
```json
{
  "hardware": {
    "true": 0.95,
    "false": 0.05
  },
  "support": {
    "true": 0.98,
    "false": 0.02
  }
}
```
Decisions: Support checkout down 100% right
```json
{
  "department": {
    "billing": 0.05,
    "technical": 0.90,
    "account": 0.03,
    "shipping": 0.01,
    "sales": 0.01
  },
  "urgency": {
    "0": 0.00,
    "1": 0.00,
    "2": 0.05,
    "3": 0.95
  },
  "outage": {
    "true": 0.99,
    "false": 0.01
  }
}
```
Decisions: Refund wrong plan 100% right
```json
{
  "department": {
    "billing": 0.95,
    "technical": 0.01,
    "account": 0.03,
    "shipping": 0.00,
    "sales": 0.01
  },
  "refund": {
    "true": 0.99,
    "false": 0.01
  },
  "tone": {
    "frustrated": 0.05,
    "calm": 0.95
  }
}
```
Decisions: Moderation doxxing 100% right
```json
{
  "policy": {
    "none": 0.1,
    "harassment": 0.8,
    "hate": 0.05,
    "spam": 0.05,
    "self_harm": 0.0
  },
  "personal_info": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Route calendar 100% right
```json
{
  "tool": {
    "web_search": 0.05,
    "calculator": 0.01,
    "calendar": 0.90,
    "email": 0.03,
    "none": 0.01
  },
  "confirm": {
    "true": 0.98,
    "false": 0.02
  }
}
```
Decisions: Doc invoice missing due 100% right
```json
{
  "doc_type": {
    "invoice": 0.95,
    "resume": 0.01,
    "contract": 0.02,
    "bank_statement": 0.01,
    "other": 0.01
  },
  "missing_due_date": {
    "true": 0.98,
    "false": 0.02
  }
}
```
Decisions: Phishing paypal 100% right
```json
{
  "phishing": {
    "true": 0.99,
    "false": 0.01
  },
  "risk": {
    "0": 0.00,
    "1": 0.00,
    "2": 0.01,
    "3": 0.99
  }
}
```
Decisions: Pii ssn email 100% right
```json
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.1,
    "financial": 0.0,
    "government_id": 0.9,
    "health": 0.0
  },
  "sensitive": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Review mixed 100% right
```json
{
  "sentiment": {
    "positive": 0.1,
    "neutral": 0.1,
    "negative": 0.8
  },
  "defect": {
    "true": 0.9,
    "false": 0.1
  },
  "recommend": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Documents: Saas escalator (hard) 30% right
```json
{
  "year2_price_per_seat_month": 47.25,
  "year3_price_per_seat_month": 47.25,
  "year1_invoice": 518400.00,
  "year2_invoice": 566400.00,
  "addon_months_billed": 10,
  "addon_invoice": 76800.00,
  "year3_invoice": 566400.00,
  "year3_discount_percent": 0.0,
  "total_contract_value": 1633600.00,
  "contract_end_date": "2027-02-28"
}
```
Documents: Expense thread 79% right
```json
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 39.24
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.52
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 2,
  "per_diem_usd": 130.00,
  "total_reimbursable_usd": 2114.53,
  "approver_email": "priya.raman@corvane.com"
}
```
Documents: Lease amendment 67% right
```json
{
  "tenants": [
    "Marcus Lin",
    "Sofia Lin"
  ],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2234.00,
  "late_fee_from_2025_06_01": 111.70,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2309.00,
  "move_in_payment": 4500.00
}
```
Documents: Ticket SLA 82% right
```json
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": [
    "billing_address"
  ],
  "affected_orders": [
    "SO-99812",
    "SO-99820",
    "SO-99827"
  ],
  "priority": "P2",
  "sla_due_local": "2025-09-13T01:30",
  "sla_due_utc": "2025-09-13T05:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
```
Documents: Sales footnotes 22% right
```json
{
  "q3_total_usd": 15300,
  "q2_total_usd": 14469,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 5.7,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": [
    "Central"
  ],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 4201
}
```

Size: 12B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.