Review · updated OCT 11

Gemma 4 31B review: great at code, weaker with documents

It scored 79 out of 100, #14 of 56. It solved 28 of 30 coding jobs and scored 65 on reading documents. Runs on a 32 GB graphics card or a Mac with 48 GB.

The short version
  • Gemma 4 31B is a free model from Google that you can run on your own computer. In our tests it's great at code, weaker with documents: 79 out of 100, #14 of 56.
  • It solved 28 of 30 coding jobs and scored 65 on reading documents. On our hardest tasks it scored 73.
  • Runs on a 32 GB graphics card or a Mac with 48 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 4 31B got 28 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 4 31B scored 65; the best model scored 100.

TestScorePublic questionsSecret questions
Coding9310091
Reading documents657961
Decisions939692

On the 18 hardest tasks (included in the scores above) it scored 73. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Gemma 4 31B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBYesQ5
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBYesQ6
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Gemma 4 31B GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Gemma 4 31B wrote for our public questions, word for word.

Coding: INI parser (hard) Worked
```python
import re

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    sections_raw = {}
    current_section = None
    last_key = None

    lines = text.splitlines()
    
    for line_num, line in enumerate(lines, 1):
        stripped_line = line.strip()
        
        # Rule 3: Ignore empty or full-line comments
        if not stripped_line or stripped_line[0] in ';#':
            continue
            
        # Rule 6: Continuation lines (start with whitespace)
        if line and not line[0].isspace() is False and current_section is not None and last_key is not None:
            # This is a continuation line
            val = _strip_inline_comment(stripped_line)
            sections_raw[current_section][last_key] += "\n" + val
            continue

        # Rule 1: Section headers
        if stripped_line.startswith('[') and stripped_line.endswith(']'):
            current_section = stripped_line[1:-1].strip()
            if current_section not in sections_raw:
                sections_raw[current_section] = {}
            last_key = None
            continue

        # Rule 2: Key-Value pairs
        delimiter = None
        eq_idx = line.find('=')
        col_idx = line.find(':')
        
        if eq_idx != -1 and (col_idx == -1 or eq_idx < col_idx):
            delimiter = '='
            split_idx = eq_idx
        elif col_idx != -1:
            delimiter = ':'
            split_idx = col_idx
        else:
            # Rule 10: Malformed line
            raise ValueError(f"Malformed line at {line_num}: {line}")

        if current_section is None:
            # Rule 10: Key before section
            raise ValueError(f"Key found before section header at line {line_num}")

        key_part = line[:split_idx].strip().lower()
        val_part = line[split_idx + 1:].strip()
        
        # Rule 4 & 5: Handle quotes and inline comments
        value = _parse_value(val_part)
        
        sections_raw[current_section][key_part] = value
        last_key = key_part

    # Rule 8: Handle DEFAULT inheritance
    default_vals = sections_raw.pop('DEFAULT', {})
    result = {}
    for section, keys in sections_raw.items():
        merged = default_vals.copy()
        merged.update(keys)
        result[section] = merged

    # Rule 9: Interpolation
    memo = {}
    path = set()

    def resolve(section, key):
        state_key = (section, key)
        if state_key in memo:
            return memo[state_key]
        if state_key in path:
            raise ValueError("Interpolation cycle detected")
        
        path.add(state_key)
        
        # Get the raw value
        try:
            val = result[section][key]
        except KeyError:
            raise KeyError(f"Key '{key}' not found in section '{section}'")

        # Recursive interpolation using regex
        def replace_match(match):
            content = match.group(1)
            if ':' in content:
                s_name, k_name = content.split(':', 1)
                if s_name not in result:
                    raise KeyError(f"Section '{s_name}' not found")
                return resolve(s_name, k_name)
            else:
                return resolve(section, content)

        final_val = re.sub(r'\$\{(.*?)\}', replace_match, val)
        path.remove(state_key)
        memo[state_key] = final_val
        return final_val

    # Final pass to resolve all values
    final_result = {}
    for section in result:
        final_result[section] = {}
        for key in result[section]:
            final_result[section][key] = resolve(section, key)

    return final_result

def _strip_inline_comment(text: str) -> str:
    # Simple strip for continuation lines as they aren't quoted wrappers
    # but still follow the " whitespace + #/; " rule
    match = re.search(r'\s+[;#]', text)
    if match:
        return text[:match.start()].strip()
    return text.strip()

def _parse_value(val: str) -> str:
    # Rule 5: Double quoted values
    if val.startswith('"') and val.endswith('"'):
        return val[1:-1]
    
    # Rule 4: Inline comments (whitespace followed by ; or #)
    # We use a regex to find the first occurrence of ' #' or ' ;'
    match = re.search(r'\s+[;#]', val)
    if match:
        return val[:match.start()].strip()
    
    return val.strip()
```
Coding: TTL LRU (hard) Worked
```python
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity, ttl, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()  # key -> (value, expiry)
        self.stats_counts = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def put(self, key, value):
        if self.capacity <= 0:
            return

        now = self.clock()
        expiry = now + self.ttl

        if key in self.cache:
            self.cache.move_to_end(key)
            self.cache[key] = (value, expiry)
            return

        # Remove all expired entries before checking capacity
        self._cleanup_expired(now)

        if len(self.cache) >= self.capacity:
            # Evict Least Recently Used (first item in OrderedDict)
            self.cache.popitem(last=False)
            self.stats_counts["evictions"] += 1

        self.cache[key] = (value, expiry)

    def get(self, key, default=None):
        if self.capacity <= 0:
            self.stats_counts["misses"] += 1
            return default

        now = self.clock()
        if key not in self.cache:
            self.stats_counts["misses"] += 1
            return default

        val, expiry = self.cache[key]
        if now >= expiry:
            del self.cache[key]
            self.stats_counts["expirations"] += 1
            self.stats_counts["misses"] += 1
            return default

        self.cache.move_to_end(key)
        self.stats_counts["hits"] += 1
        return val

    def _cleanup_expired(self, now):
        # We must iterate over a copy of keys because we are modifying the dict
        keys = list(self.cache.keys())
        for k in keys:
            _, expiry = self.cache[k]
            if now >= expiry:
                del self.cache[k]
                self.stats_counts["expirations"] += 1

    def __len__(self):
        now = self.clock()
        count = 0
        for _, expiry in self.cache.values():
            if now < expiry:
                count += 1
        return count

    def stats(self):
        return self.stats_counts.copy()
```
Coding: Parse duration Worked
```python
import re

def parse_duration(s: str) -> int:
    """
    Converts a duration string (e.g., "1h 30m 10s") into total seconds.
    Raises ValueError for invalid formats, repeated units, or incorrect order.
    """
    s = s.strip()
    if not s:
        raise ValueError("Duration string cannot be empty or blank.")

    # Pattern to identify segments of digits followed by a unit
    # This allows us to check for unexpected characters between valid parts
    pattern = re.compile(r'(\d+)([hms])', re.IGNORECASE)
    
    # Check for decimals or non-digit characters that aren't the allowed units/whitespace
    # We strip the valid matches to see if anything "illegal" remains
    remaining = s
    matches = pattern.findall(s)
    for val, unit in matches:
        remaining = remaining.replace(val + unit, '', 1)
    
    if remaining.strip():
        # If there are characters left that aren't whitespace, it's an invalid format
        # (e.g., decimals '.', unknown units 'x', or numbers without units)
        raise ValueError("Invalid duration format or unknown units encountered.")

    if not matches:
        raise ValueError("No valid duration parts found.")

    units_map = {'h': 3600, 'm': 60, 's': 1}
    unit_order = {'h': 0, 'm': 1, 's': 2}
    
    total_seconds = 0
    last_unit_index = -1
    seen_units = set()

    for value_str, unit_char in matches:
        unit_lower = unit_char.lower()
        
        if unit_lower in seen_units:
            raise ValueError(f"Repeated unit: {unit_lower}")
        
        current_unit_index = unit_order[unit_lower]
        if current_unit_index <= last_unit_index:
            raise ValueError("Units must appear in the order h, m, s.")
        
        total_seconds += int(value_str) * units_map[unit_lower]
        seen_units.add(unit_lower)
        last_unit_index = current_unit_index

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""

    # Sort and remove duplicates
    sorted_nums = sorted(set(nums))
    n = len(sorted_nums)
    result = []
    i = 0

    while i < n:
        start = i
        # Find the end of the consecutive run
        while i + 1 < n and sorted_nums[i + 1] == sorted_nums[i] + 1:
            i += 1
        
        end = i
        run_length = end - start + 1
        
        if run_length >= 3:
            # Collapse run of 3 or more
            result.append(f"{sorted_nums[start]}..{sorted_nums[end]}")
        else:
            # List individually for runs of 1 or 2
            for j in range(start, end + 1):
                result.append(str(sorted_nums[j]))
        
        i += 1

    return ",".join(result)
```
Coding: Top customers Worked
```python
import collections

def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    def parse_amount(amount_str: str) -> float:
        if not isinstance(amount_str, str):
            raise ValueError("Amount must be a string")
        # Remove currency symbols and thousands separators
        clean_str = amount_str.replace('$', '').replace(',', '').strip()
        return float(clean_str)

    # Store totals keyed by lowercase name
    totals = collections.defaultdict(float)
    # Store the original trimmed name for the first encounter of each customer
    name_map = {}

    for row in rows:
        raw_name = row.get("customer")
        raw_amount = row.get("amount")

        if not isinstance(raw_name, str) or not isinstance(raw_amount, str):
            continue

        trimmed_name = raw_name.strip()
        if not trimmed_name:
            continue

        try:
            amount = parse_amount(raw_amount)
            key = trimmed_name.lower()
            
            if key not in name_map:
                name_map[key] = trimmed_name
            
            totals[key] += amount
        except (ValueError, TypeError):
            continue

    # Prepare final list of (original_name, rounded_total)
    result = []
    for key, total in totals.items():
        result.append((name_map[key], round(total, 2)))

    # Sort by total descending, then name ascending
    result.sort(key=lambda x: (-x[1], x[0]))

    return result[:n]
```
Coding: Slugify Worked
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    # 1. Transliterate accented characters to ASCII and drop non-ASCII
    title = unicodedata.normalize('NFKD', title)
    title = title.encode('ascii', 'ignore').decode('ascii')
    
    # 2. Lowercase
    title = title.lower()
    
    # 3. Replace non-alphanumeric runs with "-", strip leading/trailing "-"
    title = re.sub(r'[^a-z0-9]+', '-', title).strip('-')
    
    # 4 & 5. Length constraints
    if len(title) <= max_len:
        return title
    
    # Try to find the longest prefix of whole words
    words = title.split('-')
    current_slug = []
    current_len = 0
    
    for word in words:
        # Length of word plus the hyphen (if not the first word)
        added_len = len(word) + (1 if current_slug else 0)
        if current_len + added_len <= max_len:
            current_slug.append(word)
            current_len += added_len
        else:
            break
            
    if current_slug:
        return '-'.join(current_slug)
    else:
        # Even the first word is longer than max_len, hard-cut it
        return title[:max_len].rstrip('-')
```
Coding: Token bucket Worked
```python
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._tokens = capacity
        self._last_update = self._clock()

    def _refill(self) -> None:
        now = self._clock()
        elapsed = now - self._last_update
        if elapsed > 0:
            new_tokens = elapsed * self._refill_per_sec
            self._tokens = min(self._capacity, self._tokens + new_tokens)
            self._last_update = now

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens
```
Decisions: Refund window (hard) 100% right
```json
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.0,
    "store_credit": 1.0,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Incident matrix (hard) 100% right
```json
{
  "severity": {
    "0": 0.0,
    "1": 1.0,
    "2": 0.0,
    "3": 0.0
  },
  "page": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Tool followup (hard) 50% right
{
  "tool": {
    "web_search": 0,
    "calculator": 0,
    "calendar": 1,
    "email": 0,
    "none": 0
  },
  "confirm": {
    "true": 0,
    "false": 1
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.0,
    "false": 1.0
  },
  "action_needed": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Meeting slot (hard) 100% right
```json
{
  "slot": {
    "A": 0,
    "B": 1,
    "C": 0,
    "D": 0
  },
  "raj_last": {
    "true": 1,
    "false": 0
  }
}
```
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 1.0,
    "false": 0.0
  },
  "support": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0.0,
    "technical": 1.0,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  },
  "outage": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 1.0,
    "technical": 0.0,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "refund": {
    "true": 1.0,
    "false": 0.0
  },
  "tone": {
    "frustrated": 0.0,
    "calm": 1.0
  }
}
Decisions: Moderation doxxing 100% right
```json
{
  "policy": {
    "none": 0.0,
    "harassment": 1.0,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.0
  },
  "personal_info": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0,
    "calculator": 0,
    "calendar": 1,
    "email": 0,
    "none": 0
  },
  "confirm": {
    "true": 1,
    "false": 0
  }
}
Decisions: Doc invoice missing due 100% right
```json
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Phishing paypal 100% right
{
  "phishing": {
    "true": 1.0,
    "false": 0.0
  },
  "risk": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0,
    "contact": 0,
    "financial": 0,
    "government_id": 1,
    "health": 0
  },
  "sensitive": {
    "true": 1,
    "false": 0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.0,
    "neutral": 0.1,
    "negative": 0.9
  },
  "defect": {
    "true": 1.0,
    "false": 0.0
  },
  "recommend": {
    "true": 0.0,
    "false": 1.0
  }
}
Documents: Saas escalator (hard) 40% right
{
  "year2_price_per_seat_month": 46.53,
  "year3_price_per_seat_month": 48.85,
  "year1_invoice": 58320.00,
  "year2_invoice": 60154.80,
  "addon_months_billed": 6,
  "addon_invoice": 24040.32,
  "year3_invoice": 159264.00,
  "year3_discount_percent": 15,
  "total_contract_value": 301779.12,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 92% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.60
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 65.00,
  "total_reimbursable_usd": 2158.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right
{
  "tenants": [
    "Marcus Lin",
    "Sofia Lin"
  ],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 4700.00
}
Documents: Ticket SLA 75% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": [
    "billing_address",
    "invoice_pdf"
  ],
  "affected_orders": [
    "SO-99812",
    "SO-99820",
    "SO-99827"
  ],
  "priority": "P2",
  "sla_due_local": "2025-09-15T11:30",
  "sla_due_utc": "2025-09-15T16:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 89% right
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": [
    "East"
  ],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 31B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.