Review · updated OCT 11

Gemma 3 27B review: not one we'd recommend right now

It scored 40 out of 100, #38 of 56. It solved 10 of 30 coding jobs and scored 46 on reading documents. Runs on a 24 GB graphics card or a Mac with 32 GB.

The short version
  • Gemma 3 27B is a free model from Google that you can run on your own computer. In our tests it's not one we'd recommend right now: 40 out of 100, #38 of 56.
  • It solved 10 of 30 coding jobs and scored 46 on reading documents. On our hardest tasks it scored 29.
  • Runs on a 24 GB graphics card or a Mac with 32 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Gemma 3 27B got 10 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Gemma 3 27B scored 46; the best model scored 100.

TestScorePublic questionsSecret questions
Coding332935
Reading documents466541
Decisions889387

On the 18 hardest tasks (included in the scores above) it scored 29. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Gemma 3 27B in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBYesQ5
Graphics card, 32 GBYesQ8
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBYesQ4
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run gemma3:27b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Gemma 3 27B wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
.
```python
import re
from typing import Dict

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    """Parses an INI string according to the specified rules."""

    result: Dict[str, Dict[str, str]] = {}
    default_section: Dict[str, str] = {}
    current_section: str = None
    current_key: str = None
    current_value: str = None
    in_string: bool = False
    interpolation_stack: list[tuple[str, str]] = []

    for line in text.splitlines():
        line = line.strip()

        if not line or line.startswith((';', '#')):
            continue

        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if current_section is not None:
                if current_key:
                    if current_value is None:
                        current_value = ''
                    result[current_section][current_key] = current_value
                
            current_section = section_name
            result[current_section] = {}
            current_key = None
            current_value = None
            in_string = False
            interpolation_stack = []
        else:
            if current_section is None:
                raise ValueError("Key line before section header")

            match = re.match(r'(\w+)\s*[:=]\s*(.+)', line)
            if not match:
                raise ValueError("Malformed line")
            
            key, value = match.groups()
            key = key.lower().strip()
            
            if current_key is None:
                current_key = key
                
                in_string = False
                
                # Extract value, handling quotes and comments
                value_start = 0
                while value_start < len(value):
                    if value[value_start] == '"':
                        value_start += 1
                        value_end = value.find('"', value_start)
                        if value_end == -1:
                            current_value = value[value_start-1:]
                            break
                        current_value = value[value_start : value_end]
                        value_start = value_end + 1
                    else:
                        comment_start = value.find(';', value_start)
                        hash_comment_start = value.find('#', value_start)
                        
                        if comment_start != -1 and (hash_comment_start == -1 or comment_start < hash_comment_start):
                            current_value = value[value_start:comment_start].strip()
                            break
                        elif hash_comment_start != -1:
                            current_value = value[value_start:hash_comment_start].strip()
                            break
                        else:
                            current_value = value[value_start:].strip()
                            break
                
            else:
                # Continuation line
                current_value += "\n" + value

    if current_section is not None and current_key is not None:
        if current_value is None:
            current_value = ''
        result[current_section][current_key] = current_value

    
    #Apply DEFAULT section
    for key, value in default_section.items():
        if key not in result.get('DEFAULT', {}):
            result['DEFAULT'] = result.get('DEFAULT',{})
            result['DEFAULT'][key] = value

    final_result: Dict[str, Dict[str, str]] = {}
    for section_name, section_data in result.items():
        if section_name != 'DEFAULT':
            #Inherit from default
            inherited_data = default_section.copy()
            for key, value in section_data.items():
                inherited_data[key] = value
            final_result[section_name] = inherited_data
        else:
            final_result[section_name] = section_data

    
    for section_name, section_data in final_result.items():
        for key, value in section_data.items():
            final_result[section_name][key] = interpolate(value, final_result)

    return final_result



def interpolate(value: str, data: Dict[str, Dict[str, str]]) -> str:
    """Interpolates values in a string using the provided data."""
    
    def replace_match(match):
        full_match = match.group(0)
        if full_match.startswith("${"):
            ref = full_match[2:-1]
            if ":" in ref:
                section, key = ref.split(":", 1)
                if section in data and key in data[section]:
                    return data[section][key]
                else:
                    raise KeyError(f"Section or key not found: {section}:{key}")
            else:
                if "DEFAULT" in data and ref in data["DEFAULT"]:
                    return data["DEFAULT"][ref]
                else:
                    raise KeyError(f"Key not found: {ref}")
        else:
            return match.group(0)

    pattern = r"\${[^}]+}"
    
    while True:
        original_value = value
        value = re.sub(pattern, replace_match, value)
        if value == original_value:
            break

    return value
```
Coding: TTL LRU (hard) Worked
```python
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.hits = 0
        self.misses = 0
        self.evictions = 0
        self.expirations = 0

    def put(self, key, value):
        if self.capacity == 0:
            return

        now = self.clock()
        expired = []

        # Remove expired entries
        for k, (v, expiry) in self.cache.items():
            if now >= expiry:
                expired.append(k)

        for k in expired:
            del self.cache[k]
            self.expirations += 1

        if key in self.cache:
            self.cache.move_to_end(key)
            self.cache[key] = (value, now + self.ttl)
            return

        if len(self.cache) >= self.capacity:
            lru_key, _ = self.cache.popitem(last=False)
            self.evictions += 1

        self.cache[key] = (value, now + self.ttl)
        self.cache.move_to_end(key)

    def get(self, key, default=None):
        if self.capacity == 0:
            self.misses += 1
            return default

        now = self.clock()
        if key in self.cache:
            value, expiry = self.cache[key]
            if now >= expiry:
                del self.cache[key]
                self.expirations += 1
                self.misses += 1
                return default
            else:
                self.hits += 1
                self.cache.move_to_end(key)
                self.cache[key] = (value, expiry)
                return value
        else:
            self.misses += 1
            return default

    def __len__(self):
        now = self.clock()
        count = 0
        for _, (value, expiry) in self.cache.items():
            if now < expiry:
                count += 1
        return count

    def stats(self):
        return {
            "hits": self.hits,
            "misses": self.misses,
            "evictions": self.evictions,
            "expirations": self.expirations,
        }
```
Coding: Parse duration Didn't work
```python
import re

def parse_duration(s: str) -> int:
    """Converts a duration string into a total number of seconds.

    Supported units: `h` (hours), `m` (minutes), `s` (seconds). Units are case-insensitive.
    Each part is a non-negative integer followed by a unit, e.g. "1h30m", "45s", "2H", "1h 5m 10s", "90m".
    Whitespace is allowed before, after and between parts.
    Units must appear in the order h, m, s and each at most once.
    Raise ValueError for empty/blank strings, numbers without a unit, unknown units, decimals, repeated units, or units out of order.
    """

    s = s.strip()
    if not s:
        raise ValueError("Empty duration string")

    parts = re.findall(r"(\d+)[hms]", s, re.IGNORECASE)
    units = re.findall(r"[hms]", s, re.IGNORECASE)

    if not parts:
        raise ValueError("Invalid duration format")

    if len(parts) != len(units):
        raise ValueError("Invalid duration format: numbers without unit")

    hours = 0
    minutes = 0
    seconds = 0

    hour_found = False
    minute_found = False
    second_found = False

    for i in range(len(parts)):
        try:
            value = int(parts[i])
        except ValueError:
            raise ValueError("Invalid duration format: non-integer value")

        unit = units[i].lower()

        if unit == 'h':
            if hour_found:
                raise ValueError("Repeated unit: hour")
            hours = value
            hour_found = True
        elif unit == 'm':
            if hour_found and not minute_found:
                pass
            elif not hour_found and minute_found:
                raise ValueError("Units out of order")
            if minute_found:
                raise ValueError("Repeated unit: minute")
            minutes = value
            minute_found = True
        elif unit == 's':
            if hour_found and minute_found and not second_found:
                pass
            elif hour_found and not minute_found and second_found:
                raise ValueError("Units out of order")
            elif not hour_found and minute_found and second_found:
                raise ValueError("Units out of order")
            if second_found:
                raise ValueError("Repeated unit: second")
            seconds = value
            second_found = True
        else:
            raise ValueError("Unknown unit")

    total_seconds = hours * 3600 + minutes * 60 + seconds

    return total_seconds
```
Coding: Summarize ranges Didn't work

```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""

    nums = sorted(list(set(nums)))
    result = []
    start = nums[0]
    end = nums[0]

    for i in range(1, len(nums)):
        if nums[i] == end + 1:
            end = nums[i]
        else:
            if start == end:
                result.append(str(start))
            else:
                result.append(f"{start}..{end}")
            start = nums[i]
            end = nums[i]

    if start == end:
        result.append(str(start))
    else:
        result.append(f"{start}..{end}")

    return ",".join(result)
```
Coding: Top customers Didn't work

```python
import re

def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    """
    Finds the top n customers based on their total spending, considering refunds.

    Args:
        rows: A list of dictionaries, each representing a transaction with "customer" and "amount" keys.
        n: The number of top customers to return.

    Returns:
        A list of (name, total) tuples, sorted by total descending, then name ascending.
    """
    customer_totals = {}
    for row in rows:
        customer = row.get("customer")
        amount_str = row.get("amount")

        if not customer or not amount_str:
            continue

        customer = customer.strip()

        try:
            amount = float(re.sub(r'[^\d\.-]', '', amount_str))
        except ValueError:
            continue

        if customer not in customer_totals:
            customer_totals[customer] = 0.0
        customer_totals[customer] += amount

    result = [(name, round(total, 2)) for name, total in customer_totals.items()]
    result.sort(key=lambda x: (-x[1], x[0]))
    return result[:n]
```
Coding: Slugify Didn't work
.
```python
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    """Builds a URL slug."""
    title = ''.join(c for c in unicodedata.normalize('NFKD', title)
                    if unicodedata.category(c) != 'Mn')
    title = ''.join(c for c in title if 0 <= ord(c) <= 0x10ffff)
    title = title.lower()
    title = ''.join(c if 'a' <= c <= 'z' or '0' <= c <= '9' else '-' for c in title)
    title = title.strip('-')
    if len(title) > max_len:
        words = title.split('-')
        current_len = 0
        result = []
        for word in words:
            if current_len + len(word) + 1 <= max_len:
                result.append(word)
                current_len += len(word) + 1
            else:
                if len(word) > max_len:
                    result.append(word[:max_len])
                break
        title = '-'.join(result).strip('-')
    return title
```
Coding: Token bucket Worked

```python
class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        """
        Initializes a TokenBucket rate limiter.

        Args:
            capacity: The maximum number of tokens the bucket can hold.
            refill_per_sec: The rate at which tokens are added per second.
            clock: A callable that returns the current time in seconds.
        """
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._tokens = capacity
        self._last_refill_time = 0.0

    def allow(self, cost: float = 1) -> bool:
        """
        Refills the bucket and attempts to subtract tokens.

        Args:
            cost: The number of tokens to subtract.

        Returns:
            True if the tokens were subtracted successfully, False otherwise.
        """
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        else:
            return False

    @property
    def tokens(self) -> float:
        """
        Returns the current number of tokens in the bucket.
        """
        self._refill()
        return self._tokens

    def _refill(self):
        """
        Refills the bucket with tokens based on elapsed time.
        """
        now = self._clock()
        time_elapsed = now - self._last_refill_time
        refill_amount = time_elapsed * self._refill_per_sec
        self._tokens = min(self._capacity, self._tokens + refill_amount)
        self._last_refill_time = now
```
Decisions: Refund window (hard) 50% right

```json
{
  "outcome": {
    "full_refund": 0.05,
    "partial_refund": 0.9,
    "store_credit": 0.03,
    "nothing": 0.02
  },
  "defective": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Incident matrix (hard) 100% right

```json
{
  "severity": {
    "0": 0.05,
    "1": 0.75,
    "2": 0.15,
    "3": 0.05
  },
  "page": {
    "true": 0.3,
    "false": 0.7
  }
}
```
Decisions: Tool followup (hard) 100% right

```json
{
  "tool": {
    "calendar": 0.95,
    "web_search": 0.01,
    "calculator": 0.01,
    "email": 0.02,
    "none": 0.01
  },
  "confirm": {
    "true": 0.85,
    "false": 0.15
  }
}
```
Decisions: Legit security alert (hard) 100% right
```json
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Meeting slot (hard) 50% right
```json
{
  "slot": {
    "A": 0.05,
    "B": 0.9,
    "C": 0.03,
    "D": 0.02
  },
  "raj_last": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Review mixed (hard) 100% right
```json
{
  "hardware": {
    "true": 0.95,
    "false": 0.05
  },
  "support": {
    "true": 0.85,
    "false": 0.15
  }
}
```
Decisions: Support checkout down 100% right

```json
{
  "department": {
    "billing": 0.1,
    "technical": 0.85,
    "account": 0.02,
    "shipping": 0.01,
    "sales": 0.02
  },
  "urgency": {
    "0": 0.01,
    "1": 0.05,
    "2": 0.2,
    "3": 0.74
  },
  "outage": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Refund wrong plan 100% right
```json
{
  "department": {
    "billing": 0.95,
    "technical": 0.01,
    "account": 0.02,
    "shipping": 0.01,
    "sales": 0.01
  },
  "refund": {
    "true": 0.98,
    "false": 0.02
  },
  "tone": {
    "frustrated": 0.2,
    "calm": 0.8
  }
}
```
Decisions: Moderation doxxing 100% right

```json
{
  "policy": {
    "none": 0.05,
    "harassment": 0.9,
    "hate": 0.01,
    "spam": 0.01,
    "self_harm": 0.03
  },
  "personal_info": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Route calendar 100% right
```json
{
  "tool": {
    "calendar": 0.95,
    "email": 0.03,
    "web_search": 0.01,
    "calculator": 0.005,
    "none": 0.005
  },
  "confirm": {
    "true": 0.99,
    "false": 0.01
  }
}
```
Decisions: Doc invoice missing due 100% right
```json
{
  "doc_type": {
    "invoice": 0.99,
    "resume": 0.001,
    "contract": 0.005,
    "bank_statement": 0.002,
    "other": 0.002
  },
  "missing_due_date": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Phishing paypal 100% right

```json
{
  "phishing": {
    "true": 0.95,
    "false": 0.05
  },
  "risk": {
    "0": 0.01,
    "1": 0.04,
    "2": 0.15,
    "3": 0.80
  }
}
```
Decisions: Pii ssn email 100% right

```json
{
  "data_kind": {
    "none": 0.05,
    "contact": 0.2,
    "financial": 0.05,
    "government_id": 0.65,
    "health": 0.05
  },
  "sensitive": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Review mixed 100% right
```json
{
  "sentiment": {
    "positive": 0.05,
    "neutral": 0.1,
    "negative": 0.85
  },
  "defect": {
    "true": 0.95,
    "false": 0.05
  },
  "recommend": {
    "true": 0.05,
    "false": 0.95
  }
}
```
Documents: Saas escalator (hard) 10% right

```json
{
  "year2_price_per_seat_month": 47.84,
  "year3_price_per_seat_month": 47.40,
  "year1_invoice": 5400.00,
  "year2_invoice": 5731.20,
  "addon_months_billed": 27,
  "addon_invoice": 7200.00,
  "year3_invoice": 5668.80,
  "year3_discount_percent": 0.15,
  "total_contract_value": 24000.00,
  "contract_end_date": "2027-02-28"
}
```
Documents: Expense thread 75% right

```json
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    { "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
    { "date": "2025-02-24", "category": "ground_transport", "amount_usd": 41.04 },
    { "date": "2025-02-25", "category": "meals", "amount_usd": 229.39 },
    { "date": "2025-02-26", "category": "lodging", "amount_usd": 468.24 }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 2118.27,
  "approver_email": "priya.raman@corvane.com"
}
```
Documents: Lease amendment 67% right

```json
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2238.00,
  "late_fee_from_2025_06_01": 111.90,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2238.00,
  "move_in_payment": 2550.00
}
```
Documents: Ticket SLA 82% right

```json
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-12T17:00",
  "sla_due_utc": "2025-09-12T21:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
```
Documents: Sales footnotes 89% right

```json
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 5.9,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}
```

Size: 27B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.