Review · updated OCT 11

Mistral Nemo review: not one we'd recommend right now

It scored 23 out of 100, #48 of 56. It solved 3 of 30 coding jobs and scored 36 on reading documents. Runs on a 12 GB graphics card or a Mac with 16 GB.

The short version
  • Mistral Nemo is a free model from Mistral AI that you can run on your own computer. In our tests it's not one we'd recommend right now: 23 out of 100, #48 of 56.
  • It solved 3 of 30 coding jobs and scored 36 on reading documents. On our hardest tasks it scored 13.
  • Runs on a 12 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Mistral Nemo got 3 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Mistral Nemo scored 36; the best model scored 100.

TestScorePublic questionsSecret questions
Coding10149
Reading documents365132
Decisions788077

On the 18 hardest tasks (included in the scores above) it scored 13. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Mistral Nemo in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBYesQ5
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ4
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run mistral-nemo

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Mistral Nemo wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict, Optional

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    sections: Dict[str, Dict[str, str]] = {}
    current_section: Optional[str] = None
    default_section: Dict[str, str] = {}
    lines = text.splitlines()
    i = 0

    while i < len(lines):
        line = lines[i].strip()

        if not line or line.startswith((';', '#')):
            i += 1
            continue

        match = re.match(r'\s*\[([^]]+)\]', line)
        if match:
            current_section = match.group(1).lower()
            if current_section == 'default':
                default_section = sections.pop('default', {})
            else:
                sections[current_section] = default_section.copy()
            i += 1
            continue

        match = re.match(r'\s*([^=:#;]+)\s*(=|:)\s*(.*)', line)
        if not match:
            raise ValueError(f"Malformed line: {line!r}")

        key, sep, value = match.groups()
        key = key.lower().strip()

        if current_section is None:
            raise ValueError(f"No section for key: {key!r}")

        if key in sections[current_section]:
            i += 1
            continue

        if value.startswith('"') and value.endswith('"'):
            value = value[1:-1]
        else:
            value = value.strip()

        sections[current_section][key] = value

        while i + 1 < len(lines) and lines[i + 1].strip() and lines[i + 1].strip()[0].isspace():
            value += '\n' + lines[i + 1].strip()
            i += 1

        i += 1

    return sections
```
Coding: TTL LRU (hard) Didn't work
```python
import time
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def put(self, key, value):
        now = self.clock()
        if key in self.cache:
            self.cache.move_to_end(key)
            self.cache[key] = (now + self.ttl, value)
        else:
            if len(self.cache) >= self.capacity:
                self._evict(now)
            self.cache[key] = (now + self.ttl, value)

    def get(self, key, default=None):
        now = self.clock()
        if key in self.cache and self.cache[key][0] > now:
            self.cache.move_to_end(key)
            return self.cache[key][1]
        else:
            self.stats["expirations"] += 1
            self.stats["misses"] += 1
            del self.cache[key]
            return default

    def __len__(self):
        now = self.clock()
        return sum(1 for k, v in self.cache.items() if v[0] > now)

    def stats(self):
        return self.stats

    def _evict(self, now):
        while self.cache:
            key = next(iter(self.cache))
            if self.cache[key][0] > now:
                break
            self.stats["evictions"] += 1
            del self.cache[key]
```
Coding: Parse duration Didn't work
```python
import re

def parse_duration(s: str) -> int:
    if not s.strip():
        raise ValueError("Empty or blank string")

    units = {'h': 3600, 'm': 60, 's': 1}
    pattern = r'^(\d+[hms]?)*$'
    if not re.match(pattern, s, re.IGNORECASE):
        raise ValueError("Invalid duration format")

    parts = re.split(r'(\d+)', s)
    parts = [p for p in parts if p]

    for i in range(0, len(parts), 2):
        value = int(parts[i])
        unit = parts[i + 1].lower()
        if unit not in units:
            raise ValueError(f"Unknown unit: {unit}")
        if i > 0 and units[unit] < units[prev_unit]:
            raise ValueError("Units out of order")
        total_seconds += value * units[unit]
        prev_unit = unit

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    nums = sorted(set(nums))
    result = []
    i = 0
    while i < len(nums):
        j = i + 1
        while j < len(nums) and nums[j] == nums[j - 1] + 1:
            j += 1
        if j - i >= 3:
            result.append(f"{nums[i]}..{nums[j - 1]}")
        elif j - i == 2:
            result.append(f"{nums[i]},{nums[j - 1]}")
        else:
            result.append(str(nums[i]))
        i = j
    return ",".join(result)
```
Coding: Top customers Didn't work
```python
from decimal import Decimal, InvalidOperation
from collections import defaultdict

def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    customers = defaultdict(Decimal)
    for row in rows:
        name = row.get("customer", "").strip()
        if not name:
            continue
        try:
            amount = Decimal(row["amount"].replace(",", ""))
        except InvalidOperation:
            continue
        customers[name.lower()] += amount
    return sorted(customers.items(), key=lambda x: (-x[1], x[0]))[:n]
```
Coding: Slugify Didn't work
```python
import re
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('utf-8')
    title = re.sub(r'[^a-z0-9]+', '-', title.lower())
    title = re.sub(r'-+', '-', title).strip('-')
    words = title.split('-')
    if len(title) > max_len:
        title = '-'.join(words[:max_len // len(words[0])]) if max_len % len(words[0]) == 0 else '-'.join(words[:max_len // len(words[0]) + 1])
        title = title[:max_len]
    return title.rstrip('-')
```
Coding: Token bucket Didn't work
```python
import time

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: callable):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.tokens = capacity
        self.last_refill = self.clock()

    def allow(self, cost: float = 1) -> bool:
        self.refill()
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self.refill()
        return self._tokens

    def refill(self) -> None:
        elapsed = self.clock() - self.last_refill
        self._tokens = min(self.capacity, self._tokens + elapsed * self.refill_per_sec)
        self.last_refill = self.clock()
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.8,
    "store_credit": 0.2,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 0% right
{
  "severity": {
    "0": 0.0,
    "1": 0.4,
    "2": 0.5,
    "3": 0.1
  },
  "page": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.0,
    "false": 1.0
  },
  "action_needed": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Meeting slot (hard) 50% right
{
  "slot": {
    "A": 0.0,
    "B": 0.0,
    "C": 1.0,
    "D": 0.0
  },
  "raj_last": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 1.0
  },
  "support": {
    "true": 1.0
  }
}
Decisions: Support checkout down 67% right
{
  "department": {
    "billing": 0.8,
    "technical": 0.2,
    "account": 0,
    "shipping": 0,
    "sales": 0
  },
  "urgency": {
    "0": 0,
    "1": 0,
    "2": 0.2,
    "3": 0.8
  },
  "outage": {
    "true": 1,
    "false": 0
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.9,
    "account": 0.1,
    "technical": 0,
    "shipping": 0,
    "sales": 0
  },
  "refund": {
    "true": 1,
    "false": 0
  },
  "tone": {
    "frustrated": 0.3,
    "calm": 0.7
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.2,
    "harassment": 0.8
  },
  "personal_info": {
    "true": 1,
    "false": 0
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Phishing paypal 50% right
{
  "phishing": {
    "true": 0.95,
    "false": 0.05
  },
  "risk": {
    "0": 0.05,
    "1": 0.2,
    "2": 0.65,
    "3": 0.1
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.0,
    "financial": 0.0,
    "government_id": 1.0,
    "health": 0.0
  },
  "sensitive": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0,
    "neutral": 0.2,
    "negative": 0.8
  },
  "defect": {
    "true": 1,
    "false": 0
  },
  "recommend": {
    "true": 0,
    "false": 1
  }
}
Documents: Saas escalator (hard) 20% right
{
  "year2_price_per_seat_month": 47.29,
  "year3_price_per_seat_month": 45.00,
  "year1_invoice": 6600.00,
  "year2_invoice": 7560.00,
  "addon_months_billed": 10,
  "addon_invoice": 7680.00,
  "year3_invoice": 8100.00,
  "year3_discount_percent": 15.00,
  "total_contract_value": 22340.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 75% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.6
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 36.0
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 230.08
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.24
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.28
    }
  ],
  "rejected_item_count": 2,
  "per_diem_days": 3,
  "per_diem_usd": 195.0,
  "total_reimbursable_usd": 2135.12,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 67% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2242,
  "late_fee_from_2025_06_01": 112.1,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2327,
  "move_in_payment": 3350
}
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-12T17:30",
  "sla_due_utc": "2025-09-12T22:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 10% right
{
  "q3_total_usd": 15356,
  "q2_total_usd": 15464,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 5.9,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": [],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 4291
}

Size: 12B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.