Review · updated OCT 11

Mistral Small 3 review: not one we'd recommend right now

It scored 41 out of 100, #37 of 56. It solved 11 of 30 coding jobs and scored 45 on reading documents. Runs on a 24 GB graphics card or a Mac with 32 GB.

The short version
  • Mistral Small 3 is a free model from Mistral AI that you can run on your own computer. In our tests it's not one we'd recommend right now: 41 out of 100, #37 of 56.
  • It solved 11 of 30 coding jobs and scored 45 on reading documents. On our hardest tasks it scored 24.
  • Runs on a 24 GB graphics card or a Mac with 32 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Mistral Small 3 got 11 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Mistral Small 3 scored 45; the best model scored 100.

TestScorePublic questionsSecret questions
Coding374335
Reading documents455842
Decisions858685

On the 18 hardest tasks (included in the scores above) it scored 24. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Mistral Small 3 in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBYesQ6
Graphics card, 32 GBYesQ8
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBYesQ5
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Mistral Small 3 GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Mistral Small 3 wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from collections import defaultdict

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    def strip_whitespace(s: str) -> str:
        return s.strip()

    def parse_value(value: str) -> str:
        if value.startswith('"') and value.endswith('"'):
            return value[1:-1]
        return value

    def parse_line(line: str, section: str, section_dict: dict[str, str], default_dict: dict[str, str]) -> None:
        line = line.strip()
        if not line or line.startswith((';', '#')):
            return
        if '=' in line:
            key, value = line.split('=', 1)
        elif ':' in line:
            key, value = line.split(':', 1)
        else:
            raise ValueError(f"Malformed line: {line}")
        key = key.strip().lower()
        value = parse_value(value.strip())
        if key in section_dict:
            section_dict[key] += f"\n{value}"
        else:
            section_dict[key] = value

    def interpolate_value(value: str, section: str, sections: dict[str, dict[str, str]]) -> str:
        def interpolate(match: re.Match) -> str:
            key = match.group(1)
            if ':' in key:
                sec, key = key.split(':', 1)
                if sec not in sections or key not in sections[sec]:
                    raise KeyError(f"Missing key: {sec}:{key}")
                return sections[sec][key]
            if key not in sections[section]:
                raise KeyError(f"Missing key: {key}")
            return sections[section][key]

        def recursive_interpolate(value: str) -> str:
            try:
                return re.sub(r'\$\{([^}]+)\}', interpolate, value)
            except KeyError as e:
                raise ValueError(f"Interpolation error: {e}")
            except ValueError as e:
                raise ValueError(f"Interpolation cycle detected: {e}")

        return recursive_interpolate(value)

    lines = text.splitlines()
    sections = {}
    current_section = None
    default_section = {}
    default_set = set()

    for line in lines:
        stripped_line = line.strip()
        if stripped_line.startswith('[') and stripped_line.endswith(']'):
            if current_section and not default_set:
                raise ValueError(f"Key line before any section header: {line}")
            section_name = stripped_line[1:-1].strip()
            if section_name == 'DEFAULT':
                current_section = 'DEFAULT'
            else:
                current_section = section_name
                sections[current_section] = defaultdict(str)
            continue

        if not current_section:
            raise ValueError(f"Key line before any section header: {line}")

        if current_section == 'DEFAULT':
            parse_line(line, current_section, default_section, default_section)
            default_set.add(line.strip().lower())
        else:
            parse_line(line, current_section, sections[current_section], default_section)

    for section in sections:
        for key, value in sections[section].items():
            try:
                sections[section][key] = interpolate_value(value, section, sections)
            except KeyError as e:
                raise ValueError(f"Interpolation error: {e}")
            except ValueError as e:
                raise ValueError(f"Interpolation cycle detected: {e}")

    return {k: dict(v) for k, v in sections.items() if k != 'DEFAULT'}
```
Coding: TTL LRU (hard) Worked
```python
from collections import OrderedDict
from typing import Callable, Any, Optional

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.hits = 0
        self.misses = 0
        self.evictions = 0
        self.expirations = 0

    def _cleanup(self):
        now = self.clock()
        keys_to_remove = []
        for key, (value, expiry) in self.cache.items():
            if now >= expiry:
                keys_to_remove.append(key)
        for key in keys_to_remove:
            del self.cache[key]
            self.expirations += 1

    def put(self, key: Any, value: Any):
        if self.capacity == 0:
            return
        now = self.clock()
        self._cleanup()
        if key in self.cache:
            del self.cache[key]
        elif len(self.cache) >= self.capacity:
            self.cache.popitem(last=False)
            self.evictions += 1
        self.cache[key] = (value, now + self.ttl)

    def get(self, key: Any, default: Optional[Any] = None) -> Any:
        if self.capacity == 0:
            self.misses += 1
            return default
        now = self.clock()
        self._cleanup()
        if key in self.cache:
            value, expiry = self.cache.pop(key)
            if now < expiry:
                self.cache[key] = (value, expiry)
                self.hits += 1
                return value
            else:
                self.expirations += 1
                self.misses += 1
                return default
        else:
            self.misses += 1
            return default

    def __len__(self) -> int:
        now = self.clock()
        return sum(1 for value, expiry in self.cache.values() if now < expiry)

    def stats(self) -> dict:
        return {
            "hits": self.hits,
            "misses": self.misses,
            "evictions": self.evictions,
            "expirations": self.expirations
        }
```
Coding: Parse duration Didn't work
```python
import re

def parse_duration(s: str) -> int:
    if not s or s.isspace():
        raise ValueError("Empty or blank input string")

    # Regular expression to match the duration parts
    pattern = re.compile(r'(\d+)(h|m|s)', re.IGNORECASE)
    parts = pattern.findall(s)

    if not parts:
        raise ValueError("No valid duration parts found")

    total_seconds = 0
    last_unit = None

    for value, unit in parts:
        value = int(value)
        unit = unit.lower()

        if unit not in {'h', 'm', 's'}:
            raise ValueError(f"Unknown unit: {unit}")

        if unit == 'h':
            total_seconds += value * 3600
        elif unit == 'm':
            total_seconds += value * 60
        elif unit == 's':
            total_seconds += value

        if last_unit and unit <= last_unit:
            raise ValueError("Units must appear in the order h, m, s and each at most once")

        last_unit = unit

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""

    nums = sorted(set(nums))
    ranges = []
    start = nums[0]

    for i in range(1, len(nums)):
        if nums[i] == nums[i - 1] + 1:
            continue
        if nums[i - 1] - start >= 2:
            ranges.append(f"{start}..{nums[i - 1]}")
        elif nums[i - 1] - start == 1:
            ranges.append(f"{start},{nums[i - 1]}")
        else:
            ranges.append(f"{nums[i - 1]}")
        start = nums[i]

    if nums[-1] - start >= 2:
        ranges.append(f"{start}..{nums[-1]}")
    elif nums[-1] - start == 1:
        ranges.append(f"{start},{nums[-1]}")
    else:
        ranges.append(f"{nums[-1]}")

    return ",".join(ranges)
```
Coding: Top customers Didn't work
```python
import re
from collections import defaultdict

def parse_amount(amount: str) -> float:
    # Remove any commas and dollar signs, then convert to float
    amount = re.sub(r'[$,]', '', amount)
    return round(float(amount), 2)

def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    customer_totals = defaultdict(float)

    for row in rows:
        customer = row.get("customer", "").strip()
        amount_str = row.get("amount", "")

        if not customer or not amount_str:
            continue

        try:
            amount = parse_amount(amount_str)
        except ValueError:
            continue

        customer = customer.lower()
        customer_totals[customer] += amount

    # Sort by total descending, then by name ascending
    sorted_customers = sorted(customer_totals.items(), key=lambda x: (-x[1], x[0]))

    # Get the top n customers
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import re
import unicodedata

def slugify(title: str, max_len: int = 50) -> str:
    # Transliterate accented characters to ASCII and drop any other non-ASCII characters
    title = unicodedata.normalize('NFKD', title)
    title = ''.join(c for c in title if unicodedata.category(c).startswith('L') or c.isdigit() or c.isspace())

    # Lowercase
    title = title.lower()

    # Replace every run of characters other than a-z and 0-9 with a single "-"
    title = re.sub(r'[^a-z0-9]+', '-', title)

    # Strip leading/trailing "-"
    title = title.strip('-')

    # Shorten the slug if it is longer than max_len
    if len(title) > max_len:
        words = title.split('-')
        shortened_slug = []
        current_length = 0

        for word in words:
            if current_length + len(word) + 1 <= max_len:
                shortened_slug.append(word)
                current_length += len(word) + 1
            else:
                break

        if current_length > max_len:
            shortened_slug = [shortened_slug[0][:max_len]]

        title = '-'.join(shortened_slug)

    return title
```
Coding: Token bucket Worked
```python
import time

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.current_tokens = capacity
        self.last_refill_time = clock()

    def _refill(self):
        now = self.clock()
        elapsed = now - self.last_refill_time
        self.current_tokens = min(self.capacity, self.current_tokens + elapsed * self.refill_per_sec)
        self.last_refill_time = now

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self.current_tokens >= cost:
            self.current_tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self.current_tokens
```
Decisions: Refund window (hard) 50% right
```json
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.9,
    "store_credit": 0.1,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Incident matrix (hard) 0% right
```json
{
  "severity": {
    "0": 0,
    "1": 0.3,
    "2": 0.7,
    "3": 0
  },
  "page": {
    "true": 0.7,
    "false": 0.3
  }
}
```
Decisions: Tool followup (hard) 100% right
```json
{
  "tool": {
    "calendar": 1.0
  },
  "confirm": {
    "true": 1.0
  }
}
```
Decisions: Legit security alert (hard) 100% right
```json
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Meeting slot (hard) 50% right
```json
{
  "slot": {
    "A": 0,
    "B": 0.9,
    "C": 0.1,
    "D": 0
  },
  "raj_last": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Review mixed (hard) 100% right
```json
{
  "hardware": {
    "true": 1.0,
    "false": 0.0
  },
  "support": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Support checkout down 100% right
```json
{
  "department": {
    "billing": 0.1,
    "technical": 0.9,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.1,
    "3": 0.9
  },
  "outage": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Refund wrong plan 100% right
```json
{
  "department": {
    "billing": 0.9,
    "technical": 0.05,
    "account": 0.02,
    "shipping": 0.01,
    "sales": 0.02
  },
  "refund": {
    "true": 0.95,
    "false": 0.05
  },
  "tone": {
    "frustrated": 0.2,
    "calm": 0.8
  }
}
```
Decisions: Moderation doxxing 100% right
```json
{
  "policy": {
    "none": 0.05,
    "harassment": 0.85,
    "hate": 0.05,
    "spam": 0.05,
    "self_harm": 0
  },
  "personal_info": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Route calendar 100% right
```json
{
  "tool": {
    "calendar": 0.9,
    "email": 0.1,
    "web_search": 0,
    "calculator": 0,
    "none": 0
  },
  "confirm": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Doc invoice missing due 100% right
```json
{
  "doc_type": {
    "invoice": 0.95,
    "resume": 0.01,
    "contract": 0.01,
    "bank_statement": 0.01,
    "other": 0.02
  },
  "missing_due_date": {
    "true": 0.90,
    "false": 0.10
  }
}
```
Decisions: Phishing paypal 100% right
```json
{
  "phishing": {
    "true": 1.0,
    "false": 0.0
  },
  "risk": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  }
}
```
Decisions: Pii ssn email 100% right
```json
{
  "data_kind": {
    "none": 0,
    "contact": 0.1,
    "financial": 0.1,
    "government_id": 0.8,
    "health": 0
  },
  "sensitive": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Review mixed 100% right
```json
{
  "sentiment": {
    "positive": 0.1,
    "neutral": 0.1,
    "negative": 0.8
  },
  "defect": {
    "true": 0.9,
    "false": 0.1
  },
  "recommend": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Documents: Saas escalator (hard) 30% right
```json
{
  "year2_price_per_seat_month": 48.33,
  "year3_price_per_seat_month": 48.33,
  "year1_invoice": 64800,
  "year2_invoice": 69840,
  "addon_months_billed": 6,
  "addon_invoice": 10884.8,
  "year3_invoice": 119280,
  "year3_discount_percent": 15,
  "total_contract_value": 253904.8,
  "contract_end_date": "2027-02-28"
}
```
Documents: Expense thread 54% right
```json
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    { "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
    { "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
    { "date": "2025-02-25", "category": "meals", "amount_usd": 228.95 },
    { "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.76 },
    { "date": "2025-02-26", "category": "lodging", "amount_usd": 466.56 }
  ],
  "rejected_item_count": 2,
  "per_diem_days": 2,
  "per_diem_usd": 130,
  "total_reimbursable_usd": 2099.25,
  "approver_email": "priya.raman@corvane.com"
}
```
Documents: Lease amendment 100% right
```json
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2236,
  "late_fee_from_2025_06_01": 111.8,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2306,
  "move_in_payment": 4700
}
```
Documents: Ticket SLA 82% right
```json
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-13T16:00",
  "sla_due_utc": "2025-09-13T21:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
```
Documents: Sales footnotes 22% right
```json
{
  "q3_total_usd": 15346,
  "q2_total_usd": 14464,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 6.2,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "Central",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 4201
}
```

Size: 24B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.