- Mistral Small 3 is a free model from Mistral AI that you can run on your own computer. In our tests it's not one we'd recommend right now: 41 out of 100, #37 of 56.
- It solved 11 of 30 coding jobs and scored 45 on reading documents. On our hardest tasks it scored 24.
- Runs on a 24 GB graphics card or a Mac with 32 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Mistral Small 3 got 11 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Mistral Small 3 scored 45; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 37 | 43 | 35 |
| Reading documents | 45 | 58 | 42 |
| Decisions | 85 | 86 | 85 |
On the 18 hardest tasks (included in the scores above) it scored 24. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Mistral Small 3 in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | Yes | Q6 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | Yes | Q5 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
Search for "Mistral Small 3 GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.
Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Mistral Small 3 wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
from collections import defaultdict
def parse_ini(text: str) -> dict[str, dict[str, str]]:
def strip_whitespace(s: str) -> str:
return s.strip()
def parse_value(value: str) -> str:
if value.startswith('"') and value.endswith('"'):
return value[1:-1]
return value
def parse_line(line: str, section: str, section_dict: dict[str, str], default_dict: dict[str, str]) -> None:
line = line.strip()
if not line or line.startswith((';', '#')):
return
if '=' in line:
key, value = line.split('=', 1)
elif ':' in line:
key, value = line.split(':', 1)
else:
raise ValueError(f"Malformed line: {line}")
key = key.strip().lower()
value = parse_value(value.strip())
if key in section_dict:
section_dict[key] += f"\n{value}"
else:
section_dict[key] = value
def interpolate_value(value: str, section: str, sections: dict[str, dict[str, str]]) -> str:
def interpolate(match: re.Match) -> str:
key = match.group(1)
if ':' in key:
sec, key = key.split(':', 1)
if sec not in sections or key not in sections[sec]:
raise KeyError(f"Missing key: {sec}:{key}")
return sections[sec][key]
if key not in sections[section]:
raise KeyError(f"Missing key: {key}")
return sections[section][key]
def recursive_interpolate(value: str) -> str:
try:
return re.sub(r'\$\{([^}]+)\}', interpolate, value)
except KeyError as e:
raise ValueError(f"Interpolation error: {e}")
except ValueError as e:
raise ValueError(f"Interpolation cycle detected: {e}")
return recursive_interpolate(value)
lines = text.splitlines()
sections = {}
current_section = None
default_section = {}
default_set = set()
for line in lines:
stripped_line = line.strip()
if stripped_line.startswith('[') and stripped_line.endswith(']'):
if current_section and not default_set:
raise ValueError(f"Key line before any section header: {line}")
section_name = stripped_line[1:-1].strip()
if section_name == 'DEFAULT':
current_section = 'DEFAULT'
else:
current_section = section_name
sections[current_section] = defaultdict(str)
continue
if not current_section:
raise ValueError(f"Key line before any section header: {line}")
if current_section == 'DEFAULT':
parse_line(line, current_section, default_section, default_section)
default_set.add(line.strip().lower())
else:
parse_line(line, current_section, sections[current_section], default_section)
for section in sections:
for key, value in sections[section].items():
try:
sections[section][key] = interpolate_value(value, section, sections)
except KeyError as e:
raise ValueError(f"Interpolation error: {e}")
except ValueError as e:
raise ValueError(f"Interpolation cycle detected: {e}")
return {k: dict(v) for k, v in sections.items() if k != 'DEFAULT'}
``` Coding: TTL LRU (hard) Worked
```python
from collections import OrderedDict
from typing import Callable, Any, Optional
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.hits = 0
self.misses = 0
self.evictions = 0
self.expirations = 0
def _cleanup(self):
now = self.clock()
keys_to_remove = []
for key, (value, expiry) in self.cache.items():
if now >= expiry:
keys_to_remove.append(key)
for key in keys_to_remove:
del self.cache[key]
self.expirations += 1
def put(self, key: Any, value: Any):
if self.capacity == 0:
return
now = self.clock()
self._cleanup()
if key in self.cache:
del self.cache[key]
elif len(self.cache) >= self.capacity:
self.cache.popitem(last=False)
self.evictions += 1
self.cache[key] = (value, now + self.ttl)
def get(self, key: Any, default: Optional[Any] = None) -> Any:
if self.capacity == 0:
self.misses += 1
return default
now = self.clock()
self._cleanup()
if key in self.cache:
value, expiry = self.cache.pop(key)
if now < expiry:
self.cache[key] = (value, expiry)
self.hits += 1
return value
else:
self.expirations += 1
self.misses += 1
return default
else:
self.misses += 1
return default
def __len__(self) -> int:
now = self.clock()
return sum(1 for value, expiry in self.cache.values() if now < expiry)
def stats(self) -> dict:
return {
"hits": self.hits,
"misses": self.misses,
"evictions": self.evictions,
"expirations": self.expirations
}
``` Coding: Parse duration Didn't work
```python
import re
def parse_duration(s: str) -> int:
if not s or s.isspace():
raise ValueError("Empty or blank input string")
# Regular expression to match the duration parts
pattern = re.compile(r'(\d+)(h|m|s)', re.IGNORECASE)
parts = pattern.findall(s)
if not parts:
raise ValueError("No valid duration parts found")
total_seconds = 0
last_unit = None
for value, unit in parts:
value = int(value)
unit = unit.lower()
if unit not in {'h', 'm', 's'}:
raise ValueError(f"Unknown unit: {unit}")
if unit == 'h':
total_seconds += value * 3600
elif unit == 'm':
total_seconds += value * 60
elif unit == 's':
total_seconds += value
if last_unit and unit <= last_unit:
raise ValueError("Units must appear in the order h, m, s and each at most once")
last_unit = unit
return total_seconds
``` Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
nums = sorted(set(nums))
ranges = []
start = nums[0]
for i in range(1, len(nums)):
if nums[i] == nums[i - 1] + 1:
continue
if nums[i - 1] - start >= 2:
ranges.append(f"{start}..{nums[i - 1]}")
elif nums[i - 1] - start == 1:
ranges.append(f"{start},{nums[i - 1]}")
else:
ranges.append(f"{nums[i - 1]}")
start = nums[i]
if nums[-1] - start >= 2:
ranges.append(f"{start}..{nums[-1]}")
elif nums[-1] - start == 1:
ranges.append(f"{start},{nums[-1]}")
else:
ranges.append(f"{nums[-1]}")
return ",".join(ranges)
``` Coding: Top customers Didn't work
```python
import re
from collections import defaultdict
def parse_amount(amount: str) -> float:
# Remove any commas and dollar signs, then convert to float
amount = re.sub(r'[$,]', '', amount)
return round(float(amount), 2)
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
customer_totals = defaultdict(float)
for row in rows:
customer = row.get("customer", "").strip()
amount_str = row.get("amount", "")
if not customer or not amount_str:
continue
try:
amount = parse_amount(amount_str)
except ValueError:
continue
customer = customer.lower()
customer_totals[customer] += amount
# Sort by total descending, then by name ascending
sorted_customers = sorted(customer_totals.items(), key=lambda x: (-x[1], x[0]))
# Get the top n customers
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import re
import unicodedata
def slugify(title: str, max_len: int = 50) -> str:
# Transliterate accented characters to ASCII and drop any other non-ASCII characters
title = unicodedata.normalize('NFKD', title)
title = ''.join(c for c in title if unicodedata.category(c).startswith('L') or c.isdigit() or c.isspace())
# Lowercase
title = title.lower()
# Replace every run of characters other than a-z and 0-9 with a single "-"
title = re.sub(r'[^a-z0-9]+', '-', title)
# Strip leading/trailing "-"
title = title.strip('-')
# Shorten the slug if it is longer than max_len
if len(title) > max_len:
words = title.split('-')
shortened_slug = []
current_length = 0
for word in words:
if current_length + len(word) + 1 <= max_len:
shortened_slug.append(word)
current_length += len(word) + 1
else:
break
if current_length > max_len:
shortened_slug = [shortened_slug[0][:max_len]]
title = '-'.join(shortened_slug)
return title
``` Coding: Token bucket Worked
```python
import time
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.clock = clock
self.current_tokens = capacity
self.last_refill_time = clock()
def _refill(self):
now = self.clock()
elapsed = now - self.last_refill_time
self.current_tokens = min(self.capacity, self.current_tokens + elapsed * self.refill_per_sec)
self.last_refill_time = now
def allow(self, cost: float = 1) -> bool:
self._refill()
if self.current_tokens >= cost:
self.current_tokens -= cost
return True
return False
@property
def tokens(self) -> float:
self._refill()
return self.current_tokens
``` Decisions: Refund window (hard) 50% right
```json
{
"outcome": {
"full_refund": 0.0,
"partial_refund": 0.9,
"store_credit": 0.1,
"nothing": 0.0
},
"defective": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Incident matrix (hard) 0% right
```json
{
"severity": {
"0": 0,
"1": 0.3,
"2": 0.7,
"3": 0
},
"page": {
"true": 0.7,
"false": 0.3
}
}
``` Decisions: Tool followup (hard) 100% right
```json
{
"tool": {
"calendar": 1.0
},
"confirm": {
"true": 1.0
}
}
``` Decisions: Legit security alert (hard) 100% right
```json
{
"phishing": {
"true": 0.05,
"false": 0.95
},
"action_needed": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Meeting slot (hard) 50% right
```json
{
"slot": {
"A": 0,
"B": 0.9,
"C": 0.1,
"D": 0
},
"raj_last": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Review mixed (hard) 100% right
```json
{
"hardware": {
"true": 1.0,
"false": 0.0
},
"support": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Support checkout down 100% right
```json
{
"department": {
"billing": 0.1,
"technical": 0.9,
"account": 0.0,
"shipping": 0.0,
"sales": 0.0
},
"urgency": {
"0": 0.0,
"1": 0.0,
"2": 0.1,
"3": 0.9
},
"outage": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Refund wrong plan 100% right
```json
{
"department": {
"billing": 0.9,
"technical": 0.05,
"account": 0.02,
"shipping": 0.01,
"sales": 0.02
},
"refund": {
"true": 0.95,
"false": 0.05
},
"tone": {
"frustrated": 0.2,
"calm": 0.8
}
}
``` Decisions: Moderation doxxing 100% right
```json
{
"policy": {
"none": 0.05,
"harassment": 0.85,
"hate": 0.05,
"spam": 0.05,
"self_harm": 0
},
"personal_info": {
"true": 0.9,
"false": 0.1
}
}
``` Decisions: Route calendar 100% right
```json
{
"tool": {
"calendar": 0.9,
"email": 0.1,
"web_search": 0,
"calculator": 0,
"none": 0
},
"confirm": {
"true": 0.9,
"false": 0.1
}
}
``` Decisions: Doc invoice missing due 100% right
```json
{
"doc_type": {
"invoice": 0.95,
"resume": 0.01,
"contract": 0.01,
"bank_statement": 0.01,
"other": 0.02
},
"missing_due_date": {
"true": 0.90,
"false": 0.10
}
}
``` Decisions: Phishing paypal 100% right
```json
{
"phishing": {
"true": 1.0,
"false": 0.0
},
"risk": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
}
}
``` Decisions: Pii ssn email 100% right
```json
{
"data_kind": {
"none": 0,
"contact": 0.1,
"financial": 0.1,
"government_id": 0.8,
"health": 0
},
"sensitive": {
"true": 0.9,
"false": 0.1
}
}
``` Decisions: Review mixed 100% right
```json
{
"sentiment": {
"positive": 0.1,
"neutral": 0.1,
"negative": 0.8
},
"defect": {
"true": 0.9,
"false": 0.1
},
"recommend": {
"true": 0.1,
"false": 0.9
}
}
``` Documents: Saas escalator (hard) 30% right
```json
{
"year2_price_per_seat_month": 48.33,
"year3_price_per_seat_month": 48.33,
"year1_invoice": 64800,
"year2_invoice": 69840,
"addon_months_billed": 6,
"addon_invoice": 10884.8,
"year3_invoice": 119280,
"year3_discount_percent": 15,
"total_contract_value": 253904.8,
"contract_end_date": "2027-02-28"
}
``` Documents: Expense thread 54% right
```json
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{ "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
{ "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
{ "date": "2025-02-25", "category": "meals", "amount_usd": 228.95 },
{ "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.76 },
{ "date": "2025-02-26", "category": "lodging", "amount_usd": 466.56 }
],
"rejected_item_count": 2,
"per_diem_days": 2,
"per_diem_usd": 130,
"total_reimbursable_usd": 2099.25,
"approver_email": "priya.raman@corvane.com"
}
``` Documents: Lease amendment 100% right
```json
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150,
"monthly_rent_from_2025_06_01": 2236,
"late_fee_from_2025_06_01": 111.8,
"security_deposit": 2150,
"total_pet_deposits": 800,
"total_monthly_payment_july_2025": 2306,
"move_in_payment": 4700
}
``` Documents: Ticket SLA 82% right
```json
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-13T16:00",
"sla_due_utc": "2025-09-13T21:00:00Z",
"reissued_invoice": "INV-2025-0812"
}
``` Documents: Sales footnotes 22% right
```json
{
"q3_total_usd": 15346,
"q2_total_usd": 14464,
"q2_central_originally_reported_usd": 3047,
"q2_to_q3_change_pct": 6.2,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "Central",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731,
"west_excluding_mountain_q3_usd": 4201
}
``` Size: 24B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.