Review · updated OCT 11

Devstral 2 2512 review: not one we'd recommend right now

It scored 53 out of 100, #28 of 56. It solved 16 of 30 coding jobs and scored 53 on reading documents. Runs on a Mac with 128 GB.

The short version
  • Devstral 2 2512 is a free model from Mistral AI that you can run on your own computer. In our tests it's not one we'd recommend right now: 53 out of 100, #28 of 56.
  • It solved 16 of 30 coding jobs and scored 53 on reading documents. On our hardest tasks it scored 24.
  • Runs on a Mac with 128 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Devstral 2 2512 got 16 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Devstral 2 2512 scored 53; the best model scored 100.

TestScorePublic questionsSecret questions
Coding537148
Reading documents537149
Decisions928694

On the 18 hardest tasks (included in the scores above) it scored 24. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Devstral 2 2512 in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBNo–
Mac, 96 GBNo–
Mac, 128 GBYesQ5

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Devstral 2 2512 GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Devstral 2 2512 wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from collections import defaultdict

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    sections: dict[str, dict[str, str]] = defaultdict(dict)
    current_section = None
    pending_key = None
    pending_value_lines: list[str] = []
    default_section: dict[str, str] = {}

    def resolve_interpolation(value: str, section: str) -> str:
        def replace(match: re.Match[str]) -> str:
            key_expr = match.group(1)
            if ':' in key_expr:
                ref_section, ref_key = key_expr.split(':', 1)
                ref_section = ref_section.strip()
                ref_key = ref_key.strip().lower()
            else:
                ref_section = section
                ref_key = key_expr.strip().lower()

            if ref_section == 'DEFAULT':
                ref_dict = default_section
            elif ref_section in sections:
                ref_dict = sections[ref_section]
            else:
                raise KeyError(f"Section '{ref_section}' not found")

            if ref_key not in ref_dict:
                raise KeyError(f"Key '{ref_key}' not found in section '{ref_section}'")
            return resolve_interpolation(ref_dict[ref_key], ref_section)

        pattern = re.compile(r'\$\{([^}]*)\}')
        while True:
            new_value = pattern.sub(replace, value)
            if new_value == value:
                return new_value
            value = new_value
            if len(value) > 1000:  # Prevent infinite recursion
                raise ValueError("Circular interpolation detected")

    for line in text.splitlines():
        stripped = line.strip()
        if not stripped or stripped.startswith(';') or stripped.startswith('#'):
            continue

        if line.startswith(' ') or line.startswith('\t'):
            if pending_key is None:
                raise ValueError("Continuation line without preceding key")
            pending_value_lines.append(line.strip())
            continue

        section_match = re.match(r'^\s*\[([^\]]*)\]\s*$', line)
        if section_match:
            current_section = section_match.group(1).strip()
            if current_section == 'DEFAULT':
                current_section = None
            pending_key = None
            pending_value_lines = []
            continue

        if pending_key is not None:
            raise ValueError("Malformed line: expected continuation or section header")

        key_value_match = re.match(r'^\s*([^:=]+)([:=])(.*)', line)
        if not key_value_match:
            raise ValueError(f"Malformed line: {line!r}")

        key = key_value_match.group(1).strip().lower()
        separator = key_value_match.group(2)
        value_part = key_value_match.group(3).strip()

        inline_comment_match = re.match(r'^(.*?)\s*[;#](.*)', value_part)
        if inline_comment_match and not value_part.startswith('"'):
            value_part = inline_comment_match.group(1).strip()

        if value_part.startswith('"') and value_part.endswith('"'):
            value = value_part[1:-1]
        else:
            value = value_part

        pending_key = key
        pending_value_lines = [value]

    if pending_key is not None:
        raise ValueError("Malformed line: unterminated key-value pair")

    for section_name, section_dict in sections.items():
        for key, value in section_dict.items():
            sections[section_name][key] = resolve_interpolation(value, section_name)

    result: dict[str, dict[str, str]] = {}
    for section_name, section_dict in sections.items():
        if section_name is None:
            continue
        inherited = default_section.copy()
        inherited.update(section_dict)
        result[section_name] = inherited

    return result
```
Coding: TTL LRU (hard) Didn't work
```python
import collections

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self._cache = collections.OrderedDict()
        self._expiry = {}
        self._hits = 0
        self._misses = 0
        self._evictions = 0
        self._expirations = 0

    def _remove_expired(self):
        now = self.clock()
        expired_keys = [key for key, expiry in self._expiry.items() if expiry <= now]
        for key in expired_keys:
            del self._cache[key]
            del self._expiry[key]
            self._expirations += 1

    def put(self, key, value):
        if self.capacity == 0:
            return
        now = self.clock()
        if key in self._cache:
            self._cache.move_to_end(key)
            self._expiry[key] = now + self.ttl
        else:
            self._remove_expired()
            if len(self._cache) >= self.capacity:
                self._cache.popitem(last=False)
                self._evictions += 1
            self._cache[key] = value
            self._expiry[key] = now + self.ttl

    def get(self, key, default=None):
        now = self.clock()
        if key in self._cache:
            if self._expiry[key] > now:
                self._cache.move_to_end(key)
                self._hits += 1
                return self._cache[key]
            else:
                del self._cache[key]
                del self._expiry[key]
                self._expirations += 1
        self._misses += 1
        return default

    def __len__(self):
        self._remove_expired()
        return len(self._cache)

    def stats(self):
        return {
            "hits": self._hits,
            "misses": self._misses,
            "evictions": self._evictions,
            "expirations": self._expirations,
        }
```
Coding: Parse duration Worked
```python
import re

def parse_duration(s: str) -> int:
    if not s.strip():
        raise ValueError("Empty or blank string")

    pattern = r'^(\s*\d+\s*[hms]\s*)+$'
    if not re.fullmatch(pattern, s, re.IGNORECASE):
        raise ValueError("Invalid format")

    parts = re.findall(r'(\d+)\s*([hms])', s, re.IGNORECASE)
    if not parts:
        raise ValueError("No valid parts found")

    seen_units = set()
    total_seconds = 0
    last_unit_order = -1
    unit_order = {'h': 0, 'm': 1, 's': 2}

    for num_str, unit in parts:
        unit_lower = unit.lower()
        current_order = unit_order[unit_lower]

        if current_order <= last_unit_order:
            raise ValueError("Units out of order or repeated")

        if unit_lower in seen_units:
            raise ValueError("Repeated unit")

        seen_units.add(unit_lower)
        last_unit_order = current_order

        num = int(num_str)
        if unit_lower == 'h':
            total_seconds += num * 3600
        elif unit_lower == 'm':
            total_seconds += num * 60
        elif unit_lower == 's':
            total_seconds += num

    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""
    nums = sorted(set(nums))
    result = []
    start = nums[0]
    prev = start
    for num in nums[1:]:
        if num == prev + 1:
            prev = num
        else:
            if prev - start >= 2:
                result.append(f"{start}..{prev}")
            else:
                result.extend(str(x) for x in range(start, prev + 1))
            start = num
            prev = num
    if prev - start >= 2:
        result.append(f"{start}..{prev}")
    else:
        result.extend(str(x) for x in range(start, prev + 1))
    return ",".join(result)
```
Coding: Top customers Worked
```python
import re
from collections import defaultdict

def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
    customer_totals = defaultdict(float)
    customer_names = {}

    for row in rows:
        customer = row.get("customer", "").strip()
        if not customer:
            continue

        amount_str = row.get("amount", "").strip()
        if not amount_str:
            continue

        try:
            amount = float(re.sub(r'[^\d.-]', '', amount_str))
        except ValueError:
            continue

        lower_customer = customer.lower()
        if lower_customer not in customer_names:
            customer_names[lower_customer] = customer

        customer_totals[lower_customer] += amount

    result = []
    for lower_customer, total in customer_totals.items():
        rounded_total = round(total, 2)
        if rounded_total == int(rounded_total):
            rounded_total = int(rounded_total)
        result.append((customer_names[lower_customer], float(rounded_total)))

    result.sort(key=lambda x: (-x[1], x[0]))
    return result[:n]
```
Coding: Slugify Worked
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    # Step 1: Transliterate and drop non-ASCII
    slug = unicodedata.normalize('NFKD', title)
    slug = ''.join(c for c in slug if not unicodedata.combining(c))
    slug = ''.join(c if c.isascii() else '-' for c in slug)

    # Step 2: Lowercase
    slug = slug.lower()

    # Step 3: Replace non-alphanumeric runs with single '-'
    slug = re.sub(r'[^a-z0-9]+', '-', slug)
    slug = slug.strip('-')

    # Step 4: Handle max_len
    if len(slug) > max_len:
        words = slug.split('-')
        if not words:
            return ''
        # Find the longest prefix of whole words <= max_len
        truncated = []
        current_len = 0
        for word in words:
            new_len = current_len + len(word) + (1 if truncated else 0)
            if new_len > max_len:
                break
            truncated.append(word)
            current_len = new_len
        slug = '-'.join(truncated)
        # If even first word is too long, hard-cut
        if not truncated and words:
            slug = words[0][:max_len]
        # Ensure we don't exceed max_len due to joining
        if len(slug) > max_len:
            slug = slug[:max_len]

    # Step 5: Ensure no trailing '-'
    slug = slug.rstrip('-')
    return slug
```
Coding: Token bucket Worked
```python
class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._tokens = capacity
        self._last_update = clock()

    def _refill(self):
        now = self._clock()
        elapsed = now - self._last_update
        self._tokens = min(self._capacity, self._tokens + elapsed * self._refill_per_sec)
        self._last_update = now

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.95,
    "store_credit": 0.05,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 0% right
{
  "severity": {"0": 0.0, "1": 0.0, "2": 0.95, "3": 0.05},
  "page": {"true": 0.95, "false": 0.05}
}
Decisions: Tool followup (hard) 100% right
{"tool": {"web_search": 0, "calculator": 0, "calendar": 1, "email": 0, "none": 0}, "confirm": {"true": 1, "false": 0}}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {"true": 0.01, "false": 0.99},
  "action_needed": {"true": 0.05, "false": 0.95}
}
Decisions: Meeting slot (hard) 50% right
{
  "slot": {"A": 0.0, "B": 1.0, "C": 0.0, "D": 0.0},
  "raj_last": {"true": 0.0, "false": 1.0}
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {"true": 1.0, "false": 0.0},
  "support": {"true": 1.0, "false": 0.0}
}
Decisions: Support checkout down 100% right
{
  "department": {"billing": 0.1, "technical": 0.85, "account": 0.01, "shipping": 0.01, "sales": 0.03},
  "urgency": {"0": 0.0, "1": 0.0, "2": 0.05, "3": 0.95},
  "outage": {"true": 0.98, "false": 0.02}
}
Decisions: Refund wrong plan 100% right
{
  "department": {"billing": 0.99, "technical": 0.001, "account": 0.005, "shipping": 0.001, "sales": 0.003},
  "refund": {"true": 0.99, "false": 0.01},
  "tone": {"frustrated": 0.05, "calm": 0.95}
}
Decisions: Moderation doxxing 100% right
{
  "policy": {"none": 0.0, "harassment": 0.95, "hate": 0.0, "spam": 0.0, "self_harm": 0.05},
  "personal_info": {"true": 0.99, "false": 0.01}
}
Decisions: Route calendar 100% right
{"tool": {"calendar": 0.9, "email": 0.1, "web_search": 0.0, "calculator": 0.0, "none": 0.0}, "confirm": {"true": 1.0, "false": 0.0}}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {"invoice": 0.99, "resume": 0.001, "contract": 0.001, "bank_statement": 0.001, "other": 0.007},
  "missing_due_date": {"true": 0.95, "false": 0.05}
}
Decisions: Phishing paypal 100% right
{
  "phishing": {"true": 0.99, "false": 0.01},
  "risk": {"0": 0.01, "1": 0.01, "2": 0.08, "3": 0.9}
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {"none": 0.0, "contact": 0.1, "financial": 0.0, "government_id": 0.9, "health": 0.0},
  "sensitive": {"true": 0.9, "false": 0.1}
}
Decisions: Review mixed 100% right
{
  "sentiment": {"positive": 0.1, "neutral": 0.1, "negative": 0.8},
  "defect": {"true": 0.9, "false": 0.1},
  "recommend": {"true": 0.05, "false": 0.95}
}
Documents: Saas escalator (hard) 20% right
{
  "year2_price_per_seat_month": 46.58,
  "year3_price_per_seat_month": 46.58,
  "year1_invoice": 52920.0,
  "year2_invoice": 54152.64,
  "addon_months_billed": 7,
  "addon_invoice": 58034.88,
  "year3_invoice": 102018.88,
  "year3_discount_percent": 15,
  "total_contract_value": 267126.4,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 71% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {"date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60},
    {"date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88},
    {"date": "2025-02-25", "category": "meals", "amount_usd": 229.39},
    {"date": "2025-02-26", "category": "lodging", "amount_usd": 466.56}
  ],
  "rejected_item_count": 2,
  "per_diem_days": 2,
  "per_diem_usd": 130.00,
  "total_reimbursable_usd": 2019.43,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 100% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.0,
  "monthly_rent_from_2025_06_01": 2236.0,
  "late_fee_from_2025_06_01": 111.8,
  "security_deposit": 2150.0,
  "total_pet_deposits": 800.0,
  "total_monthly_payment_july_2025": 2306.0,
  "move_in_payment": 4700.0
}
Documents: Ticket SLA 75% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address", "invoice_pdf"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T17:00",
  "sla_due_utc": "2025-09-15T22:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 89% right
{
  "q3_total_usd": 15346000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 6.1,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 125B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.