- Llama 3.2 3B Instruct is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 6 out of 100, #55 of 56.
- It solved 0 of 30 coding jobs and scored 13 on reading documents. On our hardest tasks it scored 2.
- Runs on an 8 GB graphics card or a Mac with 16 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.2 3B Instruct got 0 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.2 3B Instruct scored 13; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 0 | 0 | 0 |
| Reading documents | 13 | 24 | 10 |
| Decisions | 44 | 42 | 45 |
On the 18 hardest tasks (included in the scores above) it scored 2. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Llama 3.2 3B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | Yes | Q8 |
| Graphics card, 12 GB | Yes | Q8 |
| Graphics card, 16 GB | Yes | Q8 |
| Graphics card, 24 GB | Yes | Q8 |
| Graphics card, 32 GB | Yes | Q8 |
| Mac, 16 GB | Yes | Q8 |
| Mac, 24 GB | Yes | Q8 |
| Mac, 32 GB | Yes | Q8 |
| Mac, 48 GB | Yes | Q8 |
| Mac, 64 GB | Yes | Q8 |
| Mac, 96 GB | Yes | Q8 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run llama3.2:3b Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Llama 3.2 3B Instruct wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict
def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
"""
Parse an INI file.
Args:
text (str): The text of the INI file.
Returns:
Dict[str, Dict[str, str]]: A dictionary of sections, where each section is another dictionary of key-value pairs.
"""
# Split the text into lines
lines = text.splitlines()
# Initialize the result dictionary
result = {}
# Initialize the current section
current_section = None
# Initialize the current section's dictionary
current_section_dict = {}
# Initialize the current value
current_value = ""
# Iterate over each line
for line in lines:
# Remove leading and trailing whitespace
line = line.strip()
# If the line is empty, ignore it
if not line:
continue
# If the line is a section header
if line.startswith('[') and line.endswith(']'):
# Extract the section name
section_name = line[1:-1]
# If the section is DEFAULT, inherit its keys
if section_name == 'DEFAULT':
section_name = None
# Update the current section
current_section = section_name
# Update the current section's dictionary
current_section_dict = {}
# If the section does not exist in the result, add it
if current_section not in result:
result[current_section] = {}
# If the line is a key-value pair
elif '=' in line:
# Split the line into key and value
key, value = line.split('=', 1)
# Strip leading and trailing whitespace from key and value
key = key.strip()
value = value.strip()
# If the line is a comment, ignore it
if value.startswith(('#', ';'):
continue
# If the line is a quoted value
if value.startswith('"') and value.endswith('"'):
# Extract the value without quotes
value = value[1:-1]
# If the key is a reference to another section or key
if '${' in key:
# Extract the reference
reference = key[2:-1]
# If the reference is a section, use its value
if ':' in reference:
section, key = reference.split(':', 1)
section = section.strip()
if section not in result:
result[section] = {}
key = key.strip()
# If the reference is a key, use its value
else:
key = reference.strip()
# Try to get the value from the current section or the result
try:
value = result[current_section][key]
except KeyError:
# If the key is not found, raise a KeyError
raise KeyError(f"Key '{key}' not found")
# Add the key-value pair to the current section's dictionary
current_section_dict[key] = value
# If the line is a continuation
elif line.lstrip().startswith(' '):
# Strip leading whitespace from the value
value = line.lstrip().strip()
# Add the value to the current value
current_value += value + "\n"
# If the line is an inline comment
elif line.lstrip().startswith(('#', ';'):
# Ignore the line
continue
# If the line is not recognized, raise a ValueError
else:
raise ValueError(f"Unrecognized line: {line}")
# If there is a current value, add it to the current section's dictionary
if current_value:
current_section_dict[current_value.strip()] = None
# If the current section is not None, add it to the result
if current_section is not None:
result[current_section] = current_section_dict
# Return the result
return result
``` Coding: TTL LRU (hard) Didn't work
```python
import heapq
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = {}
self.heap = []
def put(self, key, value):
if not self.capacity:
return
if key in self.cache:
self.cache[key].expiry = self.clock() + self.ttl
self._mark_most_recent(key)
self.heapify()
else:
if len(self.cache) >= self.capacity:
self._evict_least_recent()
self.cache[key] = {"value": value, "expiry": self.clock() + self.ttl, "recency": self.clock()}
self._mark_most_recent(key)
self.heapify()
def get(self, key, default=None):
if key not in self.cache:
if default is None:
return None
return default
# if key in self.cache:
# del self.cache[key]
# return default
if self.cache[key]["expiry"] < self.clock():
del self.cache[key]
return default
else:
self.cache[key]["recency"] = self.clock()
return self.cache[key]["value"]
def __len__(self):
return len(self.cache)
def stats(self):
hits = sum(1 for _ in self.cache.values() if self.clock() < self.cache[_]["expiry"])
misses = len(self.cache) - hits
evictions = len(self.heap)
expirations = sum(1 for _ in self.cache.values() if self.clock() >= self.cache[_]["expiry"])
return {"hits": hits, "misses": misses, "evictions": evictions, "expirations": expirations}
def _mark_most_recent(self, key):
if key in self.heap:
index = self.heap.index(key)
self.heap[index] = (self.cache[key]["expiry"], self.cache[key]["recency"], key)
heapq.heapify(self.heap)
else:
heapq.heappush(self.heap, (self.cache[key]["expiry"], self.cache[key]["recency"], key))
def _evict_least_recent(self):
_, _, key = heapq.heappop(self.heap)
del self.cache[key]
```
Note that the `clock` function is assumed to be a zero-argument callable that returns the current time in seconds. Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
"""
Converts a duration string into a total number of seconds.
Args:
s (str): A duration string with units of h, m, s.
Returns:
int: The total number of seconds.
Raises:
ValueError: For invalid input strings.
"""
# Remove leading and trailing whitespace
s = s.strip()
# Check if the string is empty
if not s:
raise ValueError("Input string is empty")
# Split the string into parts
parts = s.split()
# Initialize variables to store the total seconds and the unit multiplier
total_seconds = 0
multipliers = {"h": 3600, "m": 60, "s": 1}
# Check if the input string contains any invalid characters
for part in parts:
# Check if the part is a digit
if not part.isdigit():
raise ValueError("Invalid character in input string")
# Check if the part has a unit
if len(part) > 1 and part[-1] in multipliers:
# Check if the unit is valid
if part[-1] not in multipliers:
raise ValueError("Unknown unit in input string")
# Extract the number and unit
num = int(part[:-1])
# Check if the number is non-negative
if num < 0:
raise ValueError("Negative number in input string")
# Add the number of seconds to the total
total_seconds += num * multipliers[part[-1]]
else:
# If the part is not a number with a unit, raise an error
raise ValueError("Input string contains invalid format")
# Check if the total number of units is valid
if len(parts) > 3:
raise ValueError("Too many units in input string")
if len(parts) < 2:
raise ValueError("Too few units in input string")
# Return the total number of seconds
return total_seconds
``` Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
"""
This function takes a list of integers, removes duplicates, sorts the numbers,
collapses runs of 3 or more consecutive integers into "a..b", and returns a string
of the resulting numbers joined with commas.
Args:
nums (list[int]): A list of integers.
Returns:
str: A string of the summarized numbers.
"""
# First, remove duplicates by converting the list to a set
nums = set(nums)
# If the set is empty, return an empty string
if not nums:
return ""
# Sort the set in ascending order
nums = sorted(nums)
# Initialize an empty list to store the result
result = []
# Initialize the start and end of the current run
start = end = nums[0]
# Iterate over the sorted numbers
for num in nums[1:]:
# If the current number is a consecutive integer to the previous one, update the end of the run
if num == end + 1:
end = num
# If the current number is not a consecutive integer, append the current run to the result and start a new run
else:
if start == end:
result.append(str(start))
else:
result.append(f"{start}..{end}")
start = end = num
# Append the last run to the result
if start == end:
result.append(str(start))
else:
result.append(f"{start}..{end}")
# Join the result with commas and return it
return ",".join(result)
``` Coding: Top customers Didn't work
```python
from collections import defaultdict
from typing import List, Dict, Tuple
def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
"""
This function takes a list of dictionaries representing customer transactions and returns the top n customers with the highest total amount spent.
:param rows: A list of dictionaries, where each dictionary has a "customer" key and an "amount" key.
:param n: The number of top customers to return.
:return: A list of tuples, where each tuple contains the customer name and their total amount spent.
"""
# Create a dictionary to store the total amount spent by each customer
customer_amounts = defaultdict(float)
# Iterate over each row in the input list
for row in rows:
# Get the customer name and amount from the current row
customer = row["customer"].strip().lower()
amount = float(row["amount"].strip("$").replace(",", ""))
# If the amount is negative, consider it as a refund
if amount < 0:
amount = -amount
# Add the amount to the customer's total
customer_amounts[customer] += amount
# Sort the customers by total amount in descending order, then by name
sorted_customers = sorted(customer_amounts.items(), key=lambda x: (-x[1], x[0]))
# Return the top n customers
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import unicodedata
import re
def slugify(title: str, max_len: int = 50) -> str:
"""
Builds a URL slug from a given title.
Args:
title (str): The title to slugify.
max_len (int): The maximum length of the slug. Defaults to 50.
Returns:
str: The slugified title.
"""
# Transliterate accented characters to ASCII and drop any other non-ASCII characters
title = ''.join(c for c in unicodedata.normalize('NFD', title) if unicodedata.category(c) != 'Mn')
# Lowercase
title = title.lower()
# Replace every run of characters other than a-z and 0-9 with a single "-"
title = re.sub(r'[^a-z0-9]+', '-', title)
# Strip leading/trailing "-"
title = title.strip('-')
# Split into words
words = title.split('-')
# Initialize the result and the prefix length
result = ''
prefix_len = 0
# Iterate over the words
for word in words:
# If adding the word would exceed max_len, break
if prefix_len + len(word) + 1 > max_len:
break
# Add the word to the result and update the prefix length
result += word + '-'
prefix_len += len(word) + 1
# Remove the trailing '-'
result = result[:-1]
return result
``` Coding: Token bucket Didn't work
```python
import time
from typing import Callable
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
"""
Initialize the token bucket.
:param capacity: The maximum number of tokens in the bucket.
:param refill_per_sec: The rate at which tokens are refilled per second.
:param clock: A function that returns the current time in seconds.
"""
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.last_refill = clock()
self.tokens = capacity
self.clock = clock
@property
def tokens(self) -> float:
"""
Get the current number of tokens in the bucket.
:return: The current number of tokens.
"""
return self._tokens
def allow(self, cost: float = 1) -> bool:
"""
Refill the bucket and attempt to allow the requested number of tokens.
:param cost: The number of tokens to attempt to get. Defaults to 1.
:return: True if tokens were successfully obtained, False otherwise.
"""
elapsed = self.clock() - self.last_refill
self.last_refill = self.clock()
self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
if self.tokens >= cost:
self.tokens -= cost
return True
return False
``` Decisions: Refund window (hard) 50% right
```json
{
"outcome": {
"full_refund": 0.6,
"partial_refund": 0.3,
"store_credit": 0.1,
"nothing": 0
},
"defective": {
"true": false,
"false": 1
}
}
``` Decisions: Incident matrix (hard) 0% right
```json
{
"severity": {
"0": 0.95,
"1": 0.05,
"2": 0,
"3": 0
},
"page": {
"true": 0.5,
"false": 0.5
}
}
``` Decisions: Tool followup (hard) 50% right
```json
{
"tool": {
"web_search": 0.0,
"calculator": 0.0,
"calendar": 0.9,
"email": 0.1,
"none": 0.0
},
"confirm": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Legit security alert (hard) 0% right
```json
{
"1": {
"true": 0.99,
"false": 0.01
},
"2": {
"true": 0.95,
"false": 0.05
}
}
``` Decisions: Meeting slot (hard) 0% right
{
"slot": {
"A": 0.0, "B": 0.0, "C": 0.0, "D": 1.0
},
"raj_last": {
"true": 0.0, "false": 1.0
} Decisions: Review mixed (hard) 50% right
```json
{
"hardware": {
"true": 1.0,
"false": 0.0
},
"support": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Support checkout down 33% right
```json
{
"department": {
"billing": 0.6,
"technical": 0.3,
"account": 0.1
},
"urgency": {
"0": 0.0,
"1": 0.4,
"2": 0.3,
"3": 0.3
},
"outage": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Refund wrong plan 100% right
```json
{
"department": {
"billing": 0.9,
"technical": 0.05,
"account": 0.05,
"shipping": 0.0,
"sales": 0.0
},
"refund": {
"true": 0.95,
"false": 0.05
},
"tone": {
"frustrated": 0.0,
"calm": 1.0
}
}
``` Decisions: Moderation doxxing 50% right
{
"policy": {
"none": 0.95,
"harassment": 0,
"hate": 0,
"spam": 0,
"self_harm": 0
},
"personal_info": {
"true": 0.5,
"false": 0.5
}
} Decisions: Route calendar 100% right
```json
{
"tool": {
"calendar": 0.9,
"email": 0.05,
"web_search": 0.05,
"none": 0.005
},
"confirm": {
"true": 0.8,
"false": 0.2
}
}
``` Decisions: Doc invoice missing due 50% right
```json
{
"doc_type": {
"invoice": 1.0,
"resume": 0.0,
"contract": 0.0,
"bank_statement": 0.0,
"other": 0.0
},
"missing_due_date": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Phishing paypal 0% right
{
"phishing": {
"true": 0.9,
"false": 0.1
},
"risk": {
"0": 0.05,
"1": 0.3,
"2": 0.4,
"3": 0.15
} Decisions: Pii ssn email 0% right
```json
{
"1": {
"none": 0.8,
"contact": 0.1,
"financial": 0.05,
"government_id": 0.05,
"health": 0.0
},
"2": {
"true": 0.1,
"false": 0.9
}
}
``` Decisions: Review mixed 100% right
```json
{
"sentiment": {
"positive": 0,
"neutral": 0,
"negative": 0.9
},
"defect": {
"true": 1,
"false": 0
},
"recommend": {
"true": 0,
"false": 1
}
}
``` Documents: Saas escalator (hard) 0% right
{
"year2_price_per_seat_month": 47.69,
"year3_price_per_seat_month": 49.13,
"year1_invoice": 5400.00,
"year2_invoice": 5688.00,
"addon_months_billed": 16,
"addon_invoice": 752.80,
"year3_invoice": 8321.60,
"year3_discount_percent": 0.00,
"total_contract_value": 219440.00,
"contract_end_date": "2029-02-28" Documents: Expense thread 67% right
```
{
"employee_id": "dana.whitfield@corvane.com",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{"date": "2025-02-24", "category": "airfare", "amount_usd": 1179.32},
{"date": "2025-02-24", "category": "ground_transport", "amount_usd": 40.00},
{"date": "2025-02-25", "category": "meals", "amount_usd": 65.00},
{"date": "2025-02-26", "category": "lodging", "amount_usd": 475.52},
{"date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.83}
],
"rejected_item_count": 0,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 1938.17,
"approver_email": "priya.raman@corvane.com"
}
``` Documents: Lease amendment 0% right
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150,
"monthly_rent_from_2025_06_01": 2250,
"late_fee_from_2025_06_01": 107.5,
"security_deposit": 2150,
"total_pet_deposits": 800,
"total_monthly_payment_july_2025": 2275,
"move_in_payment": 2150 + 2150 + 400 = 2770
} Documents: Ticket SLA 54% right
```json
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address", "invoice_pdf"],
"affected_orders": ["SO-99812", "SO-99815", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-15T17:00",
"sla_due_utc": "2025-09-15T17:00:00Z",
"reissued_invoice": "INV-2025-0812"
}
``` Documents: Sales footnotes 0% right
{
"q3_total_usd": 4902,
"q2_total_usd": 4585,
"q2_central_originally_reported_usd": 2947,
"q2_to_q3_change_pct": 5.6,
"top_region_q3": "West",
"fastest_growing_region_q1_to_q3": null,
"regions_declining_q2_to_q3": null,
"international_q3_organic_usd": 1951,
"west_excluding_mountain_q3_usd": 4280
} Size: 3.2B parameters. First tested OCT 10.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.