- Llama 3.3 70B Instruct is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 46 out of 100, #33 of 56.
- It solved 14 of 30 coding jobs and scored 45 on reading documents. On our hardest tasks it scored 27.
- Runs on a Mac with 64 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.3 70B Instruct got 14 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.3 70B Instruct scored 45; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 47 | 43 | 48 |
| Reading documents | 45 | 60 | 41 |
| Decisions | 86 | 89 | 85 |
On the 18 hardest tasks (included in the scores above) it scored 27. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Llama 3.3 70B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | No | – |
| Graphics card, 32 GB | No | – |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | No | – |
| Mac, 48 GB | No | – |
| Mac, 64 GB | Yes | Q4 |
| Mac, 96 GB | Yes | Q6 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
Search for "Llama 3.3 70B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.
Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Llama 3.3 70B Instruct wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict
def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
"""
Parse an INI file and return a dictionary of sections and their key-value pairs.
:param text: The INI file content as a string
:return: A dictionary of sections and their key-value pairs
"""
sections: Dict[str, Dict[str, str]] = {}
default_section: Dict[str, str] = {}
current_section: Dict[str, str] = default_section
current_key: str = ""
current_value: str = ""
for line in text.splitlines():
# Remove leading and trailing whitespace
line = line.strip()
# Ignore empty lines and comments
if not line or line[0] in [";", "#"]:
continue
# Check for section headers
match = re.match(r"\s*\[(.*?)\]\s*", line)
if match:
section_name = match.group(1)
if section_name == "DEFAULT":
current_section = default_section
else:
current_section = sections.setdefault(section_name, {})
current_key = ""
current_value = ""
continue
# Check for key-value pairs
match = re.match(r"\s*(\S+)\s*([:=])\s*(.*)", line)
if match:
key = match.group(1).strip().lower()
value = match.group(3).strip()
current_key = key
current_value = value
current_section[key] = value
continue
# Check for continuation lines
if line and current_key:
current_value += "\n" + line.strip()
current_section[current_key] = current_value
continue
# Raise an error for malformed lines
raise ValueError("Malformed line: " + line)
# Interpolate values
def interpolate(value: str, section: Dict[str, str]) -> str:
def replace(match: re.Match) -> str:
ref = match.group(1)
if ":" in ref:
section_name, key = ref.split(":", 1)
if section_name not in sections:
raise KeyError(f"Section '{section_name}' not found")
section_ref = sections[section_name]
if key not in section_ref:
raise KeyError(f"Key '{key}' not found in section '{section_name}'")
return section_ref[key]
else:
if ref not in section:
raise KeyError(f"Key '{ref}' not found")
return section[ref]
while True:
new_value = re.sub(r"\${(.*?)}", replace, value)
if new_value == value:
break
value = new_value
return value
for section_name, section in sections.items():
for key, value in section.items():
# Check for reference cycles
try:
sections[section_name][key] = interpolate(value, {**default_section, **section})
except RecursionError:
raise ValueError("Reference cycle detected")
return {section_name: section for section_name, section in sections.items() if section_name != "DEFAULT"}
``` Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from time import monotonic
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock=monotonic):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
def _remove_expired(self):
current_time = self.clock()
expired_keys = [key for key, (_, expiry) in self.cache.items() if current_time >= expiry]
for key in expired_keys:
del self.cache[key]
self.stats["expirations"] += 1
def put(self, key, value):
if self.capacity == 0:
return
self._remove_expired()
if key in self.cache:
del self.cache[key]
elif len(self.cache) >= self.capacity:
self.cache.popitem(last=False)
self.stats["evictions"] += 1
expiry = self.clock() + self.ttl
self.cache[key] = (value, expiry)
self.cache.move_to_end(key)
def get(self, key, default=None):
if self.capacity == 0:
self.stats["misses"] += 1
return default
self._remove_expired()
if key not in self.cache:
self.stats["misses"] += 1
return default
value, expiry = self.cache[key]
current_time = self.clock()
if current_time >= expiry:
del self.cache[key]
self.stats["expirations"] += 1
self.stats["misses"] += 1
return default
self.cache.move_to_end(key)
self.stats["hits"] += 1
return value
def __len__(self):
self._remove_expired()
return len(self.cache)
def stats(self):
return {
"hits": self.stats["hits"],
"misses": self.stats["misses"],
"evictions": self.stats["evictions"],
"expirations": self.stats["expirations"],
}
``` Coding: Parse duration Didn't work
```python
import re
from typing import Pattern
def parse_duration(s: str) -> int:
"""
Converts a duration string into a total number of seconds.
Args:
s (str): The duration string.
Returns:
int: The total number of seconds.
Raises:
ValueError: If the string is empty, contains invalid units, or has units out of order.
"""
# Remove leading and trailing whitespace
s = s.strip()
# Check for empty string
if not s:
raise ValueError("Empty string")
# Regular expression pattern to match duration parts
pattern: Pattern[str] = re.compile(r"(\d+)([hms])", re.IGNORECASE)
# Initialize total seconds and unit order check
total_seconds = 0
seen_units = set()
# Split the string into parts
parts = re.split(r"\s+", s)
# Iterate over each part
for part in parts:
# Check if the part matches the pattern
match = pattern.fullmatch(part)
# If the part does not match the pattern, raise an error
if not match:
raise ValueError(f"Invalid part: {part}")
# Get the value and unit from the match
value = int(match.group(1))
unit = match.group(2).lower()
# Check for repeated units
if unit in seen_units:
raise ValueError(f"Repeated unit: {unit}")
# Check for units out of order
if unit == 'm' and 'h' not in seen_units:
pass
elif unit == 's' and ('h' not in seen_units or 'm' not in seen_units):
pass
elif unit == 'h' and ('m' in seen_units or 's' in seen_units):
raise ValueError("Units out of order")
elif unit == 'm' and 's' in seen_units:
raise ValueError("Units out of order")
# Add the unit to the seen units
seen_units.add(unit)
# Convert the value to seconds and add to the total
if unit == 'h':
total_seconds += value * 3600
elif unit == 'm':
total_seconds += value * 60
elif unit == 's':
total_seconds += value
return total_seconds
``` Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
"""
This function takes a list of integers, sorts and removes duplicates,
then collapses runs of 3 or more consecutive integers into "a..b"
and numbers in shorter runs (1 or 2 numbers) are listed individually.
Args:
nums (list[int]): A list of integers.
Returns:
str: A string representing the summarized ranges.
"""
# If the list is empty, return an empty string
if not nums:
return ""
# Sort the numbers and remove duplicates by converting to a set
nums = sorted(set(nums))
# Initialize variables to keep track of the current run
current_run = [nums[0]]
result = []
# Iterate over the sorted numbers
for num in nums[1:]:
# If the current number is a continuation of the current run, add it to the run
if num == current_run[-1] + 1:
current_run.append(num)
# If the current number is not a continuation of the current run,
# summarize the current run and start a new one
else:
# If the current run has 3 or more numbers, summarize it as "a..b"
if len(current_run) >= 3:
result.append(f"{current_run[0]}..{current_run[-1]}")
# If the current run has 1 or 2 numbers, list them individually
else:
result.extend(map(str, current_run))
# Start a new run with the current number
current_run = [num]
# Summarize the last run
if len(current_run) >= 3:
result.append(f"{current_run[0]}..{current_run[-1]}")
else:
result.extend(map(str, current_run))
# Join everything with "," and return the result
return ",".join(result)
``` Coding: Top customers Worked
```python
from collections import defaultdict
from decimal import Decimal, InvalidOperation
import re
def top_customers(rows, n):
"""
Returns the top n customers with the highest total amount.
:param rows: A list of dictionaries containing customer information.
:param n: The number of top customers to return.
:return: A list of tuples containing the customer name and total amount.
"""
# Create a dictionary to store the total amount for each customer
customer_amounts = defaultdict(Decimal)
# Iterate over each row
for row in rows:
# Get the customer name and amount
customer = row.get("customer", "").strip()
amount = row.get("amount", "")
# Skip rows with empty customer name or invalid amount
if not customer or not amount:
continue
# Parse the amount
try:
# Remove dollar signs and commas, and convert to Decimal
amount = Decimal(re.sub(r"[^\d.-]", "", amount))
except InvalidOperation:
# Skip rows with invalid amount
continue
# Update the customer's total amount
customer_amounts[customer.lower()] += amount
# Round the total amounts to 2 decimals and store the original customer name
customer_amounts = {
next(key for key in rows if key.get("customer", "").strip().lower() == customer).get("customer", "").strip(): round(amount, 2)
for customer, amount in customer_amounts.items()
}
# Sort the customers by total amount in descending order, then by name in ascending order
sorted_customers = sorted(customer_amounts.items(), key=lambda x: (-x[1], x[0]))
# Return the top n customers
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import unicodedata
import re
def slugify(title: str, max_len: int = 50) -> str:
# Transliterate accented characters to ASCII and drop any other non-ASCII characters
title = ''.join(c for c in unicodedata.normalize('NFD', title) if unicodedata.category(c) != 'Mn' and c.isascii())
# Lowercase
title = title.lower()
# Replace every run of characters other than a-z and 0-9 with a single "-", and strip leading/trailing "-"
title = re.sub(r'[^a-z0-9]+', '-', title).strip('-')
# If the slug is longer than `max_len`, shorten it without cutting a word
if len(title) > max_len:
words = title.split('-')
slug = ''
for word in words:
if len(slug) + len(word) + 1 <= max_len:
slug += word + '-'
else:
break
slug = slug.rstrip('-')
if not slug: # If even the first word is longer than max_len, hard-cut it to max_len characters
slug = title[:max_len]
else:
slug = title
return slug
``` Coding: Token bucket Worked
```python
from typing import Callable
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
self._capacity = capacity
self._refill_per_sec = refill_per_sec
self._clock = clock
self._last_refill = clock()
self._tokens = capacity
def _refill(self) -> None:
now = self._clock()
elapsed = now - self._last_refill
self._tokens = min(self._capacity, self._tokens + elapsed * self._refill_per_sec)
self._last_refill = now
def allow(self, cost: float = 1) -> bool:
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
return False
@property
def tokens(self) -> float:
self._refill()
return self._tokens
``` Decisions: Refund window (hard) 50% right
```json
{
"outcome": {
"full_refund": 0.0,
"partial_refund": 0.8,
"store_credit": 0.2,
"nothing": 0.0
},
"defective": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Incident matrix (hard) 100% right
```json
{
"severity": {
"0": 0.0,
"1": 0.7,
"2": 0.3,
"3": 0.0
},
"page": {
"true": 0.3,
"false": 0.7
}
}
``` Decisions: Tool followup (hard) 100% right
{"tool": {"web_search": 0.0, "calculator": 0.0, "calendar": 1.0, "email": 0.0, "none": 0.0}, "confirm": {"true": 1.0, "false": 0.0}} Decisions: Legit security alert (hard) 100% right
```json
{
"phishing": {
"true": 0.01,
"false": 0.99
},
"action_needed": {
"true": 0.05,
"false": 0.95
}
}
``` Decisions: Meeting slot (hard) 0% right
```
{
"slot": {
"A": 0.0,
"B": 0.0,
"C": 0.0,
"D": 1.0
},
"raj_last": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Review mixed (hard) 100% right
{"hardware": {"true": 0.9, "false": 0.1}, "support": {"true": 0.8, "false": 0.2}} Decisions: Support checkout down 100% right
```json
{
"department": {
"billing": 0.2,
"technical": 0.7,
"account": 0.05,
"shipping": 0.0,
"sales": 0.05
},
"urgency": {
"0": 0.0,
"1": 0.0,
"2": 0.2,
"3": 0.8
},
"outage": {
"true": 0.9,
"false": 0.1
}
}
``` Decisions: Refund wrong plan 100% right
```json
{
"department": {
"billing": 0.9,
"technical": 0.05,
"account": 0.03,
"shipping": 0.01,
"sales": 0.01
},
"refund": {
"true": 0.95,
"false": 0.05
},
"tone": {
"frustrated": 0.1,
"calm": 0.9
}
}
``` Decisions: Moderation doxxing 100% right
```
{
"policy": {
"none": 0.0,
"harassment": 0.8,
"hate": 0.0,
"spam": 0.0,
"self_harm": 0.2
},
"personal_info": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Route calendar 100% right
```
{
"tool": {
"web_search": 0.0,
"calculator": 0.0,
"calendar": 0.8,
"email": 0.1,
"none": 0.1
},
"confirm": {
"true": 0.9,
"false": 0.1
}
}
``` Decisions: Doc invoice missing due 100% right
```
{
"doc_type": {
"invoice": 0.99,
"resume": 0.0,
"contract": 0.0,
"bank_statement": 0.0,
"other": 0.01
},
"missing_due_date": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Phishing paypal 100% right
```json
{
"phishing": {
"true": 0.9,
"false": 0.1
},
"risk": {
"0": 0.05,
"1": 0.05,
"2": 0.1,
"3": 0.8
}
}
``` Decisions: Pii ssn email 100% right
```
{
"data_kind": {
"none": 0.0,
"contact": 0.2,
"financial": 0.3,
"government_id": 0.5,
"health": 0.0
},
"sensitive": {
"true": 0.8,
"false": 0.2
}
}
``` Decisions: Review mixed 100% right
{"sentiment": {"positive": 0.0, "neutral": 0.0, "negative": 1.0}, "defect": {"true": 1.0, "false": 0.0}, "recommend": {"true": 0.0, "false": 1.0}} Documents: Saas escalator (hard) 20% right
```
{
"year2_price_per_seat_month": 47.53,
"year3_price_per_seat_month": 49.98,
"year1_invoice": 4860.00,
"year2_invoice": 5719.20,
"addon_months_billed": 5,
"addon_invoice": 2145.60,
"year3_invoice": 7188.80,
"year3_discount_percent": 15,
"total_contract_value": 17923.60,
"contract_end_date": "2027-02-28"
}
``` Documents: Expense thread 75% right
```
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{ "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
{ "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
{ "date": "2025-02-25", "category": "meals", "amount_usd": 229.33 },
{ "date": "2025-02-26", "category": "lodging", "amount_usd": 466.56 },
{ "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.68 }
],
"rejected_item_count": 2,
"per_diem_days": 2,
"per_diem_usd": 130.00,
"total_reimbursable_usd": 2064.05,
"approver_email": "priya.raman@corvane.com"
}
``` Documents: Lease amendment 92% right
```
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2236.00,
"late_fee_from_2025_06_01": 111.80,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2306.00,
"move_in_payment": 4650.00
}
``` Documents: Ticket SLA 82% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-16T13:30",
"sla_due_utc": "2025-09-16T18:30:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 33% right
{
"q3_total_usd": 14946,
"q2_total_usd": 14414,
"q2_central_originally_reported_usd": 3047,
"q2_to_q3_change_pct": 3.7,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731,
"west_excluding_mountain_q3_usd": 4201
} Size: 71B parameters. First tested OCT 11.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.