- Llama 3.1 70B Instruct is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 39 out of 100, #41 of 56.
- It solved 10 of 30 coding jobs and scored 44 on reading documents. On our hardest tasks it scored 22.
- Runs on a Mac with 64 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.1 70B Instruct got 10 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.1 70B Instruct scored 44; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 33 | 29 | 35 |
| Reading documents | 44 | 55 | 41 |
| Decisions | 84 | 82 | 85 |
On the 18 hardest tasks (included in the scores above) it scored 22. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download Llama 3.1 70B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | No | – |
| Graphics card, 32 GB | No | – |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | No | – |
| Mac, 48 GB | No | – |
| Mac, 64 GB | Yes | Q4 |
| Mac, 96 GB | Yes | Q6 |
| Mac, 128 GB | Yes | Q8 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
Search for "Llama 3.1 70B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.
Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Llama 3.1 70B Instruct wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
import re
def parse_ini(text: str) -> dict[str, dict[str, str]]:
"""
Parse INI text into a dictionary of sections and their key-value pairs.
:param text: The INI text to parse
:return: A dictionary of sections and their key-value pairs
"""
sections = {}
default_section = {}
current_section = None
for line in text.splitlines():
# Ignore lines that are empty or whose first non-space character is a comment
line = line.strip()
if not line or line.startswith((';', '#')):
continue
# Check for section headers
if line.startswith('[') and line.endswith(']'):
section_name = line[1:-1].strip()
if section_name == 'DEFAULT':
current_section = default_section
else:
current_section = sections.setdefault(section_name, {})
continue
# If we haven't seen a section header yet, raise ValueError
if current_section is None:
raise ValueError("Key line before any section header")
# Split the line into key and value
match = re.match(r'([^:=]+)[:=](.*)', line)
if not match:
raise ValueError("Malformed line")
key, value = match.groups()
key = key.strip().lower()
# Handle continuation lines
while True:
next_line = text.splitlines()[text.splitlines().index(line) + 1].strip()
if next_line and next_line.startswith(' '):
value += '\n' + next_line.strip()
text = text.replace(next_line, '', 1)
else:
break
# Remove inline comments and double quotes
value = re.sub(r'(;| #).*$', '', value)
if value.startswith('"') and value.endswith('"'):
value = value[1:-1]
else:
value = value.strip()
# Store the key-value pair
current_section[key] = value
# Interpolate values
for section_name, section in sections.items():
for key, value in section.items():
while True:
match = re.search(r'\${([^:]+)(:[^}]+)?}', value)
if not match:
break
ref_key, ref_section = match.groups()
if ref_section:
ref_section = ref_section[1:]
else:
ref_section = section_name
ref_section_data = sections.get(ref_section, default_section)
ref_value = ref_section_data.get(ref_key)
if ref_value is None:
raise KeyError(f"Missing key {ref_key}")
if ref_value == value:
raise ValueError("Reference cycle")
value = value.replace(match.group(0), ref_value, 1)
# Inherit DEFAULT section keys
for section in sections.values():
section.update({k: v for k, v in default_section.items() if k not in section})
return sections
``` Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from typing import Callable, Dict, Optional
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
def put(self, key, value):
now = self.clock()
if key in self.cache:
del self.cache[key]
elif self.capacity > 0 and len(self.cache) >= self.capacity:
self._remove_expired(now)
if len(self.cache) >= self.capacity:
self.cache.popitem(last=False)
self.stats["evictions"] += 1
self.cache[key] = (value, now + self.ttl)
def get(self, key, default=None):
now = self.clock()
if key in self.cache:
value, expiry = self.cache[key]
if expiry > now:
self.cache.move_to_end(key)
self.stats["hits"] += 1
return value
else:
del self.cache[key]
self.stats["expirations"] += 1
self.stats["misses"] += 1
return default
def __len__(self):
now = self.clock()
self._remove_expired(now)
return len(self.cache)
def stats(self) -> Dict[str, int]:
return self.stats.copy()
def _remove_expired(self, now: float):
while self.cache and self.cache[next(reversed(self.cache))][1] <= now:
del self.cache[next(reversed(self.cache))]
self.stats["expirations"] += 1
``` Coding: Parse duration Didn't work
```python
import re
from enum import Enum
from typing import Match
class Unit(Enum):
HOURS = 1
MINUTES = 2
SECONDS = 3
@property
def value(self):
return {
Unit.HOURS: 3600,
Unit.MINUTES: 60,
Unit.SECONDS: 1,
}.get(self)
@classmethod
def from_str(cls, s: str) -> 'Unit':
s = s.lower()
if s == 'h':
return cls.HOURS
elif s == 'm':
return cls.MINUTES
elif s == 's':
return cls.SECONDS
else:
raise ValueError(f"Unknown unit: {s}")
def parse_duration(s: str) -> int:
"""
Convert a duration string into a total number of seconds.
:param s: A duration string.
:return: The total number of seconds.
:raises ValueError: If the input string is invalid.
"""
# Remove whitespace and normalize to lowercase
s = re.sub(r'\s+', '', s).lower()
# Check for empty/blank strings
if not s:
raise ValueError("Empty string")
# Parse each part of the duration
parts = re.findall(r'(\d+)([hms])', s)
if len(parts)!= len(s):
raise ValueError("Invalid duration string")
# Validate the order and uniqueness of units
units = [Unit.from_str(unit) for _, unit in parts]
if units!= sorted(units):
raise ValueError("Units are out of order")
if len(set(units))!= len(units):
raise ValueError("Repeated units")
# Convert the duration to seconds
total_seconds = 0
for value, unit in parts:
total_seconds += int(value) * Unit.from_str(unit).value
return total_seconds
``` Coding: Summarize ranges Worked
```python
from itertools import groupby
def summarize_ranges(nums: list[int]) -> str:
"""
Summarize a list of integers into a string, collapsing runs of 3 or more consecutive integers.
Args:
- nums: A list of integers.
Returns:
- A string representing the input list, with runs of 3 or more consecutive integers collapsed.
"""
# If the input list is empty, return an empty string
if not nums:
return ""
# Sort the numbers and remove duplicates by converting the list to a set
nums = sorted(set(nums))
# Initialize variables to store the result and the current run
result = []
current_run = [nums[0]]
# Iterate over the sorted numbers
for num in nums[1:]:
# If the current number is consecutive to the last number in the current run, add it to the run
if num == current_run[-1] + 1:
current_run.append(num)
# Otherwise, process the current run and start a new one
else:
# If the current run has 3 or more numbers, collapse it into "a..b"
if len(current_run) >= 3:
result.append(f"{current_run[0]}..{current_run[-1]}")
# If the current run has 2 numbers, add them individually
elif len(current_run) == 2:
result.extend(map(str, current_run))
# If the current run has 1 number, add it individually
else:
result.append(str(current_run[0]))
# Start a new run
current_run = [num]
# Process the last run
if len(current_run) >= 3:
result.append(f"{current_run[0]}..{current_run[-1]}")
elif len(current_run) == 2:
result.extend(map(str, current_run))
else:
result.append(str(current_run[0]))
# Join the result with commas and return
return ",".join(result)
``` Coding: Top customers Didn't work
```python
from collections import defaultdict
from decimal import Decimal, ROUND_HALF_UP
from typing import List, Dict, Tuple
def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
"""
Returns the top n customers by total amount, sorted by total descending then name ascending.
Args:
rows (List[Dict]): A list of dictionaries containing customer information.
n (int): The number of top customers to return.
Returns:
List[Tuple[str, float]]: A list of tuples containing the customer name and total amount.
"""
# Create a dictionary to store the total amount for each customer
customers = defaultdict(lambda: {"total": Decimal(0), "name": ""})
# Iterate over each row in the input list
for row in rows:
# Get the customer name and amount from the row
customer_name = row.get("customer", "").strip().casefold()
amount = row.get("amount", "")
# Skip rows with empty customer names or amounts that cannot be parsed
if not customer_name or not amount:
continue
# Parse the amount string into a Decimal object
try:
amount = Decimal(amount.replace("$", "").replace(",", ""))
except ValueError:
continue
# Update the customer's total amount and store the trimmed customer name
if not customers[customer_name]["name"]:
customers[customer_name]["name"] = row["customer"].strip()
customers[customer_name]["total"] += amount
# Round the total amounts to 2 decimals
for customer in customers.values():
customer["total"] = float(customer["total"].quantize(Decimal(".01"), rounding=ROUND_HALF_UP))
# Sort the customers by total descending then name ascending
sorted_customers = sorted(
((customer["name"], customer["total"]) for customer in customers.values()),
key=lambda x: (-x[1], x[0])
)
# Return the top n customers
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import unicodedata
import re
def slugify(title: str, max_len: int = 50) -> str:
"""
Builds a URL slug from a given title.
:param title: The title to be slugified
:param max_len: The maximum length of the slug (default is 50)
:return: The slugified title
"""
# Transliterate accented characters to ASCII and drop any other non-ASCII characters
title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')
# Lowercase
title = title.lower()
# Replace every run of characters other than a-z and 0-9 with a single "-"
title = re.sub(r'[^\w\s-]', '', title).replace(' ', '-')
# Strip leading/trailing "-"
title = title.strip('-')
# If the slug is longer than max_len, shorten it without cutting a word
if len(title) > max_len:
words = title.split('-')
slug = ''
for word in words:
if len(slug) + len(word) + 1 > max_len:
break
slug += word + '-'
slug = slug.rstrip('-')
if not slug:
slug = title[:max_len]
else:
slug = title
return slug
``` Coding: Token bucket Worked
```python
from types import SimpleNamespace
from typing import Callable
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.clock = clock
self.state = SimpleNamespace(last_refill_time=self.clock(), tokens=self.capacity)
def _refill(self):
now = self.clock()
elapsed_time = now - self.state.last_refill_time
self.state.tokens = min(self.capacity, self.state.tokens + elapsed_time * self.refill_per_sec)
self.state.last_refill_time = now
def allow(self, cost: float = 1) -> bool:
self._refill()
if self.state.tokens >= cost:
self.state.tokens -= cost
return True
return False
@property
def tokens(self) -> float:
self._refill()
return self.state.tokens
``` Decisions: Refund window (hard) 50% right
{
"outcome": {
"full_refund": 0,
"partial_refund": 0.8,
"store_credit": 0.2,
"nothing": 0
},
"defective": {
"true": 0,
"false": 1
}
} Decisions: Incident matrix (hard) 0% right
{"severity": {"0": 0, "1": 0.3, "2": 0.7, "3": 0}, "page": {"true": 0.7, "false": 0.3}} Decisions: Tool followup (hard) 100% right
{
"tool": {
"web_search": 0,
"calculator": 0,
"calendar": 1,
"email": 0,
"none": 0
},
"confirm": {
"true": 1,
"false": 0
}
} Decisions: Legit security alert (hard) 100% right
{
"phishing": {
"true": 0.05,
"false": 0.95
},
"action_needed": {
"true": 0.05,
"false": 0.95
}
} Decisions: Meeting slot (hard) 0% right
{
"slot": {
"A": 0,
"B": 0,
"C": 1,
"D": 0
},
"raj_last": {
"true": 0,
"false": 1
}
} Decisions: Review mixed (hard) 100% right
{
"hardware": {
"true": 0.9,
"false": 0.1
},
"support": {
"true": 0.8,
"false": 0.2
}
} Decisions: Support checkout down 100% right
{
"department": {
"billing": 0.2,
"technical": 0.6,
"account": 0.1,
"shipping": 0,
"sales": 0.1
},
"urgency": {
"0": 0,
"1": 0,
"2": 0.2,
"3": 0.8
},
"outage": {
"true": 0.8,
"false": 0.2
}
} Decisions: Refund wrong plan 100% right
{
"department": {
"billing": 0.9,
"technical": 0.05,
"account": 0.025,
"shipping": 0.025,
"sales": 0.0
},
"refund": {
"true": 0.9,
"false": 0.1
},
"tone": {
"frustrated": 0.0,
"calm": 1.0
}
} Decisions: Moderation doxxing 100% right
{
"policy": {
"none": 0.2,
"harassment": 0.7,
"hate": 0.05,
"spam": 0.05,
"self_harm": 0
},
"personal_info": {
"true": 0.9,
"false": 0.1
}
} Decisions: Route calendar 100% right
{
"tool": {
"web_search": 0,
"calculator": 0,
"calendar": 0.9,
"email": 0,
"none": 0.1
},
"confirm": {
"true": 0.9,
"false": 0.1
}
} Decisions: Doc invoice missing due 100% right
{
"doc_type": {
"invoice": 0.95,
"resume": 0.0,
"contract": 0.0,
"bank_statement": 0.0,
"other": 0.05
},
"missing_due_date": {
"true": 0.8,
"false": 0.2
}
} Decisions: Phishing paypal 100% right
{"phishing": {"true": 0.95, "false": 0.05}, "risk": {"0": 0.0, "1": 0.0, "2": 0.05, "3": 0.95}} Decisions: Pii ssn email 100% right
```
{
"data_kind": {
"none": 0.0,
"contact": 0.2,
"financial": 0.2,
"government_id": 0.6,
"health": 0.0
},
"sensitive": {
"true": 0.6,
"false": 0.4
}
}
``` Decisions: Review mixed 100% right
{
"sentiment": {
"positive": 0.1,
"neutral": 0.2,
"negative": 0.7
},
"defect": {
"true": 1,
"false": 0
},
"recommend": {
"true": 0,
"false": 1
}
} Documents: Saas escalator (hard) 20% right
{
"year2_price_per_seat_month": 47.70,
"year3_price_per_seat_month": 47.70,
"year1_invoice": 64800.00,
"year2_invoice": 68614.40,
"addon_months_billed": 7,
"addon_invoice": 33215.60,
"year3_invoice": 68614.40,
"year3_discount_percent": 15,
"total_contract_value": 204730.40,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 75% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{ "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
{ "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
{ "date": "2025-02-25", "category": "meals", "amount_usd": 229.69 },
{ "date": "2025-02-26", "category": "lodging", "amount_usd": 467.16 },
{ "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.68 }
],
"rejected_item_count": 1,
"per_diem_days": 2,
"per_diem_usd": 130.00,
"total_reimbursable_usd": 2075.01,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 67% right
```
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2239.00,
"late_fee_from_2025_06_01": 111.95,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2309.00,
"move_in_payment": 4650.00
}
``` Documents: Ticket SLA 82% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-15T17:00",
"sla_due_utc": "2025-09-15T23:00:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 33% right
{
"q3_total_usd": 15446,
"q2_total_usd": 14364,
"q2_central_originally_reported_usd": 3047,
"q2_to_q3_change_pct": 7.4,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731,
"west_excluding_mountain_q3_usd": 4201
} Size: 71B parameters. First tested OCT 11.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.