- Llama 3.3 8B Instruct (Q4, Mac) is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 23 out of 100, #47 of 56.
- It solved 4 of 30 coding jobs and scored 33 on reading documents. On our hardest tasks it scored 10.
- Runs on an 8 GB graphics card or a Mac with 16 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.3 8B Instruct (Q4, Mac) got 4 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.3 8B Instruct (Q4, Mac) scored 33; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 13 | 14 | 13 |
| Reading documents | 33 | 51 | 28 |
| Decisions | 59 | 73 | 55 |
On the 18 hardest tasks (included in the scores above) it scored 10. This number separates the top models.
This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.
We tested the Q4 download, the version most people run, with Ollama on our Mac mini (M4, 16 GB).
Can your computer run it?
Llama 3.3 8B Instruct (Q4, Mac) comes in one standard size, which needs about 6 GB of memory.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | Yes | Q4 |
| Graphics card, 12 GB | Yes | Q4 |
| Graphics card, 16 GB | Yes | Q4 |
| Graphics card, 24 GB | Yes | Q4 |
| Graphics card, 32 GB | Yes | Q4 |
| Mac, 16 GB | Yes | Q4 |
| Mac, 24 GB | Yes | Q4 |
| Mac, 32 GB | Yes | Q4 |
| Mac, 48 GB | Yes | Q4 |
| Mac, 64 GB | Yes | Q4 |
| Mac, 96 GB | Yes | Q4 |
| Mac, 128 GB | Yes | Q4 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
The easiest way is Ollama. Install it, then run:
ollama run hf.co/bartowski/allura-forge_Llama-3.3-8B-Instruct-GGUF:Q4_K_M Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What Llama 3.3 8B Instruct (Q4, Mac) wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
```python
def parse_ini(text: str) -> dict[str, dict[str, str]]:
"""
Parse an INI file and return a dictionary of sections where each section is a dictionary of key-value pairs.
"""
sections = {}
current_section = None
default_section = {}
for line in text.splitlines():
line = line.strip()
if not line or line.startswith(('#', ';')):
continue
if line.startswith('[') and line.endswith(']'):
section_name = line[1:-1].strip()
if section_name == 'DEFAULT':
current_section = None
else:
current_section = section_name
sections[section_name] = default_section.copy()
else:
if current_section is None:
raise ValueError("Key line before any section header")
# Split the line into key and value
split_index = min(line.find('='), line.find(':'))
if split_index == -1:
raise ValueError("Malformed line")
key, value = line[:split_index].strip().lower(), line[split_index + 1:].strip()
# Check for continuation
if line.startswith(' '):
if current_section not in sections or key not in sections[current_section]:
raise ValueError("Continuation line without previous key")
value = sections[current_section][key] + '\n' + value
sections[current_section][key] = value
else:
# Replace interpolation
def replace_interpolation(value: str) -> str:
if '${' in value:
if value.startswith('${') and value.endswith('}'):
key = value[2:-1]
if ':' in key:
section, key = key.split(':')
if section not in sections:
raise KeyError(f"Section '{section}' not found")
value = value.replace('${' + key + '}', sections[section][key])
else:
if current_section is None:
raise KeyError(f"Key '{key}' not found in DEFAULT section")
value = value.replace('${' + key + '}', sections[current_section][key])
else:
raise ValueError("Invalid interpolation")
return value
sections[current_section][key] = replace_interpolation(value)
return sections
``` Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
class TTLCache:
def __init__(self, capacity: int, ttl: float, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = OrderedDict()
self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}
def put(self, key, value):
current_time = self.clock()
if key in self.cache:
self.cache[key] = (current_time, value)
else:
if len(self.cache) >= self.capacity:
self._remove_expired()
if len(self.cache) >= self.capacity:
self._evict_lru()
self.cache[key] = (current_time, value)
def get(self, key, default=None):
current_time = self.clock()
if key in self.cache:
value, expiry = self.cache[key]
if current_time >= expiry:
self.stats["expirations"] += 1
del self.cache[key]
return default
else:
self.cache[key] = (current_time, value)
self.stats["hits"] += 1
return value
else:
self.stats["misses"] += 1
return default
def __len__(self):
return sum(1 for key, (time, _) in self.cache.items() if self.clock() < time + self.ttl)
def stats(self):
return self.stats.copy()
def _remove_expired(self):
current_time = self.clock()
self.cache = OrderedDict((key, value) for key, (time, value) in self.cache.items() if current_time < time + self.ttl)
def _evict_lru(self):
self.cache.popitem(last=False)
self.stats["evictions"] += 1
``` Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
"""
Converts a duration string into a total number of seconds.
Args:
s (str): The duration string.
Returns:
int: The total number of seconds.
Raises:
ValueError: If the input string is invalid.
"""
s = s.strip()
if not s:
raise ValueError("Input string is empty")
parts = s.split()
seconds = 0
unit_order = ['h', 'm', 's']
unit_index = 0
for part in parts:
if not part:
continue
num = ''
for char in part:
if char.isdigit():
num += char
else:
break
if not num:
raise ValueError("Number without a unit")
num = int(num)
unit = part.lower()[-1]
if unit not in unit_order or unit_index > unit_order.index(unit):
raise ValueError("Unknown unit or units out of order")
unit_index += 1
if unit == 'h':
seconds += num * 3600
elif unit == 'm':
seconds += num * 60
elif unit == 's':
seconds += num
return seconds
``` Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
nums = sorted(set(nums))
result = []
start = nums[0]
count = 1
for i in range(1, len(nums)):
if nums[i] == nums[i-1] + 1:
count += 1
else:
if count >= 3:
result.append(f"{start}..{nums[i-1]}")
else:
result.append(",".join(map(str, range(start, nums[i-1] + 1))))
start = nums[i]
count = 1
if count >= 3:
result.append(f"{start}..{nums[-1]}")
else:
result.append(",".join(map(str, range(start, nums[-1] + 1))))
return ",".join(result)
``` Coding: Top customers Didn't work
```python
from typing import List, Dict, Tuple
from collections import defaultdict
import re
def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
# Create a dictionary to store the sum of amounts for each customer
customer_amounts = defaultdict(float)
# Iterate over each row
for row in rows:
customer = row["customer"].strip().lower()
amount = row["amount"]
# Check if the customer name is not empty and the amount can be parsed
if customer and re.match(r'[-+]?\d{1,3}(,\d{3})*(\.\d+)?', amount):
# Parse the amount
amount = float(re.sub(r'[,\$]', '', amount))
# If the amount is negative, it's a refund, so add it to the total
if amount < 0:
customer_amounts[customer] += amount
else:
customer_amounts[customer] += amount
# Convert the dictionary to a list of tuples and sort it
sorted_customers = sorted(customer_amounts.items(), key=lambda x: (-x[1], x[0]))
# Return the top n customers
return sorted_customers[:n]
``` Coding: Slugify Didn't work
```python
import unicodedata
import re
def slugify(title: str, max_len: int = 50) -> str:
# Transliterate accented characters to ASCII and drop any other non-ASCII characters
title = ''.join(c for c in unicodedata.normalize('NFD', title) if unicodedata.category(c) != 'Mn')
# Lowercase
title = title.lower()
# Replace every run of characters other than a-z and 0-9 with a single "-"
title = re.sub(r'[^a-z0-9]+', '-', title)
# Strip leading/trailing "-"
title = title.strip('-')
# If the slug is longer than max_len, shorten it without cutting a word
if len(title) > max_len:
words = title.split('-')
shortened_words = []
length = 0
for word in words:
if length + len(word) + 1 > max_len:
break
shortened_words.append(word)
length += len(word) + 1
title = '-'.join(shortened_words)
# If even the first word is longer than max_len, hard-cut it to max_len characters
if len(title) > max_len:
title = title[:max_len]
# The result never ends with "-"
title = title.rstrip('-')
return title
``` Coding: Token bucket Didn't work
```python
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.clock = clock
self.tokens = capacity
self.last_refill = self.clock()
def allow(self, cost: float = 1) -> bool:
now = self.clock()
elapsed = now - self.last_refill
self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
self.last_refill = now
if self.tokens >= cost:
self.tokens -= cost
return True
return False
@property
def tokens(self) -> float:
now = self.clock()
elapsed = now - self.last_refill
self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
self.last_refill = now
return self.tokens
```<|fim_end|> Decisions: Refund window (hard) 100% right
{
"outcome": {
"full_refund": 0.0,
"partial_refund": 0.0,
"store_credit": 1.0,
"nothing": 0.0
},
"defective": {
"true": 0.0,
"false": 1.0
}
} Decisions: Incident matrix (hard) 0% right
{
"0": 0,
"1": 0.2,
"2": 0.3,
"3": 0.5
}
{
"true": 0.8,
"false": 0.2
} Decisions: Tool followup (hard) 100% right
{
"tool": {
"web_search": 0,
"calculator": 0,
"calendar": 1,
"email": 0,
"none": 0
},
"confirm": {
"true": 1,
"false": 0
}
} Decisions: Legit security alert (hard) 50% right
{
"phishing": {
"true": 0,
"false": 1
},
"action_needed": {
"true": 0.9,
"false": 0.1
}
}<!DOCTYPE html|<|fim_end|> Decisions: Meeting slot (hard) 0% right
{
"A": 0.0,
"B": 0.0,
"C": 0.0,
"D": 1.0
}
{
"true": 0.0,
"false": 1.0
} Decisions: Review mixed (hard) 100% right
{
"hardware": {"true": 1, "false": 0},
"support": {"true": 1, "false": 0}
} Decisions: Support checkout down 67% right
{
"department": {
"billing": 0.4,
"technical": 0.6,
"account": 0.0,
"shipping": 0.0,
"sales": 0.0
},
"urgency": {
"0": 0.0,
"1": 0.0,
"2": 0.7,
"3": 0.3
},
"outage": {
"true": 1.0,
"false": 0.0
}
} Decisions: Refund wrong plan 100% right
{
"department": {"billing": 0.8, "technical": 0.1, "account": 0.05, "shipping": 0, "sales": 0.05},
"refund": {"true": 0.9, "false": 0.1},
"tone": {"frustrated": 0.4, "calm": 0.6}
} Decisions: Moderation doxxing 100% right
{
"policy": {
"none": 0.05,
"harassment": 0.8,
"hate": 0.05,
"spam": 0.05,
"self_harm": 0.05
},
"personal_info": {
"true": 0.9,
"false": 0.1
}
} Decisions: Route calendar 100% right
{
"tool": {
"web_search": 0,
"calculator": 0,
"calendar": 0.8,
"email": 0.1,
"none": 0.1
},
"confirm": {
"true": 1,
"false": 0
}
} Decisions: Doc invoice missing due 0% right
{
"invoice": 0.9,
"resume": 0.05,
"contract": 0.01,
"bank_statement": 0.01,
"other": 0.03,
"true": 0.6,
"false": 0.4
} Decisions: Phishing paypal 100% right
{
"phishing": {"true": 0.9, "false": 0.1},
"risk": {"0": 0.0, "1": 0.1, "2": 0.3, "3": 0.6}
} Decisions: Pii ssn email 100% right
{
"data_kind": {
"none": 0,
"contact": 0,
"financial": 0,
"government_id": 1,
"health": 0
},
"sensitive": {
"true": 1,
"false": 0
}
} Decisions: Review mixed 100% right
{
"sentiment": {"positive": 0.0, "neutral": 0.0, "negative": 1.0},
"defect": {"true": 1.0, "false": 0.0},
"recommend": {"true": 0.0, "false": 1.0}
} Documents: Saas escalator (hard) 30% right
{
"year2_price_per_seat_month": 46.31,
"year3_price_per_seat_month": 45.95,
"year1_invoice": 5400.00,
"year2_invoice": 5547.60,
"addon_months_billed": 6,
"addon_invoice": 931.2,
"year3_invoice": 5683.20,
"year3_discount_percent": 15.00,
"total_contract_value": 17631.80,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 71% right
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{ "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
{ "date": "2025-02-24", "category": "ground_transport", "amount_usd": 39.12 },
{ "date": "2025-02-25", "category": "meals", "amount_usd": 212.40 },
{ "date": "2025-02-26", "category": "lodging", "amount_usd": 466.56 },
{ "date": "2025-02-27", "category": "ground_transport", "amount_usd": 41.50 }
],
"rejected_item_count": 2,
"per_diem_days": 2,
"per_diem_usd": 65,
"total_reimbursable_usd": 1813.18,
"approver_email": "priya.raman@corvane.com"
} Documents: Lease amendment 67% right
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2243.00,
"late_fee_from_2025_06_01": 110.15,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 3278.00,
"move_in_payment": 4300.00
} Documents: Ticket SLA 75% right
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address", "invoice_pdf"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-12T17:30",
"sla_due_utc": "2025-09-12T21:30:00Z",
"reissued_invoice": "INV-2025-0812"
} Documents: Sales footnotes 11% right
{
"q3_total_usd": 49802,
"q2_total_usd": 48671,
"q2_central_originally_reported_usd": 3047,
"q2_to_q3_change_pct": 2.6,
"top_region_q3": "West",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["Central"],
"international_q3_organic_usd": 1731,
"west_excluding_mountain_q3_usd": 42900
} Size: 8B parameters. First tested OCT 11.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.