Review · updated OCT 11

Llama 3.3 8B Instruct (Q4, Mac) review: not one we'd recommend right now

It scored 23 out of 100, #47 of 56. It solved 4 of 30 coding jobs and scored 33 on reading documents. Runs on an 8 GB graphics card or a Mac with 16 GB.

The short version
  • Llama 3.3 8B Instruct (Q4, Mac) is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 23 out of 100, #47 of 56.
  • It solved 4 of 30 coding jobs and scored 33 on reading documents. On our hardest tasks it scored 10.
  • Runs on an 8 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.3 8B Instruct (Q4, Mac) got 4 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.3 8B Instruct (Q4, Mac) scored 33; the best model scored 100.

TestScorePublic questionsSecret questions
Coding131413
Reading documents335128
Decisions597355

On the 18 hardest tasks (included in the scores above) it scored 10. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the Q4 download, the version most people run, with Ollama on our Mac mini (M4, 16 GB).

Can your computer run it?

Llama 3.3 8B Instruct (Q4, Mac) comes in one standard size, which needs about 6 GB of memory.

Your computerRuns it?Version to download
Graphics card, 8 GBYesQ4
Graphics card, 12 GBYesQ4
Graphics card, 16 GBYesQ4
Graphics card, 24 GBYesQ4
Graphics card, 32 GBYesQ4
Mac, 16 GBYesQ4
Mac, 24 GBYesQ4
Mac, 32 GBYesQ4
Mac, 48 GBYesQ4
Mac, 64 GBYesQ4
Mac, 96 GBYesQ4
Mac, 128 GBYesQ4

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run hf.co/bartowski/allura-forge_Llama-3.3-8B-Instruct-GGUF:Q4_K_M

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Llama 3.3 8B Instruct (Q4, Mac) wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
def parse_ini(text: str) -> dict[str, dict[str, str]]:
    """
    Parse an INI file and return a dictionary of sections where each section is a dictionary of key-value pairs.
    """
    sections = {}
    current_section = None
    default_section = {}

    for line in text.splitlines():
        line = line.strip()
        if not line or line.startswith(('#', ';')):
            continue

        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if section_name == 'DEFAULT':
                current_section = None
            else:
                current_section = section_name
                sections[section_name] = default_section.copy()
        else:
            if current_section is None:
                raise ValueError("Key line before any section header")

            # Split the line into key and value
            split_index = min(line.find('='), line.find(':'))
            if split_index == -1:
                raise ValueError("Malformed line")
            key, value = line[:split_index].strip().lower(), line[split_index + 1:].strip()

            # Check for continuation
            if line.startswith(' '):
                if current_section not in sections or key not in sections[current_section]:
                    raise ValueError("Continuation line without previous key")
                value = sections[current_section][key] + '\n' + value
                sections[current_section][key] = value
            else:
                # Replace interpolation
                def replace_interpolation(value: str) -> str:
                    if '${' in value:
                        if value.startswith('${') and value.endswith('}'):
                            key = value[2:-1]
                            if ':' in key:
                                section, key = key.split(':')
                                if section not in sections:
                                    raise KeyError(f"Section '{section}' not found")
                                value = value.replace('${' + key + '}', sections[section][key])
                            else:
                                if current_section is None:
                                    raise KeyError(f"Key '{key}' not found in DEFAULT section")
                                value = value.replace('${' + key + '}', sections[current_section][key])
                        else:
                            raise ValueError("Invalid interpolation")
                    return value

                sections[current_section][key] = replace_interpolation(value)

    return sections
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def put(self, key, value):
        current_time = self.clock()
        if key in self.cache:
            self.cache[key] = (current_time, value)
        else:
            if len(self.cache) >= self.capacity:
                self._remove_expired()
                if len(self.cache) >= self.capacity:
                    self._evict_lru()
            self.cache[key] = (current_time, value)

    def get(self, key, default=None):
        current_time = self.clock()
        if key in self.cache:
            value, expiry = self.cache[key]
            if current_time >= expiry:
                self.stats["expirations"] += 1
                del self.cache[key]
                return default
            else:
                self.cache[key] = (current_time, value)
                self.stats["hits"] += 1
                return value
        else:
            self.stats["misses"] += 1
            return default

    def __len__(self):
        return sum(1 for key, (time, _) in self.cache.items() if self.clock() < time + self.ttl)

    def stats(self):
        return self.stats.copy()

    def _remove_expired(self):
        current_time = self.clock()
        self.cache = OrderedDict((key, value) for key, (time, value) in self.cache.items() if current_time < time + self.ttl)

    def _evict_lru(self):
        self.cache.popitem(last=False)
        self.stats["evictions"] += 1
```
Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
    """
    Converts a duration string into a total number of seconds.

    Args:
    s (str): The duration string.

    Returns:
    int: The total number of seconds.

    Raises:
    ValueError: If the input string is invalid.
    """
    s = s.strip()
    if not s:
        raise ValueError("Input string is empty")

    parts = s.split()
    seconds = 0
    unit_order = ['h', 'm', 's']
    unit_index = 0

    for part in parts:
        if not part:
            continue

        num = ''
        for char in part:
            if char.isdigit():
                num += char
            else:
                break
        if not num:
            raise ValueError("Number without a unit")

        num = int(num)
        unit = part.lower()[-1]
        if unit not in unit_order or unit_index > unit_order.index(unit):
            raise ValueError("Unknown unit or units out of order")
        unit_index += 1

        if unit == 'h':
            seconds += num * 3600
        elif unit == 'm':
            seconds += num * 60
        elif unit == 's':
            seconds += num

    return seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""

    nums = sorted(set(nums))
    result = []
    start = nums[0]
    count = 1

    for i in range(1, len(nums)):
        if nums[i] == nums[i-1] + 1:
            count += 1
        else:
            if count >= 3:
                result.append(f"{start}..{nums[i-1]}")
            else:
                result.append(",".join(map(str, range(start, nums[i-1] + 1))))
            start = nums[i]
            count = 1

    if count >= 3:
        result.append(f"{start}..{nums[-1]}")
    else:
        result.append(",".join(map(str, range(start, nums[-1] + 1))))

    return ",".join(result)
```
Coding: Top customers Didn't work
```python
from typing import List, Dict, Tuple
from collections import defaultdict
import re

def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    # Create a dictionary to store the sum of amounts for each customer
    customer_amounts = defaultdict(float)
    
    # Iterate over each row
    for row in rows:
        customer = row["customer"].strip().lower()
        amount = row["amount"]
        
        # Check if the customer name is not empty and the amount can be parsed
        if customer and re.match(r'[-+]?\d{1,3}(,\d{3})*(\.\d+)?', amount):
            # Parse the amount
            amount = float(re.sub(r'[,\$]', '', amount))
            
            # If the amount is negative, it's a refund, so add it to the total
            if amount < 0:
                customer_amounts[customer] += amount
            else:
                customer_amounts[customer] += amount
    
    # Convert the dictionary to a list of tuples and sort it
    sorted_customers = sorted(customer_amounts.items(), key=lambda x: (-x[1], x[0]))
    
    # Return the top n customers
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    # Transliterate accented characters to ASCII and drop any other non-ASCII characters
    title = ''.join(c for c in unicodedata.normalize('NFD', title) if unicodedata.category(c) != 'Mn')
    
    # Lowercase
    title = title.lower()
    
    # Replace every run of characters other than a-z and 0-9 with a single "-"
    title = re.sub(r'[^a-z0-9]+', '-', title)
    
    # Strip leading/trailing "-"
    title = title.strip('-')
    
    # If the slug is longer than max_len, shorten it without cutting a word
    if len(title) > max_len:
        words = title.split('-')
        shortened_words = []
        length = 0
        for word in words:
            if length + len(word) + 1 > max_len:
                break
            shortened_words.append(word)
            length += len(word) + 1
        title = '-'.join(shortened_words)
        # If even the first word is longer than max_len, hard-cut it to max_len characters
        if len(title) > max_len:
            title = title[:max_len]
    
    # The result never ends with "-"
    title = title.rstrip('-')
    
    return title
```
Coding: Token bucket Didn't work
```python
class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.tokens = capacity
        self.last_refill = self.clock()

    def allow(self, cost: float = 1) -> bool:
        now = self.clock()
        elapsed = now - self.last_refill
        self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
        self.last_refill = now
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        now = self.clock()
        elapsed = now - self.last_refill
        self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
        self.last_refill = now
        return self.tokens
```<|fim_end|>
Decisions: Refund window (hard) 100% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.0,
    "store_credit": 1.0,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
} 
Decisions: Incident matrix (hard) 0% right
{
  "0": 0,
  "1": 0.2,
  "2": 0.3,
  "3": 0.5
}
{
  "true": 0.8,
  "false": 0.2
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "web_search": 0,
    "calculator": 0,
    "calendar": 1,
    "email": 0,
    "none": 0
  },
  "confirm": {
    "true": 1,
    "false": 0
  }
}
Decisions: Legit security alert (hard) 50% right
{
  "phishing": {
    "true": 0,
    "false": 1
  },
  "action_needed": {
    "true": 0.9,
    "false": 0.1
  }
}<!DOCTYPE html|<|fim_end|>
Decisions: Meeting slot (hard) 0% right
{
  "A": 0.0,
  "B": 0.0,
  "C": 0.0,
  "D": 1.0
}
{
  "true": 0.0,
  "false": 1.0
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {"true": 1, "false": 0},
  "support": {"true": 1, "false": 0}
}
Decisions: Support checkout down 67% right
{
  "department": {
    "billing": 0.4,
    "technical": 0.6,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.0
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.7,
    "3": 0.3
  },
  "outage": {
    "true": 1.0,
    "false": 0.0
  }
} 
Decisions: Refund wrong plan 100% right
{
  "department": {"billing": 0.8, "technical": 0.1, "account": 0.05, "shipping": 0, "sales": 0.05},
  "refund": {"true": 0.9, "false": 0.1},
  "tone": {"frustrated": 0.4, "calm": 0.6}
} 
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.05,
    "harassment": 0.8,
    "hate": 0.05,
    "spam": 0.05,
    "self_harm": 0.05
  },
  "personal_info": {
    "true": 0.9,
    "false": 0.1
  }
} 
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0,
    "calculator": 0,
    "calendar": 0.8,
    "email": 0.1,
    "none": 0.1
  },
  "confirm": {
    "true": 1,
    "false": 0
  }
} 
Decisions: Doc invoice missing due 0% right
{
  "invoice": 0.9,
  "resume": 0.05,
  "contract": 0.01,
  "bank_statement": 0.01,
  "other": 0.03,
  "true": 0.6,
  "false": 0.4
} 
Decisions: Phishing paypal 100% right
{
  "phishing": {"true": 0.9, "false": 0.1},
  "risk": {"0": 0.0, "1": 0.1, "2": 0.3, "3": 0.6}
} 
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0,
    "contact": 0,
    "financial": 0,
    "government_id": 1,
    "health": 0
  },
  "sensitive": {
    "true": 1,
    "false": 0
  }
} 
Decisions: Review mixed 100% right
{
  "sentiment": {"positive": 0.0, "neutral": 0.0, "negative": 1.0},
  "defect": {"true": 1.0, "false": 0.0},
  "recommend": {"true": 0.0, "false": 1.0}
}
Documents: Saas escalator (hard) 30% right
{
  "year2_price_per_seat_month": 46.31,
  "year3_price_per_seat_month": 45.95,
  "year1_invoice": 5400.00,
  "year2_invoice": 5547.60,
  "addon_months_billed": 6,
  "addon_invoice": 931.2,
  "year3_invoice": 5683.20,
  "year3_discount_percent": 15.00,
  "total_contract_value": 17631.80,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 71% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    { "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
    { "date": "2025-02-24", "category": "ground_transport", "amount_usd": 39.12 },
    { "date": "2025-02-25", "category": "meals", "amount_usd": 212.40 },
    { "date": "2025-02-26", "category": "lodging", "amount_usd": 466.56 },
    { "date": "2025-02-27", "category": "ground_transport", "amount_usd": 41.50 }
  ],
  "rejected_item_count": 2,
  "per_diem_days": 2,
  "per_diem_usd": 65,
  "total_reimbursable_usd": 1813.18,
  "approver_email": "priya.raman@corvane.com"
} 
Documents: Lease amendment 67% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2243.00,
  "late_fee_from_2025_06_01": 110.15,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 3278.00,
  "move_in_payment": 4300.00
} 
Documents: Ticket SLA 75% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address", "invoice_pdf"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-12T17:30",
  "sla_due_utc": "2025-09-12T21:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 11% right
{
  "q3_total_usd": 49802,
  "q2_total_usd": 48671,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 2.6,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["Central"],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 42900
} 

Size: 8B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.