Review · updated OCT 11

Llama 3.3 70B Instruct review: not one we'd recommend right now

It scored 46 out of 100, #33 of 56. It solved 14 of 30 coding jobs and scored 45 on reading documents. Runs on a Mac with 64 GB.

The short version
  • Llama 3.3 70B Instruct is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 46 out of 100, #33 of 56.
  • It solved 14 of 30 coding jobs and scored 45 on reading documents. On our hardest tasks it scored 27.
  • Runs on a Mac with 64 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.3 70B Instruct got 14 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.3 70B Instruct scored 45; the best model scored 100.

TestScorePublic questionsSecret questions
Coding474348
Reading documents456041
Decisions868985

On the 18 hardest tasks (included in the scores above) it scored 27. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Llama 3.3 70B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBYesQ4
Mac, 96 GBYesQ6
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Llama 3.3 70B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Llama 3.3 70B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    """
    Parse an INI file and return a dictionary of sections and their key-value pairs.

    :param text: The INI file content as a string
    :return: A dictionary of sections and their key-value pairs
    """
    sections: Dict[str, Dict[str, str]] = {}
    default_section: Dict[str, str] = {}
    current_section: Dict[str, str] = default_section
    current_key: str = ""
    current_value: str = ""

    for line in text.splitlines():
        # Remove leading and trailing whitespace
        line = line.strip()

        # Ignore empty lines and comments
        if not line or line[0] in [";", "#"]:
            continue

        # Check for section headers
        match = re.match(r"\s*\[(.*?)\]\s*", line)
        if match:
            section_name = match.group(1)
            if section_name == "DEFAULT":
                current_section = default_section
            else:
                current_section = sections.setdefault(section_name, {})
            current_key = ""
            current_value = ""
            continue

        # Check for key-value pairs
        match = re.match(r"\s*(\S+)\s*([:=])\s*(.*)", line)
        if match:
            key = match.group(1).strip().lower()
            value = match.group(3).strip()
            current_key = key
            current_value = value
            current_section[key] = value
            continue

        # Check for continuation lines
        if line and current_key:
            current_value += "\n" + line.strip()
            current_section[current_key] = current_value
            continue

        # Raise an error for malformed lines
        raise ValueError("Malformed line: " + line)

    # Interpolate values
    def interpolate(value: str, section: Dict[str, str]) -> str:
        def replace(match: re.Match) -> str:
            ref = match.group(1)
            if ":" in ref:
                section_name, key = ref.split(":", 1)
                if section_name not in sections:
                    raise KeyError(f"Section '{section_name}' not found")
                section_ref = sections[section_name]
                if key not in section_ref:
                    raise KeyError(f"Key '{key}' not found in section '{section_name}'")
                return section_ref[key]
            else:
                if ref not in section:
                    raise KeyError(f"Key '{ref}' not found")
                return section[ref]

        while True:
            new_value = re.sub(r"\${(.*?)}", replace, value)
            if new_value == value:
                break
            value = new_value

        return value

    for section_name, section in sections.items():
        for key, value in section.items():
            # Check for reference cycles
            try:
                sections[section_name][key] = interpolate(value, {**default_section, **section})
            except RecursionError:
                raise ValueError("Reference cycle detected")

    return {section_name: section for section_name, section in sections.items() if section_name != "DEFAULT"}
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from time import monotonic

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock=monotonic):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def _remove_expired(self):
        current_time = self.clock()
        expired_keys = [key for key, (_, expiry) in self.cache.items() if current_time >= expiry]
        for key in expired_keys:
            del self.cache[key]
            self.stats["expirations"] += 1

    def put(self, key, value):
        if self.capacity == 0:
            return
        self._remove_expired()
        if key in self.cache:
            del self.cache[key]
        elif len(self.cache) >= self.capacity:
            self.cache.popitem(last=False)
            self.stats["evictions"] += 1
        expiry = self.clock() + self.ttl
        self.cache[key] = (value, expiry)
        self.cache.move_to_end(key)

    def get(self, key, default=None):
        if self.capacity == 0:
            self.stats["misses"] += 1
            return default
        self._remove_expired()
        if key not in self.cache:
            self.stats["misses"] += 1
            return default
        value, expiry = self.cache[key]
        current_time = self.clock()
        if current_time >= expiry:
            del self.cache[key]
            self.stats["expirations"] += 1
            self.stats["misses"] += 1
            return default
        self.cache.move_to_end(key)
        self.stats["hits"] += 1
        return value

    def __len__(self):
        self._remove_expired()
        return len(self.cache)

    def stats(self):
        return {
            "hits": self.stats["hits"],
            "misses": self.stats["misses"],
            "evictions": self.stats["evictions"],
            "expirations": self.stats["expirations"],
        }
```
Coding: Parse duration Didn't work
```python
import re
from typing import Pattern

def parse_duration(s: str) -> int:
    """
    Converts a duration string into a total number of seconds.

    Args:
    s (str): The duration string.

    Returns:
    int: The total number of seconds.

    Raises:
    ValueError: If the string is empty, contains invalid units, or has units out of order.
    """
    
    # Remove leading and trailing whitespace
    s = s.strip()
    
    # Check for empty string
    if not s:
        raise ValueError("Empty string")
    
    # Regular expression pattern to match duration parts
    pattern: Pattern[str] = re.compile(r"(\d+)([hms])", re.IGNORECASE)
    
    # Initialize total seconds and unit order check
    total_seconds = 0
    seen_units = set()
    
    # Split the string into parts
    parts = re.split(r"\s+", s)
    
    # Iterate over each part
    for part in parts:
        # Check if the part matches the pattern
        match = pattern.fullmatch(part)
        
        # If the part does not match the pattern, raise an error
        if not match:
            raise ValueError(f"Invalid part: {part}")
        
        # Get the value and unit from the match
        value = int(match.group(1))
        unit = match.group(2).lower()
        
        # Check for repeated units
        if unit in seen_units:
            raise ValueError(f"Repeated unit: {unit}")
        
        # Check for units out of order
        if unit == 'm' and 'h' not in seen_units:
            pass
        elif unit == 's' and ('h' not in seen_units or 'm' not in seen_units):
            pass
        elif unit == 'h' and ('m' in seen_units or 's' in seen_units):
            raise ValueError("Units out of order")
        elif unit == 'm' and 's' in seen_units:
            raise ValueError("Units out of order")
        
        # Add the unit to the seen units
        seen_units.add(unit)
        
        # Convert the value to seconds and add to the total
        if unit == 'h':
            total_seconds += value * 3600
        elif unit == 'm':
            total_seconds += value * 60
        elif unit == 's':
            total_seconds += value
    
    return total_seconds
```
Coding: Summarize ranges Worked
```python
def summarize_ranges(nums: list[int]) -> str:
    """
    This function takes a list of integers, sorts and removes duplicates, 
    then collapses runs of 3 or more consecutive integers into "a..b" 
    and numbers in shorter runs (1 or 2 numbers) are listed individually.
    
    Args:
    nums (list[int]): A list of integers.
    
    Returns:
    str: A string representing the summarized ranges.
    """
    
    # If the list is empty, return an empty string
    if not nums:
        return ""

    # Sort the numbers and remove duplicates by converting to a set
    nums = sorted(set(nums))
    
    # Initialize variables to keep track of the current run
    current_run = [nums[0]]
    result = []
    
    # Iterate over the sorted numbers
    for num in nums[1:]:
        # If the current number is a continuation of the current run, add it to the run
        if num == current_run[-1] + 1:
            current_run.append(num)
        # If the current number is not a continuation of the current run, 
        # summarize the current run and start a new one
        else:
            # If the current run has 3 or more numbers, summarize it as "a..b"
            if len(current_run) >= 3:
                result.append(f"{current_run[0]}..{current_run[-1]}")
            # If the current run has 1 or 2 numbers, list them individually
            else:
                result.extend(map(str, current_run))
            # Start a new run with the current number
            current_run = [num]
    
    # Summarize the last run
    if len(current_run) >= 3:
        result.append(f"{current_run[0]}..{current_run[-1]}")
    else:
        result.extend(map(str, current_run))
    
    # Join everything with "," and return the result
    return ",".join(result)
```
Coding: Top customers Worked
```python
from collections import defaultdict
from decimal import Decimal, InvalidOperation
import re

def top_customers(rows, n):
    """
    Returns the top n customers with the highest total amount.

    :param rows: A list of dictionaries containing customer information.
    :param n: The number of top customers to return.
    :return: A list of tuples containing the customer name and total amount.
    """

    # Create a dictionary to store the total amount for each customer
    customer_amounts = defaultdict(Decimal)

    # Iterate over each row
    for row in rows:
        # Get the customer name and amount
        customer = row.get("customer", "").strip()
        amount = row.get("amount", "")

        # Skip rows with empty customer name or invalid amount
        if not customer or not amount:
            continue

        # Parse the amount
        try:
            # Remove dollar signs and commas, and convert to Decimal
            amount = Decimal(re.sub(r"[^\d.-]", "", amount))
        except InvalidOperation:
            # Skip rows with invalid amount
            continue

        # Update the customer's total amount
        customer_amounts[customer.lower()] += amount

    # Round the total amounts to 2 decimals and store the original customer name
    customer_amounts = {
        next(key for key in rows if key.get("customer", "").strip().lower() == customer).get("customer", "").strip(): round(amount, 2)
        for customer, amount in customer_amounts.items()
    }

    # Sort the customers by total amount in descending order, then by name in ascending order
    sorted_customers = sorted(customer_amounts.items(), key=lambda x: (-x[1], x[0]))

    # Return the top n customers
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    # Transliterate accented characters to ASCII and drop any other non-ASCII characters
    title = ''.join(c for c in unicodedata.normalize('NFD', title) if unicodedata.category(c) != 'Mn' and c.isascii())
    
    # Lowercase
    title = title.lower()
    
    # Replace every run of characters other than a-z and 0-9 with a single "-", and strip leading/trailing "-"
    title = re.sub(r'[^a-z0-9]+', '-', title).strip('-')
    
    # If the slug is longer than `max_len`, shorten it without cutting a word
    if len(title) > max_len:
        words = title.split('-')
        slug = ''
        for word in words:
            if len(slug) + len(word) + 1 <= max_len:
                slug += word + '-'
            else:
                break
        slug = slug.rstrip('-')
        if not slug:  # If even the first word is longer than max_len, hard-cut it to max_len characters
            slug = title[:max_len]
    
    else:
        slug = title
    
    return slug
```
Coding: Token bucket Worked
```python
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._last_refill = clock()
        self._tokens = capacity

    def _refill(self) -> None:
        now = self._clock()
        elapsed = now - self._last_refill
        self._tokens = min(self._capacity, self._tokens + elapsed * self._refill_per_sec)
        self._last_refill = now

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens
```
Decisions: Refund window (hard) 50% right
```json
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.8,
    "store_credit": 0.2,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Incident matrix (hard) 100% right
```json
{
  "severity": {
    "0": 0.0,
    "1": 0.7,
    "2": 0.3,
    "3": 0.0
  },
  "page": {
    "true": 0.3,
    "false": 0.7
  }
}
```
Decisions: Tool followup (hard) 100% right
{"tool": {"web_search": 0.0, "calculator": 0.0, "calendar": 1.0, "email": 0.0, "none": 0.0}, "confirm": {"true": 1.0, "false": 0.0}}
Decisions: Legit security alert (hard) 100% right
```json
{
  "phishing": {
    "true": 0.01,
    "false": 0.99
  },
  "action_needed": {
    "true": 0.05,
    "false": 0.95
  }
}
```
Decisions: Meeting slot (hard) 0% right
```
{
  "slot": {
    "A": 0.0,
    "B": 0.0,
    "C": 0.0,
    "D": 1.0
  },
  "raj_last": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Review mixed (hard) 100% right
{"hardware": {"true": 0.9, "false": 0.1}, "support": {"true": 0.8, "false": 0.2}}
Decisions: Support checkout down 100% right
```json
{
  "department": {
    "billing": 0.2,
    "technical": 0.7,
    "account": 0.05,
    "shipping": 0.0,
    "sales": 0.05
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.2,
    "3": 0.8
  },
  "outage": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Refund wrong plan 100% right
```json
{
  "department": {
    "billing": 0.9,
    "technical": 0.05,
    "account": 0.03,
    "shipping": 0.01,
    "sales": 0.01
  },
  "refund": {
    "true": 0.95,
    "false": 0.05
  },
  "tone": {
    "frustrated": 0.1,
    "calm": 0.9
  }
}
```
Decisions: Moderation doxxing 100% right
```
{
  "policy": {
    "none": 0.0,
    "harassment": 0.8,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.2
  },
  "personal_info": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Route calendar 100% right
```
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 0.8,
    "email": 0.1,
    "none": 0.1
  },
  "confirm": {
    "true": 0.9,
    "false": 0.1
  }
}
```
Decisions: Doc invoice missing due 100% right
```
{
  "doc_type": {
    "invoice": 0.99,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.01
  },
  "missing_due_date": {
    "true": 0.99,
    "false": 0.01
  }
}
```
Decisions: Phishing paypal 100% right
```json
{
  "phishing": {
    "true": 0.9,
    "false": 0.1
  },
  "risk": {
    "0": 0.05,
    "1": 0.05,
    "2": 0.1,
    "3": 0.8
  }
}
```
Decisions: Pii ssn email 100% right
```
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.2,
    "financial": 0.3,
    "government_id": 0.5,
    "health": 0.0
  },
  "sensitive": {
    "true": 0.8,
    "false": 0.2
  }
}
```
Decisions: Review mixed 100% right
{"sentiment": {"positive": 0.0, "neutral": 0.0, "negative": 1.0}, "defect": {"true": 1.0, "false": 0.0}, "recommend": {"true": 0.0, "false": 1.0}}
Documents: Saas escalator (hard) 20% right
```
{
  "year2_price_per_seat_month": 47.53,
  "year3_price_per_seat_month": 49.98,
  "year1_invoice": 4860.00,
  "year2_invoice": 5719.20,
  "addon_months_billed": 5,
  "addon_invoice": 2145.60,
  "year3_invoice": 7188.80,
  "year3_discount_percent": 15,
  "total_contract_value": 17923.60,
  "contract_end_date": "2027-02-28"
}
```
Documents: Expense thread 75% right
```
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    { "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
    { "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
    { "date": "2025-02-25", "category": "meals", "amount_usd": 229.33 },
    { "date": "2025-02-26", "category": "lodging", "amount_usd": 466.56 },
    { "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.68 }
  ],
  "rejected_item_count": 2,
  "per_diem_days": 2,
  "per_diem_usd": 130.00,
  "total_reimbursable_usd": 2064.05,
  "approver_email": "priya.raman@corvane.com"
}
```
Documents: Lease amendment 92% right
```
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 4650.00
}
```
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-16T13:30",
  "sla_due_utc": "2025-09-16T18:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 33% right
{
  "q3_total_usd": 14946,
  "q2_total_usd": 14414,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 3.7,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 4201
}

Size: 71B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.