Review · updated OCT 11

Llama 3.2 3B Instruct review: not one we'd recommend right now

It scored 6 out of 100, #55 of 56. It solved 0 of 30 coding jobs and scored 13 on reading documents. Runs on an 8 GB graphics card or a Mac with 16 GB.

The short version
  • Llama 3.2 3B Instruct is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 6 out of 100, #55 of 56.
  • It solved 0 of 30 coding jobs and scored 13 on reading documents. On our hardest tasks it scored 2.
  • Runs on an 8 GB graphics card or a Mac with 16 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.2 3B Instruct got 0 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.2 3B Instruct scored 13; the best model scored 100.

TestScorePublic questionsSecret questions
Coding000
Reading documents132410
Decisions444245

On the 18 hardest tasks (included in the scores above) it scored 2. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Llama 3.2 3B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBYesQ8
Graphics card, 12 GBYesQ8
Graphics card, 16 GBYesQ8
Graphics card, 24 GBYesQ8
Graphics card, 32 GBYesQ8
Mac, 16 GBYesQ8
Mac, 24 GBYesQ8
Mac, 32 GBYesQ8
Mac, 48 GBYesQ8
Mac, 64 GBYesQ8
Mac, 96 GBYesQ8
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

The easiest way is Ollama. Install it, then run:

ollama run llama3.2:3b

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Llama 3.2 3B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    """
    Parse an INI file.

    Args:
    text (str): The text of the INI file.

    Returns:
    Dict[str, Dict[str, str]]: A dictionary of sections, where each section is another dictionary of key-value pairs.
    """

    # Split the text into lines
    lines = text.splitlines()

    # Initialize the result dictionary
    result = {}

    # Initialize the current section
    current_section = None

    # Initialize the current section's dictionary
    current_section_dict = {}

    # Initialize the current value
    current_value = ""

    # Iterate over each line
    for line in lines:
        # Remove leading and trailing whitespace
        line = line.strip()

        # If the line is empty, ignore it
        if not line:
            continue

        # If the line is a section header
        if line.startswith('[') and line.endswith(']'):
            # Extract the section name
            section_name = line[1:-1]

            # If the section is DEFAULT, inherit its keys
            if section_name == 'DEFAULT':
                section_name = None

            # Update the current section
            current_section = section_name

            # Update the current section's dictionary
            current_section_dict = {}

            # If the section does not exist in the result, add it
            if current_section not in result:
                result[current_section] = {}

        # If the line is a key-value pair
        elif '=' in line:
            # Split the line into key and value
            key, value = line.split('=', 1)

            # Strip leading and trailing whitespace from key and value
            key = key.strip()
            value = value.strip()

            # If the line is a comment, ignore it
            if value.startswith(('#', ';'):
                continue

            # If the line is a quoted value
            if value.startswith('"') and value.endswith('"'):
                # Extract the value without quotes
                value = value[1:-1]

            # If the key is a reference to another section or key
            if '${' in key:
                # Extract the reference
                reference = key[2:-1]

                # If the reference is a section, use its value
                if ':' in reference:
                    section, key = reference.split(':', 1)
                    section = section.strip()
                    if section not in result:
                        result[section] = {}
                    key = key.strip()
                # If the reference is a key, use its value
                else:
                    key = reference.strip()

                # Try to get the value from the current section or the result
                try:
                    value = result[current_section][key]
                except KeyError:
                    # If the key is not found, raise a KeyError
                    raise KeyError(f"Key '{key}' not found")

            # Add the key-value pair to the current section's dictionary
            current_section_dict[key] = value

        # If the line is a continuation
        elif line.lstrip().startswith(' '):
            # Strip leading whitespace from the value
            value = line.lstrip().strip()

            # Add the value to the current value
            current_value += value + "\n"

        # If the line is an inline comment
        elif line.lstrip().startswith(('#', ';'):
            # Ignore the line
            continue

        # If the line is not recognized, raise a ValueError
        else:
            raise ValueError(f"Unrecognized line: {line}")

    # If there is a current value, add it to the current section's dictionary
    if current_value:
        current_section_dict[current_value.strip()] = None

    # If the current section is not None, add it to the result
    if current_section is not None:
        result[current_section] = current_section_dict

    # Return the result
    return result
```
Coding: TTL LRU (hard) Didn't work
```python
import heapq

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = {}
        self.heap = []

    def put(self, key, value):
        if not self.capacity:
            return

        if key in self.cache:
            self.cache[key].expiry = self.clock() + self.ttl
            self._mark_most_recent(key)
            self.heapify()
        else:
            if len(self.cache) >= self.capacity:
                self._evict_least_recent()
            self.cache[key] = {"value": value, "expiry": self.clock() + self.ttl, "recency": self.clock()}
            self._mark_most_recent(key)
            self.heapify()

    def get(self, key, default=None):
        if key not in self.cache:
            if default is None:
                return None
            return default
            # if key in self.cache:
            #     del self.cache[key]
            # return default

        if self.cache[key]["expiry"] < self.clock():
            del self.cache[key]
            return default
        else:
            self.cache[key]["recency"] = self.clock()
            return self.cache[key]["value"]

    def __len__(self):
        return len(self.cache)

    def stats(self):
        hits = sum(1 for _ in self.cache.values() if self.clock() < self.cache[_]["expiry"])
        misses = len(self.cache) - hits
        evictions = len(self.heap)
        expirations = sum(1 for _ in self.cache.values() if self.clock() >= self.cache[_]["expiry"])
        return {"hits": hits, "misses": misses, "evictions": evictions, "expirations": expirations}

    def _mark_most_recent(self, key):
        if key in self.heap:
            index = self.heap.index(key)
            self.heap[index] = (self.cache[key]["expiry"], self.cache[key]["recency"], key)
            heapq.heapify(self.heap)
        else:
            heapq.heappush(self.heap, (self.cache[key]["expiry"], self.cache[key]["recency"], key))

    def _evict_least_recent(self):
        _, _, key = heapq.heappop(self.heap)
        del self.cache[key]
```

Note that the `clock` function is assumed to be a zero-argument callable that returns the current time in seconds.
Coding: Parse duration Didn't work
```python
def parse_duration(s: str) -> int:
    """
    Converts a duration string into a total number of seconds.

    Args:
    s (str): A duration string with units of h, m, s.

    Returns:
    int: The total number of seconds.

    Raises:
    ValueError: For invalid input strings.
    """

    # Remove leading and trailing whitespace
    s = s.strip()

    # Check if the string is empty
    if not s:
        raise ValueError("Input string is empty")

    # Split the string into parts
    parts = s.split()

    # Initialize variables to store the total seconds and the unit multiplier
    total_seconds = 0
    multipliers = {"h": 3600, "m": 60, "s": 1}

    # Check if the input string contains any invalid characters
    for part in parts:
        # Check if the part is a digit
        if not part.isdigit():
            raise ValueError("Invalid character in input string")

        # Check if the part has a unit
        if len(part) > 1 and part[-1] in multipliers:
            # Check if the unit is valid
            if part[-1] not in multipliers:
                raise ValueError("Unknown unit in input string")
            # Extract the number and unit
            num = int(part[:-1])
            # Check if the number is non-negative
            if num < 0:
                raise ValueError("Negative number in input string")
            # Add the number of seconds to the total
            total_seconds += num * multipliers[part[-1]]
        else:
            # If the part is not a number with a unit, raise an error
            raise ValueError("Input string contains invalid format")

    # Check if the total number of units is valid
    if len(parts) > 3:
        raise ValueError("Too many units in input string")
    if len(parts) < 2:
        raise ValueError("Too few units in input string")

    # Return the total number of seconds
    return total_seconds
```
Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
    """
    This function takes a list of integers, removes duplicates, sorts the numbers, 
    collapses runs of 3 or more consecutive integers into "a..b", and returns a string 
    of the resulting numbers joined with commas.

    Args:
        nums (list[int]): A list of integers.

    Returns:
        str: A string of the summarized numbers.
    """

    # First, remove duplicates by converting the list to a set
    nums = set(nums)
    
    # If the set is empty, return an empty string
    if not nums:
        return ""

    # Sort the set in ascending order
    nums = sorted(nums)

    # Initialize an empty list to store the result
    result = []

    # Initialize the start and end of the current run
    start = end = nums[0]

    # Iterate over the sorted numbers
    for num in nums[1:]:
        # If the current number is a consecutive integer to the previous one, update the end of the run
        if num == end + 1:
            end = num
        # If the current number is not a consecutive integer, append the current run to the result and start a new run
        else:
            if start == end:
                result.append(str(start))
            else:
                result.append(f"{start}..{end}")
            start = end = num

    # Append the last run to the result
    if start == end:
        result.append(str(start))
    else:
        result.append(f"{start}..{end}")

    # Join the result with commas and return it
    return ",".join(result)
```
Coding: Top customers Didn't work
```python
from collections import defaultdict
from typing import List, Dict, Tuple

def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    """
    This function takes a list of dictionaries representing customer transactions and returns the top n customers with the highest total amount spent.
    
    :param rows: A list of dictionaries, where each dictionary has a "customer" key and an "amount" key.
    :param n: The number of top customers to return.
    :return: A list of tuples, where each tuple contains the customer name and their total amount spent.
    """

    # Create a dictionary to store the total amount spent by each customer
    customer_amounts = defaultdict(float)

    # Iterate over each row in the input list
    for row in rows:
        # Get the customer name and amount from the current row
        customer = row["customer"].strip().lower()
        amount = float(row["amount"].strip("$").replace(",", ""))

        # If the amount is negative, consider it as a refund
        if amount < 0:
            amount = -amount

        # Add the amount to the customer's total
        customer_amounts[customer] += amount

    # Sort the customers by total amount in descending order, then by name
    sorted_customers = sorted(customer_amounts.items(), key=lambda x: (-x[1], x[0]))

    # Return the top n customers
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    """
    Builds a URL slug from a given title.

    Args:
    title (str): The title to slugify.
    max_len (int): The maximum length of the slug. Defaults to 50.

    Returns:
    str: The slugified title.
    """

    # Transliterate accented characters to ASCII and drop any other non-ASCII characters
    title = ''.join(c for c in unicodedata.normalize('NFD', title) if unicodedata.category(c) != 'Mn')

    # Lowercase
    title = title.lower()

    # Replace every run of characters other than a-z and 0-9 with a single "-"
    title = re.sub(r'[^a-z0-9]+', '-', title)

    # Strip leading/trailing "-"
    title = title.strip('-')

    # Split into words
    words = title.split('-')

    # Initialize the result and the prefix length
    result = ''
    prefix_len = 0

    # Iterate over the words
    for word in words:
        # If adding the word would exceed max_len, break
        if prefix_len + len(word) + 1 > max_len:
            break

        # Add the word to the result and update the prefix length
        result += word + '-'
        prefix_len += len(word) + 1

    # Remove the trailing '-'
    result = result[:-1]

    return result
```
Coding: Token bucket Didn't work
```python
import time
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        """
        Initialize the token bucket.

        :param capacity: The maximum number of tokens in the bucket.
        :param refill_per_sec: The rate at which tokens are refilled per second.
        :param clock: A function that returns the current time in seconds.
        """
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.last_refill = clock()
        self.tokens = capacity
        self.clock = clock

    @property
    def tokens(self) -> float:
        """
        Get the current number of tokens in the bucket.

        :return: The current number of tokens.
        """
        return self._tokens

    def allow(self, cost: float = 1) -> bool:
        """
        Refill the bucket and attempt to allow the requested number of tokens.

        :param cost: The number of tokens to attempt to get. Defaults to 1.
        :return: True if tokens were successfully obtained, False otherwise.
        """
        elapsed = self.clock() - self.last_refill
        self.last_refill = self.clock()
        self.tokens = min(self.capacity, self.tokens + elapsed * self.refill_per_sec)
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False
```
Decisions: Refund window (hard) 50% right
```json
{
  "outcome": {
    "full_refund": 0.6,
    "partial_refund": 0.3,
    "store_credit": 0.1,
    "nothing": 0
  },
  "defective": {
    "true": false,
    "false": 1
  }
}
```
Decisions: Incident matrix (hard) 0% right
```json
{
  "severity": {
    "0": 0.95,
    "1": 0.05,
    "2": 0,
    "3": 0
  },
  "page": {
    "true": 0.5,
    "false": 0.5
  }
}
```
Decisions: Tool followup (hard) 50% right
```json
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 0.9,
    "email": 0.1,
    "none": 0.0
  },
  "confirm": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Legit security alert (hard) 0% right
```json
{
  "1": {
    "true": 0.99,
    "false": 0.01
  },
  "2": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Meeting slot (hard) 0% right
{
  "slot": {
    "A": 0.0, "B": 0.0, "C": 0.0, "D": 1.0
  },
  "raj_last": {
    "true": 0.0, "false": 1.0
  }
Decisions: Review mixed (hard) 50% right
```json
{
  "hardware": {
    "true": 1.0,
    "false": 0.0
  },
  "support": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Support checkout down 33% right
```json
{
  "department": {
    "billing": 0.6,
    "technical": 0.3,
    "account": 0.1
  },
  "urgency": {
    "0": 0.0,
    "1": 0.4,
    "2": 0.3,
    "3": 0.3
  },
  "outage": {
    "true": 1.0,
    "false": 0.0
  }
}
```
Decisions: Refund wrong plan 100% right
```json
{
  "department": {
    "billing": 0.9,
    "technical": 0.05,
    "account": 0.05,
    "shipping": 0.0,
    "sales": 0.0
  },
  "refund": {
    "true": 0.95,
    "false": 0.05
  },
  "tone": {
    "frustrated": 0.0,
    "calm": 1.0
  }
}
```
Decisions: Moderation doxxing 50% right
{
  "policy": {
    "none": 0.95,
    "harassment": 0,
    "hate": 0,
    "spam": 0,
    "self_harm": 0
  },
  "personal_info": {
    "true": 0.5,
    "false": 0.5
  }
}
Decisions: Route calendar 100% right
```json
{
  "tool": {
    "calendar": 0.9,
    "email": 0.05,
    "web_search": 0.05,
    "none": 0.005
  },
  "confirm": {
    "true": 0.8,
    "false": 0.2
  }
}
```
Decisions: Doc invoice missing due 50% right
```json
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 0.0,
    "false": 1.0
  }
}
```
Decisions: Phishing paypal 0% right
{
  "phishing": {
    "true": 0.9,
    "false": 0.1
  },
  "risk": {
    "0": 0.05,
    "1": 0.3,
    "2": 0.4,
    "3": 0.15
  }
Decisions: Pii ssn email 0% right
```json
{
  "1": {
    "none": 0.8,
    "contact": 0.1,
    "financial": 0.05,
    "government_id": 0.05,
    "health": 0.0
  },
  "2": {
    "true": 0.1,
    "false": 0.9
  }
}
```
Decisions: Review mixed 100% right
```json
{
  "sentiment": {
    "positive": 0,
    "neutral": 0,
    "negative": 0.9
  },
  "defect": {
    "true": 1,
    "false": 0
  },
  "recommend": {
    "true": 0,
    "false": 1
  }
}
```
Documents: Saas escalator (hard) 0% right
{
  "year2_price_per_seat_month": 47.69,
  "year3_price_per_seat_month": 49.13,
  "year1_invoice": 5400.00,
  "year2_invoice": 5688.00,
  "addon_months_billed": 16,
  "addon_invoice": 752.80,
  "year3_invoice": 8321.60,
  "year3_discount_percent": 0.00,
  "total_contract_value": 219440.00,
  "contract_end_date": "2029-02-28"
Documents: Expense thread 67% right
```
{
  "employee_id": "dana.whitfield@corvane.com",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {"date": "2025-02-24", "category": "airfare", "amount_usd": 1179.32},
    {"date": "2025-02-24", "category": "ground_transport", "amount_usd": 40.00},
    {"date": "2025-02-25", "category": "meals", "amount_usd": 65.00},
    {"date": "2025-02-26", "category": "lodging", "amount_usd": 475.52},
    {"date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.83}
  ],
  "rejected_item_count": 0,
  "per_diem_days": 3,
  "per_diem_usd": 195.00,
  "total_reimbursable_usd": 1938.17,
  "approver_email": "priya.raman@corvane.com"
}
```
Documents: Lease amendment 0% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150,
  "monthly_rent_from_2025_06_01": 2250,
  "late_fee_from_2025_06_01": 107.5,
  "security_deposit": 2150,
  "total_pet_deposits": 800,
  "total_monthly_payment_july_2025": 2275,
  "move_in_payment": 2150 + 2150 + 400 = 2770
}
Documents: Ticket SLA 54% right
```json
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address", "invoice_pdf"],
  "affected_orders": ["SO-99812", "SO-99815", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T17:00",
  "sla_due_utc": "2025-09-15T17:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
```
Documents: Sales footnotes 0% right
{
  "q3_total_usd": 4902,
  "q2_total_usd": 4585,
  "q2_central_originally_reported_usd": 2947,
  "q2_to_q3_change_pct": 5.6,
  "top_region_q3": "West",
  "fastest_growing_region_q1_to_q3": null,
  "regions_declining_q2_to_q3": null,
  "international_q3_organic_usd": 1951,
  "west_excluding_mountain_q3_usd": 4280
}

Size: 3.2B parameters. First tested OCT 10.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.