Review · updated OCT 11

Llama 3.1 70B Instruct review: not one we'd recommend right now

It scored 39 out of 100, #41 of 56. It solved 10 of 30 coding jobs and scored 44 on reading documents. Runs on a Mac with 64 GB.

The short version
  • Llama 3.1 70B Instruct is a free model from Meta that you can run on your own computer. In our tests it's not one we'd recommend right now: 39 out of 100, #41 of 56.
  • It solved 10 of 30 coding jobs and scored 44 on reading documents. On our hardest tasks it scored 22.
  • Runs on a Mac with 64 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Llama 3.1 70B Instruct got 10 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Llama 3.1 70B Instruct scored 44; the best model scored 100.

TestScorePublic questionsSecret questions
Coding332935
Reading documents445541
Decisions848285

On the 18 hardest tasks (included in the scores above) it scored 22. This number separates the top models.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Llama 3.1 70B Instruct in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBYesQ4
Mac, 96 GBYesQ6
Mac, 128 GBYesQ8

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Llama 3.1 70B Instruct GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Llama 3.1 70B Instruct wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re

def parse_ini(text: str) -> dict[str, dict[str, str]]:
    """
    Parse INI text into a dictionary of sections and their key-value pairs.

    :param text: The INI text to parse
    :return: A dictionary of sections and their key-value pairs
    """

    sections = {}
    default_section = {}

    current_section = None

    for line in text.splitlines():
        # Ignore lines that are empty or whose first non-space character is a comment
        line = line.strip()
        if not line or line.startswith((';', '#')):
            continue

        # Check for section headers
        if line.startswith('[') and line.endswith(']'):
            section_name = line[1:-1].strip()
            if section_name == 'DEFAULT':
                current_section = default_section
            else:
                current_section = sections.setdefault(section_name, {})
            continue

        # If we haven't seen a section header yet, raise ValueError
        if current_section is None:
            raise ValueError("Key line before any section header")

        # Split the line into key and value
        match = re.match(r'([^:=]+)[:=](.*)', line)
        if not match:
            raise ValueError("Malformed line")
        key, value = match.groups()
        key = key.strip().lower()

        # Handle continuation lines
        while True:
            next_line = text.splitlines()[text.splitlines().index(line) + 1].strip()
            if next_line and next_line.startswith(' '):
                value += '\n' + next_line.strip()
                text = text.replace(next_line, '', 1)
            else:
                break

        # Remove inline comments and double quotes
        value = re.sub(r'(;| #).*$', '', value)
        if value.startswith('"') and value.endswith('"'):
            value = value[1:-1]
        else:
            value = value.strip()

        # Store the key-value pair
        current_section[key] = value

    # Interpolate values
    for section_name, section in sections.items():
        for key, value in section.items():
            while True:
                match = re.search(r'\${([^:]+)(:[^}]+)?}', value)
                if not match:
                    break
                ref_key, ref_section = match.groups()
                if ref_section:
                    ref_section = ref_section[1:]
                else:
                    ref_section = section_name
                ref_section_data = sections.get(ref_section, default_section)
                ref_value = ref_section_data.get(ref_key)
                if ref_value is None:
                    raise KeyError(f"Missing key {ref_key}")
                if ref_value == value:
                    raise ValueError("Reference cycle")
                value = value.replace(match.group(0), ref_value, 1)

    # Inherit DEFAULT section keys
    for section in sections.values():
        section.update({k: v for k, v in default_section.items() if k not in section})

    return sections
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from typing import Callable, Dict, Optional

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache = OrderedDict()
        self.stats = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def put(self, key, value):
        now = self.clock()
        if key in self.cache:
            del self.cache[key]
        elif self.capacity > 0 and len(self.cache) >= self.capacity:
            self._remove_expired(now)
            if len(self.cache) >= self.capacity:
                self.cache.popitem(last=False)
                self.stats["evictions"] += 1
        self.cache[key] = (value, now + self.ttl)

    def get(self, key, default=None):
        now = self.clock()
        if key in self.cache:
            value, expiry = self.cache[key]
            if expiry > now:
                self.cache.move_to_end(key)
                self.stats["hits"] += 1
                return value
            else:
                del self.cache[key]
                self.stats["expirations"] += 1
        self.stats["misses"] += 1
        return default

    def __len__(self):
        now = self.clock()
        self._remove_expired(now)
        return len(self.cache)

    def stats(self) -> Dict[str, int]:
        return self.stats.copy()

    def _remove_expired(self, now: float):
        while self.cache and self.cache[next(reversed(self.cache))][1] <= now:
            del self.cache[next(reversed(self.cache))]
            self.stats["expirations"] += 1
```
Coding: Parse duration Didn't work
```python
import re
from enum import Enum
from typing import Match

class Unit(Enum):
    HOURS = 1
    MINUTES = 2
    SECONDS = 3

    @property
    def value(self):
        return {
            Unit.HOURS: 3600,
            Unit.MINUTES: 60,
            Unit.SECONDS: 1,
        }.get(self)

    @classmethod
    def from_str(cls, s: str) -> 'Unit':
        s = s.lower()
        if s == 'h':
            return cls.HOURS
        elif s == 'm':
            return cls.MINUTES
        elif s == 's':
            return cls.SECONDS
        else:
            raise ValueError(f"Unknown unit: {s}")


def parse_duration(s: str) -> int:
    """
    Convert a duration string into a total number of seconds.

    :param s: A duration string.
    :return: The total number of seconds.
    :raises ValueError: If the input string is invalid.
    """

    # Remove whitespace and normalize to lowercase
    s = re.sub(r'\s+', '', s).lower()

    # Check for empty/blank strings
    if not s:
        raise ValueError("Empty string")

    # Parse each part of the duration
    parts = re.findall(r'(\d+)([hms])', s)
    if len(parts)!= len(s):
        raise ValueError("Invalid duration string")

    # Validate the order and uniqueness of units
    units = [Unit.from_str(unit) for _, unit in parts]
    if units!= sorted(units):
        raise ValueError("Units are out of order")
    if len(set(units))!= len(units):
        raise ValueError("Repeated units")

    # Convert the duration to seconds
    total_seconds = 0
    for value, unit in parts:
        total_seconds += int(value) * Unit.from_str(unit).value

    return total_seconds
```
Coding: Summarize ranges Worked
```python
from itertools import groupby

def summarize_ranges(nums: list[int]) -> str:
    """
    Summarize a list of integers into a string, collapsing runs of 3 or more consecutive integers.

    Args:
    - nums: A list of integers.

    Returns:
    - A string representing the input list, with runs of 3 or more consecutive integers collapsed.
    """
    
    # If the input list is empty, return an empty string
    if not nums:
        return ""

    # Sort the numbers and remove duplicates by converting the list to a set
    nums = sorted(set(nums))

    # Initialize variables to store the result and the current run
    result = []
    current_run = [nums[0]]

    # Iterate over the sorted numbers
    for num in nums[1:]:
        # If the current number is consecutive to the last number in the current run, add it to the run
        if num == current_run[-1] + 1:
            current_run.append(num)
        # Otherwise, process the current run and start a new one
        else:
            # If the current run has 3 or more numbers, collapse it into "a..b"
            if len(current_run) >= 3:
                result.append(f"{current_run[0]}..{current_run[-1]}")
            # If the current run has 2 numbers, add them individually
            elif len(current_run) == 2:
                result.extend(map(str, current_run))
            # If the current run has 1 number, add it individually
            else:
                result.append(str(current_run[0]))
            # Start a new run
            current_run = [num]

    # Process the last run
    if len(current_run) >= 3:
        result.append(f"{current_run[0]}..{current_run[-1]}")
    elif len(current_run) == 2:
        result.extend(map(str, current_run))
    else:
        result.append(str(current_run[0]))

    # Join the result with commas and return
    return ",".join(result)
```
Coding: Top customers Didn't work
```python
from collections import defaultdict
from decimal import Decimal, ROUND_HALF_UP
from typing import List, Dict, Tuple

def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    """
    Returns the top n customers by total amount, sorted by total descending then name ascending.

    Args:
    rows (List[Dict]): A list of dictionaries containing customer information.
    n (int): The number of top customers to return.

    Returns:
    List[Tuple[str, float]]: A list of tuples containing the customer name and total amount.
    """

    # Create a dictionary to store the total amount for each customer
    customers = defaultdict(lambda: {"total": Decimal(0), "name": ""})

    # Iterate over each row in the input list
    for row in rows:
        # Get the customer name and amount from the row
        customer_name = row.get("customer", "").strip().casefold()
        amount = row.get("amount", "")

        # Skip rows with empty customer names or amounts that cannot be parsed
        if not customer_name or not amount:
            continue

        # Parse the amount string into a Decimal object
        try:
            amount = Decimal(amount.replace("$", "").replace(",", ""))
        except ValueError:
            continue

        # Update the customer's total amount and store the trimmed customer name
        if not customers[customer_name]["name"]:
            customers[customer_name]["name"] = row["customer"].strip()
        customers[customer_name]["total"] += amount

    # Round the total amounts to 2 decimals
    for customer in customers.values():
        customer["total"] = float(customer["total"].quantize(Decimal(".01"), rounding=ROUND_HALF_UP))

    # Sort the customers by total descending then name ascending
    sorted_customers = sorted(
        ((customer["name"], customer["total"]) for customer in customers.values()),
        key=lambda x: (-x[1], x[0])
    )

    # Return the top n customers
    return sorted_customers[:n]
```
Coding: Slugify Didn't work
```python
import unicodedata
import re

def slugify(title: str, max_len: int = 50) -> str:
    """
    Builds a URL slug from a given title.

    :param title: The title to be slugified
    :param max_len: The maximum length of the slug (default is 50)
    :return: The slugified title
    """

    # Transliterate accented characters to ASCII and drop any other non-ASCII characters
    title = unicodedata.normalize('NFKD', title).encode('ascii', 'ignore').decode('ascii')

    # Lowercase
    title = title.lower()

    # Replace every run of characters other than a-z and 0-9 with a single "-"
    title = re.sub(r'[^\w\s-]', '', title).replace(' ', '-')

    # Strip leading/trailing "-"
    title = title.strip('-')

    # If the slug is longer than max_len, shorten it without cutting a word
    if len(title) > max_len:
        words = title.split('-')
        slug = ''
        for word in words:
            if len(slug) + len(word) + 1 > max_len:
                break
            slug += word + '-'
        slug = slug.rstrip('-')
        if not slug:
            slug = title[:max_len]
    else:
        slug = title

    return slug
```
Coding: Token bucket Worked
```python
from types import SimpleNamespace
from typing import Callable

class TokenBucket:
    def __init__(self, capacity: float, refill_per_sec: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.refill_per_sec = refill_per_sec
        self.clock = clock
        self.state = SimpleNamespace(last_refill_time=self.clock(), tokens=self.capacity)

    def _refill(self):
        now = self.clock()
        elapsed_time = now - self.state.last_refill_time
        self.state.tokens = min(self.capacity, self.state.tokens + elapsed_time * self.refill_per_sec)
        self.state.last_refill_time = now

    def allow(self, cost: float = 1) -> bool:
        self._refill()
        if self.state.tokens >= cost:
            self.state.tokens -= cost
            return True
        return False

    @property
    def tokens(self) -> float:
        self._refill()
        return self.state.tokens
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "full_refund": 0,
    "partial_refund": 0.8,
    "store_credit": 0.2,
    "nothing": 0
  },
  "defective": {
    "true": 0,
    "false": 1
  }
}
Decisions: Incident matrix (hard) 0% right
{"severity": {"0": 0, "1": 0.3, "2": 0.7, "3": 0}, "page": {"true": 0.7, "false": 0.3}}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "web_search": 0,
    "calculator": 0,
    "calendar": 1,
    "email": 0,
    "none": 0
  },
  "confirm": {
    "true": 1,
    "false": 0
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.05,
    "false": 0.95
  }
}
Decisions: Meeting slot (hard) 0% right
{
  "slot": {
    "A": 0,
    "B": 0,
    "C": 1,
    "D": 0
  },
  "raj_last": {
    "true": 0,
    "false": 1
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 0.9,
    "false": 0.1
  },
  "support": {
    "true": 0.8,
    "false": 0.2
  }
}
Decisions: Support checkout down 100% right
{
  "department": {
    "billing": 0.2,
    "technical": 0.6,
    "account": 0.1,
    "shipping": 0,
    "sales": 0.1
  },
  "urgency": {
    "0": 0,
    "1": 0,
    "2": 0.2,
    "3": 0.8
  },
  "outage": {
    "true": 0.8,
    "false": 0.2
  }
}
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.9,
    "technical": 0.05,
    "account": 0.025,
    "shipping": 0.025,
    "sales": 0.0
  },
  "refund": {
    "true": 0.9,
    "false": 0.1
  },
  "tone": {
    "frustrated": 0.0,
    "calm": 1.0
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.2,
    "harassment": 0.7,
    "hate": 0.05,
    "spam": 0.05,
    "self_harm": 0
  },
  "personal_info": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0,
    "calculator": 0,
    "calendar": 0.9,
    "email": 0,
    "none": 0.1
  },
  "confirm": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 0.95,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.05
  },
  "missing_due_date": {
    "true": 0.8,
    "false": 0.2
  }
}
Decisions: Phishing paypal 100% right
{"phishing": {"true": 0.95, "false": 0.05}, "risk": {"0": 0.0, "1": 0.0, "2": 0.05, "3": 0.95}}
Decisions: Pii ssn email 100% right
```
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.2,
    "financial": 0.2,
    "government_id": 0.6,
    "health": 0.0
  },
  "sensitive": {
    "true": 0.6,
    "false": 0.4
  }
}
```
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.1,
    "neutral": 0.2,
    "negative": 0.7
  },
  "defect": {
    "true": 1,
    "false": 0
  },
  "recommend": {
    "true": 0,
    "false": 1
  }
}
Documents: Saas escalator (hard) 20% right
{
  "year2_price_per_seat_month": 47.70,
  "year3_price_per_seat_month": 47.70,
  "year1_invoice": 64800.00,
  "year2_invoice": 68614.40,
  "addon_months_billed": 7,
  "addon_invoice": 33215.60,
  "year3_invoice": 68614.40,
  "year3_discount_percent": 15,
  "total_contract_value": 204730.40,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 75% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    { "date": "2025-02-24", "category": "airfare", "amount_usd": 1184.60 },
    { "date": "2025-02-24", "category": "ground_transport", "amount_usd": 38.88 },
    { "date": "2025-02-25", "category": "meals", "amount_usd": 229.69 },
    { "date": "2025-02-26", "category": "lodging", "amount_usd": 467.16 },
    { "date": "2025-02-27", "category": "ground_transport", "amount_usd": 44.68 }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 2,
  "per_diem_usd": 130.00,
  "total_reimbursable_usd": 2075.01,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 67% right
```
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2239.00,
  "late_fee_from_2025_06_01": 111.95,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2309.00,
  "move_in_payment": 4650.00
}
```
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T17:00",
  "sla_due_utc": "2025-09-15T23:00:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 33% right
{
  "q3_total_usd": 15446,
  "q2_total_usd": 14364,
  "q2_central_originally_reported_usd": 3047,
  "q2_to_q3_change_pct": 7.4,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731,
  "west_excluding_mountain_q3_usd": 4201
}

Size: 71B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.