Review · updated OCT 11

Command A review: not one we'd recommend right now

It scored 45 out of 100, #34 of 56. It solved 10 of 30 coding jobs and scored 56 on reading documents. Runs on a Mac with 96 GB.

The short version
  • Command A is a free model from Cohere that you can run on your own computer. In our tests it's not one we'd recommend right now: 45 out of 100, #34 of 56.
  • It solved 10 of 30 coding jobs and scored 56 on reading documents. On our hardest tasks it scored 27.
  • Runs on a Mac with 96 GB.

Coding

Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. Command A got 10 of 30 right. The best local coders solved 29 of 30.

Reading documents

The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. Command A scored 56; the best model scored 100.

TestScorePublic questionsSecret questions
Coding332935
Reading documents567452
Decisions888987

On the 18 hardest tasks (included in the scores above) it scored 27. This number separates the top models.

This model did much better on our public questions than on the secret ones. Sometimes that's luck; it can also mean the model has seen the public questions before.

We tested the full-size version online. The smaller versions you download usually score a little lower.

Can your computer run it?

You can download Command A in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.

Your computerRuns it?Version to download
Graphics card, 8 GBNo–
Graphics card, 12 GBNo–
Graphics card, 16 GBNo–
Graphics card, 24 GBNo–
Graphics card, 32 GBNo–
Mac, 16 GBNo–
Mac, 24 GBNo–
Mac, 32 GBNo–
Mac, 48 GBNo–
Mac, 64 GBNo–
Mac, 96 GBYesQ4
Mac, 128 GBYesQ6

For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out

How to try it

Search for "Command A GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.

Want the strongest all-rounder instead? Qwen3.6 27B scored 97.

Its actual answers

What Command A wrote for our public questions, word for word.

Coding: INI parser (hard) Didn't work
```python
import re
from typing import Dict

def parse_ini(text: str) -> Dict[str, Dict[str, str]]:
    lines = text.splitlines()
    result: Dict[str, Dict[str, str]] = {}
    current_section: Dict[str, str] = {}
    section_stack: list[str] = []
    default_section: Dict[str, str] = {}
    in_section = False

    def strip_comments(line: str) -> str:
        if '"' not in line:
            return re.split(r'\s*(?<!\\)[;#]', line)[0].rstrip()
        parts = re.split(r'(?<!\\)"', line)
        for i in range(1, len(parts), 2):
            parts[i] = re.split(r'\s*(?<!\\)[;#]', parts[i])[0]
        return ''.join(parts).rstrip()

    def interpolate(value: str, section_name: str, seen: set[str] = None) -> str:
        if seen is None:
            seen = set()
        if value in seen:
            raise ValueError("Interpolation cycle detected")
        seen.add(value)
        def replace(match: re.Match) -> str:
            key_ref = match.group(1)
            if ':' in key_ref:
                ref_section, ref_key = key_ref.split(':', 1)
                if ref_section not in result:
                    raise KeyError(f"Section {ref_section} not found")
                if ref_key not in result[ref_section]:
                    raise KeyError(f"Key {ref_key} not found in section {ref_section}")
                return interpolate(result[ref_section][ref_key], ref_section, seen)
            else:
                if key_ref not in current_section:
                    if key_ref not in default_section:
                        raise KeyError(f"Key {key_ref} not found in section {section_name}")
                    return interpolate(default_section[key_ref], 'DEFAULT', seen)
                return interpolate(current_section[key_ref], section_name, seen)
        value = re.sub(r'\${(.*?)}', replace, value)
        seen.remove(value)
        return value

    for line_no, line in enumerate(lines, 1):
        line = strip_comments(line).rstrip()
        if not line or line.isspace():
            continue
        stripped_line = line.strip()
        if stripped_line.startswith('[') and stripped_line.endswith(']'):
            if not in_section and section_stack:
                raise ValueError(f"Malformed line {line_no}: {line}")
            section_name = stripped_line[1:-1].strip()
            if section_name == 'DEFAULT':
                current_section = default_section
            else:
                current_section = result.setdefault(section_name, {})
                for key, value in default_section.items():
                    if key not in current_section:
                        current_section[key] = value
            in_section = True
            section_stack.append(section_name)
            continue
        if not in_section:
            raise ValueError(f"Key outside section at line {line_no}: {line}")
        if line.startswith(' ') or line.startswith('\t'):
            if not section_stack:
                raise ValueError(f"Malformed continuation line {line_no}: {line}")
            last_section = section_stack[-1]
            last_key = next(reversed(result[last_section]), None)
            if last_key is None:
                raise ValueError(f"Malformed continuation line {line_no}: {line}")
            result[last_section][last_key] += '\n' + line.strip()
            continue
        match = re.match(r'\s*([^\s=:#]+)\s*[:=]\s*(.*)\s*$', line)
        if not match:
            raise ValueError(f"Malformed line {line_no}: {line}")
        key, value = match.groups()
        key = key.strip().lower()
        if value.startswith('"') and value.endswith('"'):
            value = value[1:-1].replace('\\"', '"')
        else:
            value = value.strip()
        if section_stack:
            current_section[key] = value
        else:
            default_section[key] = value

    for section_name, section in result.items():
        for key, value in section.items():
            try:
                result[section_name][key] = interpolate(value, section_name)
            except (KeyError, ValueError) as e:
                raise ValueError(f"Error interpolating key {key} in section {section_name}: {e}")

    if 'DEFAULT' in result:
        del result['DEFAULT']

    return result
```
Coding: TTL LRU (hard) Didn't work
```python
from collections import OrderedDict
from typing import Any, Callable, Dict, Optional

class TTLCache:
    def __init__(self, capacity: int, ttl: float, clock: Callable[[], float]):
        self.capacity = capacity
        self.ttl = ttl
        self.clock = clock
        self.cache: OrderedDict[Any, tuple[Any, float]] = OrderedDict()
        self.stats: Dict[str, int] = {"hits": 0, "misses": 0, "evictions": 0, "expirations": 0}

    def _remove_expired(self) -> None:
        now = self.clock()
        expired = [key for key, (_, expiry) in self.cache.items() if expiry <= now]
        for key in expired:
            del self.cache[key]
            self.stats["expirations"] += 1

    def put(self, key: Any, value: Any) -> None:
        if self.capacity == 0:
            return

        self._remove_expired()

        now = self.clock()
        expiry = now + self.ttl

        if key in self.cache:
            self.cache.move_to_end(key)
        elif len(self.cache) >= self.capacity:
            self.cache.popitem(last=False)
            self.stats["evictions"] += 1

        self.cache[key] = (value, expiry)

    def get(self, key: Any, default: Optional[Any] = None) -> Any:
        if self.capacity == 0:
            self.stats["misses"] += 1
            return default

        self._remove_expired()

        now = self.clock()
        if key in self.cache:
            value, expiry = self.cache[key]
            if expiry > now:
                self.cache.move_to_end(key)
                self.stats["hits"] += 1
                return value
            else:
                del self.cache[key]
                self.stats["expirations"] += 1

        self.stats["misses"] += 1
        return default

    def __len__(self) -> int:
        now = self.clock()
        return sum(1 for _, expiry in self.cache.values() if expiry > now)

    def stats(self) -> Dict[str, int]:
        return self.stats.copy()
```
Coding: Parse duration Didn't work
```python
import re
from typing import Dict, List, Tuple

def parse_duration(s: str) -> int:
    """
    Converts a duration string into a total number of seconds.

    Args:
        s: A string representing a duration, e.g., "1h30m", "45s", "2H", "1h 5m 10s", "90m".

    Returns:
        The total number of seconds represented by the input string.

    Raises:
        ValueError: If the input string is invalid (empty, contains unknown units, decimals, repeated units, or units out of order).
    """

    # Remove leading and trailing whitespace
    s = s.strip()
    if not s:
        raise ValueError("Input string is empty or blank")

    # Define the order of units and their conversion factors
    units: List[str] = ['h', 'm', 's']
    factors: Dict[str, int] = {'h': 3600, 'm': 60, 's': 1}

    # Split the input string into parts based on whitespace
    parts: List[str] = re.split(r'\s+', s)

    # Initialize variables to track the total seconds and the last seen unit
    total_seconds: int = 0
    last_unit_index: int = -1

    # Iterate over each part
    for part in parts:
        # Use regular expression to match the number and unit
        match = re.match(r'^(\d+)([hms])$', part, re.IGNORECASE)
        if not match:
            raise ValueError(f"Invalid part: {part}")

        # Extract the number and unit
        num: int = int(match.group(1))
        unit: str = match.group(2).lower()

        # Check if the unit is valid and in order
        current_unit_index: int = units.index(unit)
        if current_unit_index <= last_unit_index:
            raise ValueError(f"Units out of order or repeated: {unit}")
        last_unit_index = current_unit_index

        # Add the converted value to the total seconds
        total_seconds += num * factors[unit]

    return total_seconds
```
Coding: Summarize ranges Didn't work
```python
def summarize_ranges(nums: list[int]) -> str:
    if not nums:
        return ""

    nums = sorted(set(nums))
    result = []
    start = nums[0]

    for i in range(1, len(nums)):
        if nums[i] != nums[i - 1] + 1:
            if nums[i - 1] != start:
                result.append(f"{start}..{nums[i - 1]}")
            else:
                result.append(str(start))
                if nums[i - 1] != nums[i] - 1:
                    result.append(str(nums[i] - 1))
            start = nums[i]

    if nums[-1] != start:
        result.append(f"{start}..{nums[-1]}")
    else:
        result.append(str(start))

    return ",".join(result)
```
Coding: Top customers Didn't work
```python
import re
from collections import defaultdict
from typing import Dict, List, Tuple


def top_customers(rows: List[Dict], n: int) -> List[Tuple[str, float]]:
    """
    Returns the top n customers by total amount spent.

    Args:
        rows (List[Dict]): List of dictionaries containing customer and amount information.
        n (int): Number of top customers to return.

    Returns:
        List[Tuple[str, float]]: List of tuples containing customer name and total amount spent,
            sorted by total amount descending and then by name ascending.
    """

    # Initialize a dictionary to store customer totals
    customer_totals: Dict[str, float] = defaultdict(float)
    # Dictionary to store the first valid name spelling for each customer
    customer_names: Dict[str, str] = {}

    # Regular expression to parse amounts
    amount_pattern = re.compile(r"^\s*[\$]?(-?\d{1,3}(?:,\d{3})*(?:\.\d{1,2})?)\s*$")

    for row in rows:
        customer = row.get("customer", "").strip()
        amount_str = row.get("amount", "").strip()

        # Skip rows with empty customer name or invalid amount
        if not customer or not amount_pattern.match(amount_str):
            continue

        # Parse the amount
        amount_str = amount_str.replace(",", "").replace("$", "")
        try:
            amount = float(amount_str)
        except ValueError:
            continue

        # Store the first valid name spelling for the customer
        customer_lower = customer.lower()
        if customer_lower not in customer_names:
            customer_names[customer_lower] = customer

        # Update the customer's total
        customer_totals[customer_lower] += amount

    # Prepare the result list with proper name spelling and rounded totals
    result = [
        (customer_names[customer], round(total, 2))
        for customer, total in customer_totals.items()
    ]

    # Sort by total descending, then by name ascending
    result.sort(key=lambda x: (-x[1], x[0]))

    # Return the top n customers
    return result[:n]
```
Coding: Slugify Worked
```python
import re
import unicodedata


def slugify(title: str, max_len: int = 50) -> str:
    """
    Builds a URL slug from a given title.

    Args:
        title (str): The title to be converted into a slug.
        max_len (int, optional): Maximum length of the slug. Defaults to 50.

    Returns:
        str: The generated URL slug.
    """

    # Transliterate accented characters to ASCII and drop any other non-ASCII characters
    ascii_title = (
        unicodedata.normalize("NFKD", title)
        .encode("ascii", "ignore")
        .decode("ascii")
        .lower()
    )

    # Replace every run of characters other than a-z and 0-9 with a single "-", and strip leading/trailing "-"
    slug = re.sub(r"[^a-z0-9]+", "-", ascii_title).strip("-")

    # If the slug is longer than max_len, shorten it without cutting a word
    if len(slug) > max_len:
        words = slug.split("-")
        shortened_slug = ""
        for word in words:
            if len(shortened_slug) + len(word) + (1 if shortened_slug else 0) <= max_len:
                shortened_slug += ("-" if shortened_slug else "") + word
            else:
                # If even the first word is longer than max_len, hard-cut it to max_len characters
                if not shortened_slug:
                    return word[:max_len]
                break
        slug = shortened_slug

    return slug
```
Coding: Token bucket Worked
```python
from typing import Callable

class TokenBucket:
    def __init__(
        self, capacity: float, refill_per_sec: float, clock: Callable[[], float]
    ) -> None:
        self._capacity = capacity
        self._refill_per_sec = refill_per_sec
        self._clock = clock
        self._tokens = capacity
        self._last_refill_time = clock()

    @property
    def tokens(self) -> float:
        self._refill()
        return self._tokens

    def allow(self, cost: float = 1.0) -> bool:
        self._refill()
        if self._tokens >= cost:
            self._tokens -= cost
            return True
        return False

    def _refill(self) -> None:
        now = self._clock()
        elapsed = now - self._last_refill_time
        refill_amount = elapsed * self._refill_per_sec
        self._tokens = min(self._capacity, self._tokens + refill_amount)
        self._last_refill_time = now
```
Decisions: Refund window (hard) 50% right
{
  "outcome": {
    "full_refund": 0.0,
    "partial_refund": 0.8,
    "store_credit": 0.2,
    "nothing": 0.0
  },
  "defective": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Incident matrix (hard) 100% right
{
  "severity": {
    "0": 0.0,
    "1": 0.9,
    "2": 0.1,
    "3": 0.0
  },
  "page": {
    "true": 0.1,
    "false": 0.9
  }
}
Decisions: Tool followup (hard) 100% right
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 1.0,
    "email": 0.0,
    "none": 0.0
  },
  "confirm": {
    "true": 1.0,
    "false": 0.0
  }
}
Decisions: Legit security alert (hard) 100% right
{
  "phishing": {
    "true": 0.05,
    "false": 0.95
  },
  "action_needed": {
    "true": 0.1,
    "false": 0.9
  }
}
Decisions: Meeting slot (hard) 0% right
{
  "slot": {
    "A": 0.0,
    "B": 0.0,
    "C": 0.8,
    "D": 0.2
  },
  "raj_last": {
    "true": 0.0,
    "false": 1.0
  }
}
Decisions: Review mixed (hard) 100% right
{
  "hardware": {
    "true": 0.9,
    "false": 0.1
  },
  "support": {
    "true": 0.95,
    "false": 0.05
  }
}
Decisions: Support checkout down 100% right
```json
{
  "department": {
    "billing": 0.0,
    "technical": 0.95,
    "account": 0.0,
    "shipping": 0.0,
    "sales": 0.05
  },
  "urgency": {
    "0": 0.0,
    "1": 0.0,
    "2": 0.0,
    "3": 1.0
  },
  "outage": {
    "true": 0.95,
    "false": 0.05
  }
}
```
Decisions: Refund wrong plan 100% right
{
  "department": {
    "billing": 0.9,
    "technical": 0.0,
    "account": 0.05,
    "shipping": 0.0,
    "sales": 0.05
  },
  "refund": {
    "true": 0.95,
    "false": 0.05
  },
  "tone": {
    "frustrated": 0.05,
    "calm": 0.95
  }
}
Decisions: Moderation doxxing 100% right
{
  "policy": {
    "none": 0.0,
    "harassment": 0.95,
    "hate": 0.0,
    "spam": 0.0,
    "self_harm": 0.05
  },
  "personal_info": {
    "true": 0.95,
    "false": 0.05
  }
}
Decisions: Route calendar 100% right
{
  "tool": {
    "web_search": 0.0,
    "calculator": 0.0,
    "calendar": 0.9,
    "email": 0.1,
    "none": 0.0
  },
  "confirm": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Doc invoice missing due 100% right
{
  "doc_type": {
    "invoice": 1.0,
    "resume": 0.0,
    "contract": 0.0,
    "bank_statement": 0.0,
    "other": 0.0
  },
  "missing_due_date": {
    "true": 0.9,
    "false": 0.1
  }
}
Decisions: Phishing paypal 100% right
{
  "phishing": {
    "true": 0.95,
    "false": 0.05
  },
  "risk": {
    "0": 0.01,
    "1": 0.04,
    "2": 0.05,
    "3": 0.9
  }
}
Decisions: Pii ssn email 100% right
{
  "data_kind": {
    "none": 0.0,
    "contact": 0.2,
    "financial": 0.0,
    "government_id": 0.8,
    "health": 0.0
  },
  "sensitive": {
    "true": 0.8,
    "false": 0.2
  }
}
Decisions: Review mixed 100% right
{
  "sentiment": {
    "positive": 0.0,
    "neutral": 0.0,
    "negative": 1.0
  },
  "defect": {
    "true": 1.0,
    "false": 0.0
  },
  "recommend": {
    "true": 0.0,
    "false": 1.0
  }
}
Documents: Saas escalator (hard) 30% right
{
  "year2_price_per_seat_month": 46.83,
  "year3_price_per_seat_month": 46.83,
  "year1_invoice": 62640.00,
  "year2_invoice": 138624.00,
  "addon_months_billed": 6,
  "addon_invoice": 42096.00,
  "year3_invoice": 138624.00,
  "year3_discount_percent": 15.00,
  "total_contract_value": 382984.00,
  "contract_end_date": "2027-02-28"
}
Documents: Expense thread 88% right
{
  "employee_id": "EMP-20417",
  "destination_city": "Lisbon",
  "trip_start": "2025-02-24",
  "trip_end": "2025-02-27",
  "approved_items": [
    {
      "date": "2025-02-24",
      "category": "airfare",
      "amount_usd": 1184.6
    },
    {
      "date": "2025-02-24",
      "category": "ground_transport",
      "amount_usd": 38.88
    },
    {
      "date": "2025-02-25",
      "category": "meals",
      "amount_usd": 229.39
    },
    {
      "date": "2025-02-26",
      "category": "lodging",
      "amount_usd": 466.56
    },
    {
      "date": "2025-02-27",
      "category": "ground_transport",
      "amount_usd": 44.82
    }
  ],
  "rejected_item_count": 1,
  "per_diem_days": 2,
  "per_diem_usd": 65.00,
  "total_reimbursable_usd": 2134.25,
  "approver_email": "priya.raman@corvane.com"
}
Documents: Lease amendment 92% right
{
  "tenants": ["Marcus Lin", "Sofia Lin"],
  "landlord": "Ridgeline Property Group LLC",
  "zip": "97205",
  "lease_end": "2025-11-30",
  "original_monthly_rent": 2150.00,
  "monthly_rent_from_2025_06_01": 2236.00,
  "late_fee_from_2025_06_01": 111.80,
  "security_deposit": 2150.00,
  "total_pet_deposits": 800.00,
  "total_monthly_payment_july_2025": 2306.00,
  "move_in_payment": 2950.00
}
Documents: Ticket SLA 82% right
{
  "ticket_id": "48213",
  "account_id": "ACC-7731",
  "open_issue": "inventory_sync",
  "resolved_issues": ["billing_address"],
  "affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
  "priority": "P2",
  "sla_due_local": "2025-09-15T11:30",
  "sla_due_utc": "2025-09-15T16:30:00Z",
  "reissued_invoice": "INV-2025-0812"
}
Documents: Sales footnotes 78% right
{
  "q3_total_usd": 15246000,
  "q2_total_usd": 14464000,
  "q2_central_originally_reported_usd": 3047000,
  "q2_to_q3_change_pct": 5.4,
  "top_region_q3": "East",
  "fastest_growing_region_q1_to_q3": "International",
  "regions_declining_q2_to_q3": ["East"],
  "international_q3_organic_usd": 1731000,
  "west_excluding_mountain_q3_usd": 4201000
}

Size: 111B parameters. First tested OCT 11.

Models that scored about the same

Comments

Sign in with GitHub to comment. Spam and abuse are hidden automatically.