- GLM 4.5 Air is a free model from Zhipu AI that you can run on your own computer. In our tests it's decent, with trade-offs: 76 out of 100, #18 of 56.
- It solved 21 of 30 coding jobs and scored 82 on reading documents. On our hardest tasks it scored 52.
- Runs on a Mac with 96 GB.
Coding
Our coding test is 30 programming jobs, from small ones like reading time durations or cleaning up messy data to harder ones like a config-file parser or a double-entry ledger. We run each answer against tests the model never sees, and a job only counts if everything passes. GLM 4.5 Air got 21 of 30 right. The best local coders solved 29 of 30.
Reading documents
The second test hands the model things like an expense claim thread, a pay stub or an insurance statement, and asks for specific numbers and dates. Many questions need a bit of math, or noticing a correction further down the email. GLM 4.5 Air scored 82; the best model scored 100.
| Test | Score | Public questions | Secret questions |
|---|---|---|---|
| Coding | 70 | 71 | 70 |
| Reading documents | 82 | 82 | 82 |
| Decisions | 91 | 98 | 88 |
On the 18 hardest tasks (included in the scores above) it scored 52. This number separates the top models.
We tested the full-size version online. The smaller versions you download usually score a little lower.
Can your computer run it?
You can download GLM 4.5 Air in several sizes. Smaller versions fit on more computers and lose a little quality, so pick the biggest one that fits.
| Your computer | Runs it? | Version to download |
|---|---|---|
| Graphics card, 8 GB | No | – |
| Graphics card, 12 GB | No | – |
| Graphics card, 16 GB | No | – |
| Graphics card, 24 GB | No | – |
| Graphics card, 32 GB | No | – |
| Mac, 16 GB | No | – |
| Mac, 24 GB | No | – |
| Mac, 32 GB | No | – |
| Mac, 48 GB | No | – |
| Mac, 64 GB | No | – |
| Mac, 96 GB | Yes | Q4 |
| Mac, 128 GB | Yes | Q6 |
For normal-length chats. Long documents need more memory. Q4 or Q8? · How we worked this out
How to try it
Search for "GLM 4.5 Air GGUF" in LM Studio or on Hugging Face and pick the version that fits your computer. New to this? Start with our Ollama guide.
Want the strongest all-rounder instead? Qwen3.6 27B scored 97.
Its actual answers
What GLM 4.5 Air wrote for our public questions, word for word.
Coding: INI parser (hard) Didn't work
To solve this problem, we need to parse an INI text file according to specific rules and return a dictionary of dictionaries representing the sections and their key-value pairs. The solution involves several steps: parsing the text into sections and key-value pairs, handling continuations, comments, and quoted values, and finally resolving interpolations in the values.
### Approach
1. **Initialization**: We start by initializing a dictionary to hold all sections (including the special `DEFAULT` section) and variables to track the current section and the last processed key in that section.
2. **Line Processing**: We process each line of the input text:
- **Empty and Comment Lines**: Skip lines that are empty or start with `;` or `#` after optional whitespace.
- **Section Headers**: Lines enclosed in `[ ]` define a new section. The section name is stripped of surrounding whitespace.
- **Key-Value Lines**: For lines containing `=` or `:`, split into key and value. The key is stripped and lowercased. The value is processed to remove surrounding whitespace, inline comments, and quotes.
- **Continuation Lines**: Lines starting with whitespace are appended to the previous value with a newline character.
3. **Interpolation**: After parsing, we resolve interpolations in values. Interpolations can reference keys in the same section or other sections using `${key}` or `${section:key}`. This is done recursively, checking for circular references and missing keys.
4. **Result Construction**: The final result excludes the `DEFAULT` section but includes all other sections with their resolved key-value pairs.
### Solution Code
```python
def parse_ini(text: str) -> dict[str, dict[str, str]]:
sections = {}
current_section = None
last_key = None
lines = text.splitlines()
i = 0
n = len(lines)
while i < n:
line = lines[i]
stripped_line = line.strip()
if stripped_line == '' or stripped_line[0] in ';#':
i += 1
continue
if line.lstrip().startswith('[') and ']' in line:
start_index = line.index('[')
end_index = line.index(']', start_index+1)
section_name = line[start_index+1:end_index].strip()
current_section = section_name
sections[current_section] = {}
last_key = None
i += 1
continue
if current_section is None:
raise ValueError("Key before any section header")
if line.startswith(' ') or line.startswith('\t'):
if last_key is None:
raise ValueError("Continuation line without a previous key")
value = sections[current_section][last_key]
value += "\n" + line.strip()
sections[current_section][last_key] = value
i += 1
continue
pos_equal = line.find('=')
pos_colon = line.find(':')
if pos_equal == -1 and pos_colon == -1:
raise ValueError("Malformed line")
if pos_equal == -1:
pos = pos_colon
elif pos_colon == -1:
pos = pos_equal
else:
pos = min(pos_equal, pos_colon)
key_str = line[:pos]
key = key_str.strip().lower()
value_str = line[pos+1:].lstrip()
if len(value_str) >= 2 and value_str[0] == '"' and value_str[-1] == '"':
value = value_str[1:-1]
else:
indices = []
idx_semicolon = value_str.find(' ;')
idx_hash = value_str.find(' #')
if idx_semicolon != -1:
indices.append(idx_semicolon)
if idx_hash != -1:
indices.append(idx_hash)
if indices:
idx = min(indices)
value_str = value_str[:idx]
value = value_str.strip()
j = i + 1
current_value = value
while j < n:
next_line = lines[j]
if next_line.startswith(' ') or next_line.startswith('\t'):
current_value += "\n" + next_line.strip()
j += 1
else:
break
sections[current_section][key] = current_value
last_key = key
i = j
def get_value(section_name, key_name, current_section, sections):
if section_name == current_section:
if current_section in sections and key_name in sections[current_section]:
return sections[current_section][key_name]
if 'DEFAULT' in sections and key_name in sections['DEFAULT']:
return sections['DEFAULT'][key_name]
raise KeyError(f"Key {key_name} not found in section {current_section} or DEFAULT")
else:
if section_name in sections and key_name in sections[section_name]:
return sections[section_name][key_name]
raise KeyError(f"Key {key_name} not found in section {section_name}")
def resolve_value(value, current_section, sections, visited=None):
if visited is None:
visited = set()
start = value.find('${')
if start == -1:
return value
end = value.find('}', start+2)
if end == -1:
return value
inner = value[start+2:end]
if ':' in inner:
section_name, key_name = inner.split(':', 1)
else:
section_name = current_section
key_name = inner
if (section_name, key_name) in visited:
raise ValueError(f"Circular reference: {section_name}:{key_name}")
visited.add((section_name, key_name))
try:
base_value = get_value(section_name, key_name, current_section, sections)
resolved_base_value = resolve_value(base_value, current_section, sections, visited)
except (KeyError, ValueError) as e:
visited.remove((section_name, key_name))
raise e
visited.remove((section_name, key_name))
new_value = value[:start] + resolved_base_value + value[end+1:]
return resolve_value(new_value, current_section, sections, visited)
for section_name, section in sections.items():
for key, value in section.items():
sections[section_name][key] = resolve_value(value, section_name, sections, set())
result = {}
for section_name, section in sections.items():
if section_name != 'DEFAULT':
result[section_name] = section
return result
```
### Explanation
1. **Initialization**: The `sections` dictionary stores all parsed sections, including `DEFAULT`. `current_section` tracks the active section, and `last_key` remembers the last processed key in the current section.
2. **Line Processing**:
- **Section Headers**: Lines starting with `[ ]` define new sections. The section name is extracted and stored in `sections`.
- **Key-Value Lines**: Lines containing `=` or `:` are split into key and value. The key is stripped and lowercased. The value is processed to remove quotes and inline comments.
- **Continuation Lines**: Lines starting with whitespace are appended to the previous value with a newline.
3. **Interpolation**: The `resolve_value` function recursively resolves `${key}` and `${section:key}` interpolations. It checks for circular references and missing keys using a `visited` set.
4. **Result Construction**: The final result excludes the `DEFAULT` section but includes all other sections with their resolved key-value pairs after interpolation.
This approach efficiently parses the INI text according to the specified rules, handling comments, continuations, quotes, and interpolations while ensuring correctness and error handling. Coding: TTL LRU (hard) Worked
To solve this problem, we need to implement a TTL (Time-To-Live) cache with LRU (Least Recently Used) eviction policy. The cache should store key-value pairs where each entry has a time-to-live (TTL) after which it expires. The cache should efficiently handle insertions, retrievals, and removals of expired entries while maintaining the LRU order for evictions when the cache exceeds its capacity.
### Approach
1. **Data Structures**:
- **Doubly Linked List**: To maintain the LRU order. The most recently accessed items are near the head, and the least recently accessed items are near the tail.
- **Dictionary (Hash Map)**: To map keys to their corresponding nodes in the linked list for O(1) access time.
2. **Key Operations**:
- **Put Operation**:
- If the key exists, update its value and expiry time, then move it to the head (most recently used).
- If the key is new:
- First, remove all expired entries.
- If the cache is still at or above capacity, evict the least recently used item (tail of the linked list).
- Add the new key-value pair with the current time plus TTL as the expiry time and move it to the head.
- **Get Operation**:
- If the key exists and is not expired, move it to the head (most recently used) and return its value.
- If the key is expired or does not exist, remove it (if expired), increment the miss count, and return the default value.
- **Remove Expired Entries**: Iterate through all entries to remove those that have expired based on the current time from the clock.
- **Length Calculation**: Count all non-expired entries by checking each entry's expiry against the current time.
- **Statistics Tracking**: Maintain counts for hits, misses, evictions, and expirations.
3. **Edge Cases**:
- **Zero Capacity**: The cache does not store any entries; all operations are no-ops or return defaults.
- **Expired Entries**: These are removed during `get` or `put` operations, and their removal counts as an expiration.
### Solution Code
```python
class TTLCache:
class Node:
__slots__ = ('key', 'value', 'expiry', 'prev', 'next')
def __init__(self, key, value, expiry):
self.key = key
self.value = value
self.expiry = expiry
self.prev = None
self.next = None
def __init__(self, capacity: int, ttl: float, clock):
self.capacity = capacity
self.ttl = ttl
self.clock = clock
self.cache = {}
self.head = None
self.tail = None
self.hits = 0
self.misses = 0
self.evictions = 0
self.expirations = 0
def _add_node_at_head(self, node):
node.prev = None
node.next = self.head
if self.head is not None:
self.head.prev = node
self.head = node
if self.tail is None:
self.tail = node
def _remove_node(self, node):
if node.prev is not None:
node.prev.next = node.next
else:
self.head = node.next
if node.next is not None:
node.next.prev = node.prev
else:
self.tail = node.prev
def _move_to_head(self, node):
self._remove_node(node)
self._add_node_at_head(node)
def _remove_all_expired(self):
current_time = self.clock()
expired_keys = []
for key, node in self.cache.items():
if current_time >= node.expiry:
expired_keys.append(key)
for key in expired_keys:
node = self.cache[key]
self._remove_node(node)
del self.cache[key]
self.expirations += 1
def put(self, key, value):
if self.capacity == 0:
return
if key in self.cache:
node = self.cache[key]
node.value = value
node.expiry = self.clock() + self.ttl
self._move_to_head(node)
else:
self._remove_all_expired()
if len(self.cache) >= self.capacity:
if self.tail is not None:
tail_node = self.tail
self._remove_node(tail_node)
del self.cache[tail_node.key]
self.evictions += 1
current_time = self.clock()
node = self.Node(key, value, current_time + self.ttl)
self.cache[key] = node
self._add_node_at_head(node)
def get(self, key, default=None):
if self.capacity == 0:
self.misses += 1
return default
if key in self.cache:
node = self.cache[key]
current_time = self.clock()
if current_time < node.expiry:
self._move_to_head(node)
self.hits += 1
return node.value
else:
self._remove_node(node)
del self.cache[key]
self.expirations += 1
self.misses += 1
return default
else:
self.misses += 1
return default
def __len__(self):
if self.capacity == 0:
return 0
current_time = self.clock()
count = 0
for node in self.cache.values():
if current_time < node.expiry:
count += 1
return count
def stats(self):
return {
'hits': self.hits,
'misses': self.misses,
'evictions': self.evictions,
'expirations': self.expirations
}
```
### Explanation
- **Initialization**: The cache is initialized with a given capacity, TTL duration, and a clock function to get the current time.
- **Linked List Operations**: The `_add_node_at_head`, `_remove_node`, and `_move_to_head` methods manage the doubly linked list to maintain LRU order.
- **Put Operation**: Updates existing entries or inserts new ones. For new entries, it first removes expired entries and evicts the least recently used item if necessary before adding the new entry.
- **Get Operation**: Retrieves the value if the key exists and is not expired, updates its recency, and returns it. If the key is expired or missing, it removes the expired entry (if applicable) and returns the default value.
- **Length Calculation**: Counts non-expired entries by checking each entry's expiry against the current time.
- **Statistics**: Tracks the number of hits, misses, evictions, and expirations to monitor cache performance.
This approach efficiently combines LRU eviction with TTL expiration, ensuring optimal performance for both insertions and retrievals while maintaining correct statistics. Coding: Parse duration Didn't work
Coding: Summarize ranges Worked
To solve this problem, we need to process a list of integers by removing duplicates, sorting the remaining numbers, and then formatting consecutive sequences. Specifically, runs of three or more consecutive integers should be collapsed into a range format (e.g., "a..b"), while shorter runs (one or two numbers) should be listed individually. The results should be joined into a single string separated by commas without any spaces.
### Approach
1. **Remove Duplicates and Sort**: First, we convert the input list into a set to eliminate duplicates, then sort the resulting set to ensure the numbers are in ascending order.
2. **Identify Consecutive Runs**: Traverse the sorted list to identify sequences of consecutive integers. For each sequence, determine its start and end points.
3. **Format Runs**:
- For sequences of three or more consecutive integers, format them as "start..end".
- For sequences of one or two integers, list each number individually.
4. **Combine Results**: Join all formatted parts into a single string separated by commas.
### Solution Code
```python
def summarize_ranges(nums: list[int]) -> str:
if not nums:
return ""
sorted_nums = sorted(set(nums))
n = len(sorted_nums)
i = 0
result = []
while i < n:
j = i
while j < n - 1 and sorted_nums[j + 1] == sorted_nums[j] + 1:
j += 1
start = sorted_nums[i]
end = sorted_nums[j]
if end - start >= 2:
result.append(f"{start}..{end}")
else:
for num in range(start, end + 1):
result.append(str(num))
i = j + 1
return ",".join(result)
```
### Explanation
1. **Initial Checks**: The function first checks if the input list is empty and returns an empty string if true.
2. **Processing the List**: The input list is converted to a set to remove duplicates, then sorted to ensure ascending order.
3. **Traversal and Grouping**: The sorted list is traversed to find sequences of consecutive integers. For each sequence:
- The start and end of the sequence are identified.
- If the sequence length is three or more (determined by `end - start >= 2`), it is formatted as "start..end".
- Shorter sequences (one or two numbers) are added to the result list as individual strings.
4. **Result Construction**: The formatted parts are joined into a single string separated by commas, which is returned as the final result.
This approach efficiently processes the input list by leveraging sorting and a linear traversal to identify and format consecutive sequences, ensuring optimal performance and correctness. Coding: Top customers Worked
To solve this problem, we need to process a list of transaction records, aggregate the total amounts per customer, and then return the top `n` customers based on their total amounts. The solution involves several steps including data cleaning, aggregation, sorting, and formatting the results.
### Approach
1. **Data Cleaning**: For each transaction record, we first clean and validate the customer name and the amount. The customer name is stripped of leading/trailing whitespace, and empty names are skipped. The amount string is processed to remove dollar signs, commas, and spaces before being converted to a float. If the amount string cannot be converted, the record is skipped.
2. **Aggregation**: We use a dictionary to aggregate the total amounts per customer. The dictionary keys are case-insensitive versions of the customer names (to handle case variations), while the values store the original customer name (as first encountered) and the cumulative total amount.
3. **Rounding and Sorting**: After processing all records, we round each customer's total amount to two decimal places. The customers are then sorted in descending order of their total amounts. If two customers have the same total, they are sorted alphabetically by name.
4. **Top N Selection**: Finally, we return the top `n` customers from the sorted list. If there are fewer than `n` customers, all customers are returned.
### Solution Code
```python
def top_customers(rows: list[dict], n: int) -> list[tuple[str, float]]:
customer_data = {}
for row in rows:
customer = row.get('customer', '').strip()
amount_str = row.get('amount', '').strip()
if not customer or not amount_str:
continue
clean_amount = amount_str.replace('$', '').replace(',', '').replace(' ', '')
try:
amount = float(clean_amount)
except ValueError:
continue
customer_lower = customer.lower()
if customer_lower not in customer_data:
customer_data[customer_lower] = (customer, amount)
else:
original_name, current_total = customer_data[customer_lower]
customer_data[customer_lower] = (original_name, current_total + amount)
result_list = []
for original_name, total in customer_data.values():
rounded_total = round(total, 2)
result_list.append((original_name, rounded_total))
sorted_list = sorted(result_list, key=lambda x: (-x[1], x[0]))
return sorted_list[:n]
```
### Explanation
1. **Data Cleaning**: The customer name is stripped of any leading or trailing whitespace. The amount string is processed to remove any dollar signs, commas, and spaces, making it suitable for conversion to a float. If the amount string is invalid, the record is skipped.
2. **Aggregation**: A dictionary (`customer_data`) is used where keys are the lowercase versions of customer names to ensure case-insensitive matching. The first occurrence of a customer name (original case) is stored along with their cumulative total amount. Subsequent amounts for the same customer are added to this total.
3. **Rounding and Sorting**: Each customer's total amount is rounded to two decimal places. The customers are sorted based on their total amounts in descending order. If two customers have the same total, they are sorted alphabetically by name in ascending order.
4. **Top N Selection**: The sorted list of customers is then sliced to return the top `n` customers. If the list contains fewer than `n` customers, the entire list is returned.
This approach efficiently processes the transaction data, handles case variations in customer names, and ensures the results are formatted and sorted as required. Coding: Slugify Worked
To solve this problem, we need to create a URL slug from a given title string, adhering to specific formatting rules. The slug must be suitable for URLs, ensuring it is ASCII-compatible, lowercase, and uses hyphens to separate words. Additionally, the slug should not exceed a specified maximum length without cutting words unless necessary.
### Approach
1. **Transliterate and Convert to ASCII**: Convert accented characters to their ASCII equivalents (e.g., 'é' becomes 'e') and remove any non-ASCII characters. This is done using Unicode normalization (NFKD) followed by encoding to ASCII while ignoring non-ASCII characters.
2. **Lowercase the String**: Convert the entire string to lowercase to maintain consistency.
3. **Replace Non-Alphanumeric Characters**: Replace sequences of characters that are not letters or digits with a single hyphen. Also, remove any leading or trailing hyphens to ensure the slug starts and ends with alphanumeric characters.
4. **Truncate if Necessary**: If the resulting slug exceeds the specified maximum length, truncate it without cutting words. This involves:
- Splitting the slug into words separated by hyphens.
- Building the longest possible prefix of whole words (each word is a sequence of alphanumeric characters) that fits within the maximum length.
- If the first word alone exceeds the maximum length, hard-cut it to the maximum length characters.
5. **Ensure No Trailing Hyphen**: The final slug should not end with a hyphen, which is inherently handled by the truncation process.
### Solution Code
```python
import unicodedata
import re
def slugify(title: str, max_len: int = 50) -> str:
# Step 1: Normalize the string to decompose characters and remove non-ASCII
normalized = unicodedata.normalize('NFKD', title)
ascii_title = normalized.encode('ascii', 'ignore').decode('ascii')
# Step 2: Convert to lowercase
lower_title = ascii_title.lower()
# Step 3: Replace non-alphanumeric runs with a single hyphen and strip hyphens
slug = re.sub(r'[^a-z0-9]+', '-', lower_title)
slug = slug.strip('-')
# If the slug is empty after processing, return empty string
if not slug:
return ''
# Step 4: Truncate the slug if it exceeds max_len
if len(slug) <= max_len:
return slug
words = slug.split('-')
result = []
current_length = 0
for word in words:
if not result:
if len(word) <= max_len:
result.append(word)
current_length = len(word)
else:
result.append(word[:max_len])
current_length = max_len
break
else:
required_length = current_length + 1 + len(word)
if required_length <= max_len:
result.append(word)
current_length = required_length
else:
break
return '-'.join(result)
```
### Explanation
1. **Normalization and ASCII Conversion**: The input string is normalized to decompose accented characters (e.g., 'é' becomes 'e' followed by a combining acute accent). Non-ASCII characters are then removed by encoding to ASCII and ignoring any characters that cannot be represented.
2. **Lowercasing**: The entire string is converted to lowercase to ensure uniformity.
3. **Hyphen Replacement**: Sequences of non-alphanumeric characters are replaced with a single hyphen. Leading and trailing hyphens are stripped to ensure the slug starts and ends with alphanumeric characters.
4. **Truncation**: The slug is split into words separated by hyphens. The algorithm then builds the longest possible prefix of whole words that fits within the maximum length. If the first word exceeds the maximum length, it is truncated to the specified length. Subsequent words are only added if they fit within the remaining length when combined with existing words and hyphens.
5. **Result Construction**: The selected words are joined with hyphens to form the final slug, ensuring no trailing hyphen and adherence to the maximum length constraint.
This approach efficiently processes the input string into a URL-friendly slug while adhering to all specified constraints. Coding: Token bucket Worked
To solve this problem, we need to implement a token bucket rate limiter that refills tokens continuously based on elapsed time. The bucket starts full with a specified capacity and refills tokens at a given rate per second. The `allow` method checks if a certain number of tokens (cost) are available and subtracts them if possible, while the `tokens` property returns the current token count after refilling.
### Approach
1. **Initialization**: The token bucket is initialized with a maximum capacity, a refill rate (tokens per second), and a clock function to get the current time. The bucket starts full, and the last update time is set to the current time.
2. **Refilling Mechanism**: The `_refill` method calculates the elapsed time since the last update. It then adds the new tokens based on the elapsed time and the refill rate, ensuring the token count does not exceed the capacity. The last update time is then set to the current time.
3. **Allow Method**: The `allow` method first refills the bucket. If the available tokens are sufficient to cover the cost, it subtracts the cost and returns `True`; otherwise, it returns `False` without modifying the token count.
4. **Tokens Property**: The `tokens` property refills the bucket and returns the current token count. This ensures that any access to the token count reflects the most up-to-date value after refilling.
### Solution Code
```python
class TokenBucket:
def __init__(self, capacity: float, refill_per_sec: float, clock: callable):
self.capacity = capacity
self.refill_per_sec = refill_per_sec
self.clock = clock
self._tokens = capacity
self.last_update = clock()
def _refill(self):
current_time = self.clock()
elapsed = current_time - self.last_update
if elapsed < 0:
elapsed = 0
new_tokens = self._tokens + elapsed * self.refill_per_sec
self._tokens = min(new_tokens, self.capacity)
self.last_update = current_time
def allow(self, cost: float = 1) -> bool:
self._refill()
if self._tokens >= cost:
self._tokens -= cost
return True
return False
@property
def tokens(self) -> float:
self._refill()
return self._tokens
```
### Explanation
- **Initialization**: The `__init__` method sets up the bucket with the given capacity, refill rate, and clock function. The initial token count is set to the capacity, and the last update time is initialized to the current time.
- **Refilling**: The `_refill` method calculates the time elapsed since the last update. It then computes the new tokens by multiplying the elapsed time by the refill rate and adds this to the current tokens, capping the result at the capacity. The last update time is then updated to the current time.
- **Allow Method**: The `allow` method first refills the bucket to update the token count. If the available tokens are at least the requested cost, the cost is deducted from the token count, and the method returns `True`; otherwise, it returns `False`.
- **Tokens Property**: The `tokens` property ensures the bucket is refilled to the current state before returning the token count. This provides a real-time view of the available tokens without modifying the bucket's state.
This approach efficiently manages the token bucket by continuously refilling tokens based on elapsed time and provides methods to check for available tokens and consume them when possible. Decisions: Refund window (hard) 100% right
```json
{
"outcome": {
"full_refund": 0.0,
"partial_refund": 0.0,
"store_credit": 1.0,
"nothing": 0.0
},
"defective": {
"true": 0.0,
"false": 1.0
}
}
``` Decisions: Incident matrix (hard) 100% right
```json
{
"severity": {
"0": 0.0,
"1": 0.95,
"2": 0.05,
"3": 0.0
},
"page": {
"true": 0.05,
"false": 0.95
}
}
``` Decisions: Tool followup (hard) 100% right
```json
{
"tool": {
"web_search": 0.01,
"calculator": 0.01,
"calendar": 0.98,
"email": 0.01,
"none": 0.01
},
"confirm": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Legit security alert (hard) 100% right
```json
{
"phishing": {
"true": 0.001,
"false": 0.999
},
"action_needed": {
"true": 0.001,
"false": 0.999
}
}
``` Decisions: Meeting slot (hard) 100% right
{
"slot": {
"A": 0.0,
"B": 1.0,
"C": 0.0,
"D": 0.0
},
"raj_last": {
"true": 1.0,
"false": 0.0
}
} Decisions: Review mixed (hard) 100% right
```json
{
"hardware": {
"true": 0.95,
"false": 0.05
},
"support": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Support checkout down 67% right
```json
{
"department": {
"billing": 0.95,
"technical": 0.04,
"account": 0.00,
"shipping": 0.00,
"sales": 0.01
},
"urgency": {
"0": 0.00,
"1": 0.00,
"2": 0.05,
"3": 0.95
},
"outage": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Refund wrong plan 100% right
```json
{
"department": {
"billing": 0.95,
"technical": 0.01,
"account": 0.01,
"shipping": 0.01,
"sales": 0.02
},
"refund": {
"true": 0.99,
"false": 0.01
},
"tone": {
"frustrated": 0.05,
"calm": 0.95
}
}
``` Decisions: Moderation doxxing 100% right
```json
{
"policy": {
"none": 0.0,
"harassment": 0.95,
"hate": 0.0,
"spam": 0.0,
"self_harm": 0.05
},
"personal_info": {
"true": 0.99,
"false": 0.01
}
}
``` Decisions: Route calendar 100% right
```json
{
"tool": {
"web_search": 0.0,
"calculator": 0.0,
"calendar": 1.0,
"email": 0.0,
"none": 0.0
},
"confirm": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Doc invoice missing due 100% right
```json
{
"doc_type": {
"invoice": 1.0,
"resume": 0.0,
"contract": 0.0,
"bank_statement": 0.0,
"other": 0.0
},
"missing_due_date": {
"true": 1.0,
"false": 0.0
}
}
``` Decisions: Phishing paypal 100% right
```json
{
"phishing": {
"true": 0.99,
"false": 0.01
},
"risk": {
"0": 0.01,
"1": 0.02,
"2": 0.07,
"3": 0.90
}
}
``` Decisions: Pii ssn email 100% right
```json
{
"data_kind": {
"none": 0,
"contact": 0.1,
"financial": 0,
"government_id": 0.9,
"health": 0
},
"sensitive": {
"true": 1,
"false": 0
}
}
``` Decisions: Review mixed 100% right
```json
{
"sentiment": {
"positive": 0.1,
"neutral": 0.1,
"negative": 0.8
},
"defect": {
"true": 0.95,
"false": 0.05
},
"recommend": {
"true": 0.05,
"false": 0.95
}
}
``` Documents: Saas escalator (hard) 20% right
{
"year2_price_per_seat_month": 46.53,
"year3_price_per_seat_month": 48.86,
"year1_invoice": 58320.00,
"year2_invoice": 60298.56,
"addon_months_billed": 960,
"addon_invoice": 39937.76,
"year3_invoice": 63322.56,
"year3_discount_percent": 10.00,
"total_contract_value": 221878.88,
"contract_end_date": "2027-02-28"
} Documents: Expense thread 100% right
```json
{
"employee_id": "EMP-20417",
"destination_city": "Lisbon",
"trip_start": "2025-02-24",
"trip_end": "2025-02-27",
"approved_items": [
{
"date": "2025-02-24",
"category": "airfare",
"amount_usd": 1184.60
},
{
"date": "2025-02-24",
"category": "ground_transport",
"amount_usd": 38.88
},
{
"date": "2025-02-25",
"category": "meals",
"amount_usd": 229.39
},
{
"date": "2025-02-26",
"category": "lodging",
"amount_usd": 466.56
},
{
"date": "2025-02-27",
"category": "ground_transport",
"amount_usd": 44.82
}
],
"rejected_item_count": 1,
"per_diem_days": 3,
"per_diem_usd": 195.00,
"total_reimbursable_usd": 2159.25,
"approver_email": "priya.raman@corvane.com"
}
``` Documents: Lease amendment 100% right
```json
{
"tenants": ["Marcus Lin", "Sofia Lin"],
"landlord": "Ridgeline Property Group LLC",
"zip": "97205",
"lease_end": "2025-11-30",
"original_monthly_rent": 2150.00,
"monthly_rent_from_2025_06_01": 2236.00,
"late_fee_from_2025_06_01": 111.80,
"security_deposit": 2150.00,
"total_pet_deposits": 800.00,
"total_monthly_payment_july_2025": 2306.00,
"move_in_payment": 4700.00
}
``` Documents: Ticket SLA 91% right
```json
{
"ticket_id": "48213",
"account_id": "ACC-7731",
"open_issue": "inventory_sync",
"resolved_issues": ["billing_address"],
"affected_orders": ["SO-99812", "SO-99820", "SO-99827"],
"priority": "P2",
"sla_due_local": "2025-09-15T15:30",
"sla_due_utc": "2025-09-15T21:30:00Z",
"reissued_invoice": "INV-2025-0812"
}
``` Documents: Sales footnotes 100% right
{
"q3_total_usd": 15346000,
"q2_total_usd": 14464000,
"q2_central_originally_reported_usd": 3047000,
"q2_to_q3_change_pct": 6.1,
"top_region_q3": "East",
"fastest_growing_region_q1_to_q3": "International",
"regions_declining_q2_to_q3": ["East"],
"international_q3_organic_usd": 1731000,
"west_excluding_mountain_q3_usd": 4201000
} Size: 110B parameters. First tested OCT 11.
Comments
Sign in with GitHub to comment. Spam and abuse are hidden automatically.