SB‑Bench

run
full-2026-08-16
tasks
37 11 hard
channels
6
spend
$5.58
recorded
2026-08-16 13:49

Channels

The local endpoint records 68.3% of the available signal against 99.5% from the strongest paid channel, a gap of 31.2 points. Each tick is one task, full height a clean pass; the short and missing ticks are where a channel dropped out.

grok-4.6

$1.061 · 40.3s p50 · 70.6 tok/s

99.5%

qwen3.8-max

$3.577 · 235.4s p50 · 48.9 tok/s

98.4%

gpt-5.6-luna

$0.036 · 10.7s p50 · 107.0 tok/s

94.5%

glm-5.2

$0.513 · 17.1s p50 · 84.9 tok/s

94.3%

deepseek-v4-pro

$0.396 · 59.0s p50 · 72.7 tok/s

93.4%

qwen3.8-27b-local

free · local · 10.9s p50 · 35.6 tok/s

68.3%
buildfixdiagnosereviewassert

clean passpartialzerohard tier

Not recorded this run: fable. Shown nowhere below rather than counted as zero.

Core against hard

The hard tier was added to separate channels the core tier could not. It did that for one of them: qwen3.8-27b-local gives up 0.15 moving to hard, while 5 of the paid channels hold their core score to within 0.02. The hard tasks are harder — they just are not hard for frontier models.

grok-4.6core1.00hard0.99-0.00
qwen3.8-maxcore0.99hard0.97-0.01
gpt-5.6-lunacore0.93hard0.99+0.06
glm-5.2core0.94hard0.95+0.02
deepseek-v4-procore0.91hard0.98+0.07
qwen3.8-27b-localcore0.73hard0.58-0.15

rightmost column: change from core to hard, so a negative number is a channel losing ground on the harder tier

By task type

Mean score per channel. Build and fix tasks execute hidden tests; diagnose and review are graded blind against written ground truth.

channelbuildfixdiagnosereviewassert
grok-4.60.991.001.000.981.00
qwen3.8-max0.961.001.001.001.00
gpt-5.6-luna0.941.000.921.001.00
glm-5.20.990.750.921.001.00
deepseek-v4-pro0.981.000.830.971.00
qwen3.8-27b-local0.540.750.750.880.96

Where qwen3.8-27b-local loses

19 of 37 tasks scored below a clean pass, 6 of them at zero. A zero is rarely diffuse weakness: it is usually one small defect that costs the whole task.

tasktypetierlocalfieldwhat happened
B10buildcore0.001.000/9 hidden tests
B2buildcore0.001.000/12 hidden tests
B8buildcore0.001.000/16 hidden tests
H2buildhard0.000.970/14 hidden tests
H6buildhard0.001.000/15 hidden tests
F4fixcore0.001.000/6 hidden tests
H4buildhard0.070.901/14 hidden tests
B6buildcore0.100.721/10 hidden tests
D1diagnosecore0.501.00right subsystem, wrong mechanism
D3diagnosecore0.501.00right subsystem, wrong mechanism
D4diagnosecore0.500.60right subsystem, wrong mechanism
D7diagnosecore0.501.00right subsystem, wrong mechanism
HD2diagnosehard0.500.90right subsystem, wrong mechanism
HD4diagnosehard0.501.00right subsystem, wrong mechanism
H3buildhard0.571.008/14 hidden tests
HR1reviewhard0.730.99found=['P1', 'P3', 'P5', 'P7'] fp=0
R1reviewcore0.921.00found=['P1', 'P2', 'P3', 'P4', 'P5', 'P6'] fp=1
I2assertcore0.931.0013/14 hidden tests
B1buildcore0.941.0015/16 hidden tests

Gap by type

build+0.43
fix+0.20
diagnose+0.18
review+0.11
assert+0.04

Gap by tier

core+0.23
hard+0.40

localfield mean

What the score costs

Quality against spend for the whole run. gpt-5.6-luna reaches 94.5% for $0.036; the local channel is free but sits 31.2 points under the best paid channel.

65707580859095100free$0.05$0.30$1$2grok-4.6qwen3.8-maxgpt-5.6-lunaglm-5.2deepseek-v4-proqwen3.8-27b-local

horizontal: run cost, log scale · vertical: % of available points

Every task

The full record. Task prompts and hidden tests stay unpublished: a task set that leaks stops measuring anything.

tasktypetiergrok-4.6qwen3.8-maxgpt-5.6-lunaglm-5.2deepseek-v4-proqwen-local
B1buildcore1.001.001.001.001.000.94
B10buildcore1.001.001.001.001.000.00
B2buildcore1.001.001.001.001.000.00
B3buildcore1.001.001.001.001.001.00
B4buildcore1.001.001.001.001.001.00
B5buildcore1.001.001.001.001.001.00
B6buildcore0.900.700.100.901.000.10
B7buildcore1.001.001.001.000.861.00
B8buildcore1.001.001.001.001.000.00
B9buildcore1.001.001.001.001.001.00
H1buildhard1.001.001.001.001.001.00
H2buildhard1.001.001.001.000.860.00
H3buildhard1.001.001.001.001.000.57
H4buildhard1.000.710.861.000.930.07
H5buildhard1.001.001.001.001.001.00
H6buildhard1.001.001.001.001.000.00
F1fixcore1.001.001.001.001.001.00
F2fixcore1.001.001.000.001.001.00
F3fixcore1.001.001.001.001.001.00
F4fixcore1.001.001.001.001.000.00
D1diagnosecore1.001.001.001.001.000.50
D2diagnosecore1.001.001.001.000.501.00
D3diagnosecore1.001.001.001.001.000.50
D4diagnosecore1.001.000.001.000.000.50
D5diagnosecore1.001.001.000.500.501.00
D6diagnosecore1.001.001.001.001.001.00
D7diagnosecore1.001.001.001.001.000.50
D8diagnosecore1.001.001.001.001.001.00
HD1diagnosehard1.001.001.001.001.001.00
HD2diagnosehard1.001.001.000.501.000.50
HD3diagnosehard1.001.001.001.001.001.00
HD4diagnosehard1.001.001.001.001.000.50
R1reviewcore1.001.001.001.001.000.92
R2reviewcore1.001.001.001.000.911.00
HR1reviewhard0.931.001.001.001.000.73
I1assertcore1.001.001.001.001.001.00
I2assertcore1.001.001.001.001.000.93