Channels
The local endpoint records 68.3% of the available signal against 99.5% from the strongest paid channel, a gap of 31.2 points. Each tick is one task, full height a clean pass; the short and missing ticks are where a channel dropped out.
qwen3.8-max
gpt-5.6-luna
glm-5.2
deepseek-v4-pro
qwen3.8-27b-local
clean passpartialzerohard tier
Not recorded this run: fable. Shown nowhere below rather than counted as zero.
Core against hard
The hard tier was added to separate channels the core tier could not. It did that for one of them: qwen3.8-27b-local gives up 0.15 moving to hard, while 5 of the paid channels hold their core score to within 0.02. The hard tasks are harder — they just are not hard for frontier models.
rightmost column: change from core to hard, so a negative number is a channel losing ground on the harder tier
By task type
Mean score per channel. Build and fix tasks execute hidden tests; diagnose and review are graded blind against written ground truth.
| channel | build | fix | diagnose | review | assert |
|---|---|---|---|---|---|
| grok-4.6 | 0.99 | 1.00 | 1.00 | 0.98 | 1.00 |
| qwen3.8-max | 0.96 | 1.00 | 1.00 | 1.00 | 1.00 |
| gpt-5.6-luna | 0.94 | 1.00 | 0.92 | 1.00 | 1.00 |
| glm-5.2 | 0.99 | 0.75 | 0.92 | 1.00 | 1.00 |
| deepseek-v4-pro | 0.98 | 1.00 | 0.83 | 0.97 | 1.00 |
| qwen3.8-27b-local | 0.54 | 0.75 | 0.75 | 0.88 | 0.96 |
Where qwen3.8-27b-local loses
19 of 37 tasks scored below a clean pass, 6 of them at zero. A zero is rarely diffuse weakness: it is usually one small defect that costs the whole task.
| task | type | tier | local | field | what happened |
|---|---|---|---|---|---|
| B10 | build | core | 0.00 | 1.00 | 0/9 hidden tests |
| B2 | build | core | 0.00 | 1.00 | 0/12 hidden tests |
| B8 | build | core | 0.00 | 1.00 | 0/16 hidden tests |
| H2 | build | hard | 0.00 | 0.97 | 0/14 hidden tests |
| H6 | build | hard | 0.00 | 1.00 | 0/15 hidden tests |
| F4 | fix | core | 0.00 | 1.00 | 0/6 hidden tests |
| H4 | build | hard | 0.07 | 0.90 | 1/14 hidden tests |
| B6 | build | core | 0.10 | 0.72 | 1/10 hidden tests |
| D1 | diagnose | core | 0.50 | 1.00 | right subsystem, wrong mechanism |
| D3 | diagnose | core | 0.50 | 1.00 | right subsystem, wrong mechanism |
| D4 | diagnose | core | 0.50 | 0.60 | right subsystem, wrong mechanism |
| D7 | diagnose | core | 0.50 | 1.00 | right subsystem, wrong mechanism |
| HD2 | diagnose | hard | 0.50 | 0.90 | right subsystem, wrong mechanism |
| HD4 | diagnose | hard | 0.50 | 1.00 | right subsystem, wrong mechanism |
| H3 | build | hard | 0.57 | 1.00 | 8/14 hidden tests |
| HR1 | review | hard | 0.73 | 0.99 | found=['P1', 'P3', 'P5', 'P7'] fp=0 |
| R1 | review | core | 0.92 | 1.00 | found=['P1', 'P2', 'P3', 'P4', 'P5', 'P6'] fp=1 |
| I2 | assert | core | 0.93 | 1.00 | 13/14 hidden tests |
| B1 | build | core | 0.94 | 1.00 | 15/16 hidden tests |
Gap by type
Gap by tier
localfield mean
What the score costs
Quality against spend for the whole run. gpt-5.6-luna reaches 94.5% for $0.036; the local channel is free but sits 31.2 points under the best paid channel.
horizontal: run cost, log scale · vertical: % of available points
Every task
The full record. Task prompts and hidden tests stay unpublished: a task set that leaks stops measuring anything.
| task | type | tier | grok-4.6 | qwen3.8-max | gpt-5.6-luna | glm-5.2 | deepseek-v4-pro | qwen-local |
|---|---|---|---|---|---|---|---|---|
| B1 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.94 |
| B10 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.00 |
| B2 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.00 |
| B3 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| B4 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| B5 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| B6 | build | core | 0.90 | 0.70 | 0.10 | 0.90 | 1.00 | 0.10 |
| B7 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 0.86 | 1.00 |
| B8 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.00 |
| B9 | build | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| H1 | build | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| H2 | build | hard | 1.00 | 1.00 | 1.00 | 1.00 | 0.86 | 0.00 |
| H3 | build | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.57 |
| H4 | build | hard | 1.00 | 0.71 | 0.86 | 1.00 | 0.93 | 0.07 |
| H5 | build | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| H6 | build | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.00 |
| F1 | fix | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| F2 | fix | core | 1.00 | 1.00 | 1.00 | 0.00 | 1.00 | 1.00 |
| F3 | fix | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| F4 | fix | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.00 |
| D1 | diagnose | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.50 |
| D2 | diagnose | core | 1.00 | 1.00 | 1.00 | 1.00 | 0.50 | 1.00 |
| D3 | diagnose | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.50 |
| D4 | diagnose | core | 1.00 | 1.00 | 0.00 | 1.00 | 0.00 | 0.50 |
| D5 | diagnose | core | 1.00 | 1.00 | 1.00 | 0.50 | 0.50 | 1.00 |
| D6 | diagnose | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| D7 | diagnose | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.50 |
| D8 | diagnose | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| HD1 | diagnose | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| HD2 | diagnose | hard | 1.00 | 1.00 | 1.00 | 0.50 | 1.00 | 0.50 |
| HD3 | diagnose | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| HD4 | diagnose | hard | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.50 |
| R1 | review | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.92 |
| R2 | review | core | 1.00 | 1.00 | 1.00 | 1.00 | 0.91 | 1.00 |
| HR1 | review | hard | 0.93 | 1.00 | 1.00 | 1.00 | 1.00 | 0.73 |
| I1 | assert | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| I2 | assert | core | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.93 |