A coding benchmark of one
Coding models measured on tasks mined from one engineer’s own work: scripts written here, production incidents debugged here, defects found in real review. Executable tasks are scored by running hidden tests; the rest are graded blind against written ground truth. The question each run answers is whether a free model on a single home GPU can do this work, and what the paid alternatives cost.
Task prompts and hidden tests are deliberately unpublished. A benchmark whose questions leak stops measuring anything.