The comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
That’s a nice discussion. Some people say that with current model capabilities, the real differentiator is the harness. What are the best harnesses you guys are using?
This approach of not only producing the benchmark tasks, but also focusing on creating a data engine that will improve over time and produce up-to-date tasks that challenge the cutting-edge models is very interesting and valuable.
thorough work, good stuff.. it even runs a selection-bias analysis against their own benchmark and reports that some tasks that were disproportionately hard for a model. Rare to see a benchmark paper attack itself like that.