Kimi K3: second only to Fable 5 on AA-Briefcase(artificialanalysis.ai) |
Kimi K3: second only to Fable 5 on AA-Briefcase(artificialanalysis.ai) |
Back in 2025 it was common to test models in a different harnesses.
I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.
Why did it come out of fashion ?
EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
Specifically they use this harness: https://github.com/ArtificialAnalysis/Stirrup