Companies are optimizing models for specific benchmarks Openai is optimizing for gpqa diamond and anthropic is optimizing for humanity last exam.
gpt 5.6 wins on gpqa and opus 5 wins on humanity last exam |
No comments yet
Companies are optimizing models for specific benchmarks Openai is optimizing for gpqa diamond and anthropic is optimizing for humanity last exam.
gpt 5.6 wins on gpqa and opus 5 wins on humanity last exam |