Raising the bar from correctness to quality #
Todayโs coding benchmarks have established that models can write correct code. But as AI-generated code becomes the dominant path to production, correctness is now table stakes. The question that we should be asking is: can models actually write good code?
Weโre excited to introduce FrontierCode, a benchmark that measures how well models can truly meet the standards of high-quality production codebases. What sets us apart:
Our benchmark provides the strongest available signal of a modelโs ability to write high-quality, maintainable code. We find that even todayโs most capable models struggle on this new standard.
Manually reviewed by Cognition researchers
First-ever benchmark measuring code quality