Post

HN
Hacker News

FrontierCode

Raising the bar from correctness to quality #

Todayโ€™s coding benchmarks have established that models can write correct code. But as AI-generated code becomes the dominant path to production, correctness is now table stakes. The question that we should be asking is: can models actually write good code?

Weโ€™re excited to introduce FrontierCode, a benchmark that measures how well models can truly meet the standards of high-quality production codebases. What sets us apart:

Our benchmark provides the strongest available signal of a modelโ€™s ability to write high-quality, maintainable code. We find that even todayโ€™s most capable models struggle on this new standard.

Manually reviewed by Cognition researchers

First-ever benchmark measuring code quality