Scott Wu, CEO of Cognition, puts it plainly: "When every model aces the test, the test stops telling you anything useful." His company’s AI agent Devin is nearing flawless performance on the same benchmarks that once defined the competitive scene. Benchmarks that once separated the good from the great now fail to differentiate, as top-tier AI models like Devin surmount nearly every challenge.

Devin's rise has been remarkable. When it launched in early 2024, it managed only about 13% on SWE-Bench, a standard for evaluating AI coding skills. Fast forward to mid-2026, Devin is hitting around 90% on that original benchmark, and close to 80% on an enhanced version named SWE-Bench Pro. That leap shows how quickly AI models have matured, pushing the limits of traditional testing methods. Cognition has responded by creating proprietary metrics such as FrontierCode 1.1, alongside internal tests crafted to mimic the messy, less predictable nature of actual software engineering tasks rather than the neat problems that traditional benchmarks favor.

Beyond scores, Wu warns that some common evaluation metrics, like token efficiency, can misrepresent the true output quality of AI coding systems. Such nuances matter especially given Cognition's recent milestones. The company raised $1 billion in funding this May, valuing itself at $26 billion less than three years after its founding in late 2023. Impressively, Devin now contributes approximately 89% of the committed code written by Cognition's engineering team. To strengthen its position amid a growing market for autonomous coding tools, Cognition also acquired a competitor, Windsurf, signaling tough rivalry ahead.

The evolution of AI benchmarks is a vivid example of the challenges in accurately gauging next-generation technologies. As these tools entwine more deeply with real-world productivity, companies like Cognition emphasize outcomes over numbers on a leaderboard shifting the conversation to what AI actually delivers in practice.