Grok 4.5 scored 9 out of 10 in a head-to-head game-building test this week, outpacing Anthropic's Fable 5 at 7.5 and OpenAI's GPT-5.5 at 7. The evaluation, run by Command Code using the /design command with an identical prompt across all three models, asked each one to build a game from scratch under the same conditions.
The result caught a lot of people off guard. Fable 5 leads on most general intelligence benchmarks and has posted strong numbers in agentic knowledge work. But in this test, Grok 4.5 produced something that looked and felt like a finished mobile game, not a rough prototype. Fable 5 and GPT-5.5 both delivered functional output, but reviewers flagged both as feeling rushed, the kind of thing you'd need to significantly rework before shipping. Grok 4.5 already leads on cost efficiency by a wide margin, so topping the quality ranking in a practical build test on top of that changes its position in the competitive picture considerably.
What GPT-5.6 Changes
Before the game-build comparison had even finished circulating, OpenAI launched GPT-5.6. It comes in three variants: Sol, Terra, and Luna. Sol is the one to watch on benchmarks. It gained roughly 5 points over GPT-5.5 on the CritPt physics benchmark and beat Fable 5 by around 4 points on the same test. That's a meaningful gap in a domain where incremental improvements usually dominate the headlines.
Alongside GPT-5.6, OpenAI rolled out an autonomous agent called ChatGPT Work, available across ChatGPT, Codex, and the API. The agent is designed to handle multi-step tasks without human prompting at each stage, which puts it squarely in the territory Anthropic has been building toward with Fable 5's agentic capabilities.
- Grok 4.5: 9/10 in game-build test, cost-efficiency leader
- Fable 5: 7.5/10, strong on general intelligence benchmarks
- GPT-5.5: 7/10 in the same test
- GPT-5.6 Sol: +5 points over GPT-5.5 on CritPt physics, beats Fable 5 by ~4 points
The pace here is the real story. These three comparisons, Grok 4.5 vs. Fable 5 vs. GPT-5.5, then GPT-5.6 dropping before that evaluation had settled, happened inside a single week. Labs are shipping faster than independent reviewers can fully process what's already out. GPT-5.6 Sol leads the CritPt physics benchmark right now, but the next release is probably already in testing somewhere.
This article is for informational purposes only and does not constitute financial or investment advice.



