
Quick Summary
ChatGPT 3 defeated Grok 4 in a unique AI chess tournament for general-purpose LLMs, with Gemini taking third place. While Grok appeared dominant early on, repeated blunders cost it the crown, showcasing how strategic focus still matters, even for highly capable AI.
How ChatGPT secured the AI chess crown
OpenAI’s ChatGPT o3 model has emerged as the strongest AI chess player in a recent three-day tournament, beating Elon Musk’s xAI Grok 4 in a decisive final match. The event, aimed at testing the chess capabilities of general-purpose large language models (LLMs) rather than specialized chess engines, also saw Google’s Gemini clinch third place after defeating another OpenAI model.
The final match was a stark turnaround from earlier rounds, where Grok appeared unstoppable. According to Pedro Pinhata of Chess.com, Grok’s dominance “up until the semifinals” suggested a likely win, until it faltered under pressure.
Why Grok stumbled in the finals
Chess grandmaster Hikaru Nakamura, streaming commentary on the event, called Grok’s performance “blundering” and “unrecognizable” compared to its earlier form. Grok reportedly made multiple critical errors, including repeated losses of its queen, a fundamental mistake at any competitive level.
Elon Musk had downplayed Grok’s chess focus before the final, saying its earlier victories were a “side effect” of the team’s minimal effort in chess-specific training. However, the final showed that without targeted optimization, even advanced LLMs can struggle under strategic stress.
A different kind of chess competition
Unlike traditional computer chess matches, where engines like Stockfish or AlphaZero dominate, this tournament was unique. It exclusively featured general-purpose AI models, the same tools people use for writing, coding, and problem-solving.
These systems aren’t trained solely for chess mastery, but the tournament offered a window into:
- Reasoning ability under strict rules
- Long-term planning skills
- Error correction during play
This made the event a benchmark not just for raw calculation power, but for strategic thinking in non-specialist AI systems.
The historical significance of chess as an AI benchmark
For decades, chess has been a proving ground for AI. IBM’s Deep Blue famously defeated Garry Kasparov in 1997, a milestone for machine intelligence. However, today’s LLM-based contest shows a shift in priorities: rather than building a single-task chess machine, companies now test AIs designed for everything.
This raises questions about future AI competitions. Could we soon see tournaments for creative writing, coding challenges, or real-time negotiation between AIs?
Key takeaways from the tournament
- ChatGPT o3’s adaptability: It maintained strong play under pressure, capitalizing on Grok’s mistakes.
- Specialization still matters: Grok’s lack of chess-focused tuning became evident in high-stakes games.
- AI evaluation is evolving: Performance in chess is now part of a broader skills assessment for LLMs.
What does this mean for the AI industry
The win gives OpenAI bragging rights in an increasingly competitive AI race. Both OpenAI and xAI claim their latest models are the most advanced, but this result adds a competitive edge to OpenAI’s narrative, especially in high-visibility, skill-based challenges.
For Google’s Gemini, a third-place finish may not dominate headlines, but it reinforces its position among the top-performing AI systems. In a marketplace where public perception matters as much as technical ability, such rankings influence how users choose between AI platforms.