go-bench.com
Who plays the best Go?
Decision models on a 9×9 board. Every legal point, no hints.
Players
- OpenAI Decisions API joins when its preview opens
- TypeSafe AI Jev pilot games under way
Why Go
Go is a closed chapter for AI labs. Engines beat the best human players in 2016, and since then no lab has had a reason to train its models to play it. That is what makes it a fair test now: how well a model plays a game its makers never aimed at.
A 9×9 board keeps games short, cheap to run and quick to score, while still asking for real reading: fights, life and death, and counting at the end.
How a game works
- The model gets the rules, the board and the legal points. Nothing says which point is good.
- Each turn it answers two questions: place a stone or pass, and which of the legal points. An illegal move can't happen, so the result measures play, not rule-following.
- Each model plays both colours against a fixed ladder of reference opponents, from GNU Go to KataGo. Ratings are set against that ladder, so a new model's score is comparable to every earlier one.
- KataGo will review every move. Alongside the rating we will report points lost per move, blunders, cost per game and time per move.
The exact rules text, the two questions and how games are scored are in the protocol.
Baselines
Chat models will play the same games for comparison, once with reasoning off and once with it on, and are reported separately.
Versions are shown as each API reports them, with the first results. Games, move reviews and the code will be public.