Puzzles and games have long served as testing grounds for artificial intelligence development, from checkers in the 1950s to chess and Go in recent decades. Measuring progress on these tasks provides a window into both the strengths and limitations of AI systems.
Performance on word and logic puzzles has accelerated sharply in recent months. In late 2024, even the most advanced models could solve only 18 percent of New York Times Connections puzzles, but by early 2025, some models were solving them nearly perfectly every time.
The rapid gains reveal where AI systems excel and where human reasoning still outperforms them. Researchers have used a collection of seven puzzles to test where models succeed and fail against human solvers.
