GPT, Gemini, and Claude All Pick the Same Right Answers
A new study of 270 puzzles with multiple correct answers finds AI models cluster tightly together and diverge sharply from how humans pick solutions.
- New paper studies human vs AI answer distributions on 270 multi-solution puzzles across five families.
- Models cluster tightly together but diverge sharply from human answer distributions.
- Human distributions are significantly closer to uniform across valid solutions than any model.
- Model distributions are less diverse than humans in every puzzle family tested.
- Raising reasoning effort or human-persona prompts does not close the gap with humans.
- Data and code available at hai-discrepancies.github.io.
AI Models Converge on the Same Valid Answers
A new study by Zihao Wang, Francesco Insulla, and Andrea Montanari finds that frontier AI models cluster around the same valid solutions to multi-answer reasoning puzzles, while people distribute their choices more broadly. Across 270 puzzles, models from different vendors resembled one another more than they resembled humans.
Accuracy conceals solution preferences
A solver can achieve perfect accuracy while repeatedly choosing only a small subset of the available correct answers. Sampling a model many times, or aggregating answers from many people, produces a distribution that reveals which solutions dominate, which receive little attention, and how broadly each group explores the valid answer space.
That distribution matters for code suggestions, synthetic data, creative generation, policy drafts, red-teaming, and other tasks with several defensible outputs. Standard accuracy scores record whether an answer is correct. They leave shared preferences and concentrated output patterns unmeasured.
Five puzzle families, 270 controlled tests
The benchmark contains 270 puzzles drawn from five families:
- Arithmetic
- Maze
- Rooks
- Minesweeper
- Sudoku
Each puzzle has between three and eight valid solutions. The low difficulty keeps failure rates down, allowing the analysis to focus on preferences among correct answers. The researchers collected responses from human participants and repeated samples from three frontier model families represented by GPT, Gemini, and Claude systems.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.