Ask an LLM to Free-Associate and It Thinks Smaller Than You Do

The Core · TL;DR
- A new study compares semantic word-association patterns between 82 human participants and three LLMs (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5)
- Humans consistently showed higher entropy, larger semantic leaps, and broader conceptual dispersion than all three models tested
- Temperature tuning improved alignment on individual metrics but no configuration matched the full human profile across all dimensions
- The paper, submitted to arXiv on July 13, 2026, is set for presentation at the Cognitive Science Society's annual meeting
Give a person and a chatbot the same word association task, and the human mind wanders further. That's the central finding of a new comparative study measuring how GPT-4o, Gemini-2.5-Pro and Claude-Sonnet-4.5 navigate meaning compared to 82 human participants performing verbal fluency tasks, the classic psychological exercise of naming related words in sequence.
The paper, submitted to arXiv on July 13, 2026 and accepted for the Proceedings of the Annual Meeting of the Cognitive Science Society, treats word generation as a kind of walk through semantic space. Instead of just scoring whether the outputs looked human-like, the researchers broke the "walk" into three measurable dimensions: entropy, which captures how predictable each step is; distance to next, which measures how big a conceptual leap each new word represents relative to the last one; and distance to centroid, which tracks how far the whole sequence drifts from its own average, a proxy for overall dispersion.
Across all three metrics, humans came out ahead. Participants showed higher entropy, took larger semantic steps between consecutive words, and produced sequences with wider overall spread than any of the three models tested. In practical terms, human associative thinking looked less predictable and more exploratory, while the LLMs tended to stay closer to well-trodden conceptual paths.
Turning Up the Temperature Doesn't Fix It
The researchers didn't stop at a single model configuration. They also tested whether adjusting temperature, the parameter that controls how much randomness a model injects into its outputs, could close the gap. It helped, but only partially. Raising temperature nudged individual metrics closer to human levels in some cases, yet no single setting reproduced the full human profile across entropy, step distance, and dispersion simultaneously. Push the randomness up enough to widen dispersion, and predictability metrics might still miss the mark, or vice versa.
That partial-alignment result is arguably the more consequential finding for anyone building on top of these models. It suggests the divergence between human and machine associative thinking isn't simply a matter of dialing in the right sampling parameters. The three models, despite differing architectures and training approaches from OpenAI, Google, and Anthropic respectively, all converge on a similarly narrower style of semantic exploration relative to humans, and no amount of temperature tuning fully erases that signature.
The implications reach beyond cognitive science curiosity. Verbal fluency and associative range are often used as proxies for creativity, idea generation, and even cognitive flexibility in clinical settings. If LLMs systematically underexplore semantic space compared to humans, that has direct relevance for anything from brainstorming assistants to educational tools that lean on these models to simulate open-ended thinking. It also raises a methodological flag for researchers using LLMs as stand-ins for human subjects in psycholinguistic experiments, a practice that has grown increasingly common as models have gotten more fluent. This study suggests such substitutions may quietly bake in a narrower cognitive style than the humans they're meant to approximate.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
