Why AI Needs Trillions of Words to Learn What a Child Learns From Millions

LLMsResearch
The Similarities between the Learning of Children and AI

cogmed.com · News coverage photograph, editorial use approved

The Core · TL;DR

  • Meta's Llama 3.1 was pretrained on 15 trillion tokens, while children may hear only around 100-300 million words by adulthood.
  • Georgetown's Ethan Gotlieb Wilcox estimates frontier models may use 10x more data than Llama 3.1, roughly 150 trillion tokens.
  • Despite the massive data gap, models like Claude, DeepSeek, and GPT are fluent enough to pass as human in conversation.
  • Stanford's Michael C. Frank studies how children achieve language competence with vastly less input, a key open question for AI efficiency research.

A child begins to grasp language after hearing roughly 100 million words by adolescence, and perhaps 300 million by adulthood once reading is factored in. Meta's Llama 3.1, by comparison, was pretrained on 15 trillion tokens. That gap, several orders of magnitude wide, sits at the center of what researchers call the data efficiency problem in language learning.

The comparison isn't a rounding error. Children typically start showing signs of understanding language around their first birthday, well before they've been exposed to anything close to the volume of text modern language models consume. Machines, by contrast, need staggering quantities of data to reach fluency, and even then they learn differently than humans do.

Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University who studies data efficiency in language models, estimates that today's frontier systems may be trained on ten times more data than Llama 3.1's 15 trillion tokens. If accurate, that would push some models into the range of 150 trillion tokens, a figure with no equivalent in human experience.

Fluent, but not human-like

Despite this reliance on vastly more input, models like Claude, DeepSeek, and OpenAI's GPT family have become fluent and flexible enough to pass convincingly as human conversational partners. That fluency raises an odd tension: the outputs look increasingly human, even as the learning process behind them looks nothing like a child's.

Michael C. Frank, a cognitive scientist at Stanford University, studies this efficiency gap from the human side, examining how children extract so much linguistic competence from comparatively little input. His work, alongside Wilcox's, treats the disparity not as a curiosity but as a genuine open question for both fields.

For AI architects, the gap is a practical challenge: building models that learn more like children could mean systems that require far less data, computing power, and energy to reach similar capabilities. For cognitive scientists, it's a window into what makes human language acquisition so efficient in the first place, whether it's innate structure, rich multimodal context, social interaction, or some combination not yet captured in any training pipeline.

Neither field has closed that gap yet. But the scale of the mismatch, trillions of tokens against a few hundred million words, suggests current LLM architectures are solving a fundamentally different problem than the one a toddler solves every day.

Original reporting and research used to synthesize this article.

  1. 1Kids outlearn AI—and we still don’t know whytechnologyreview.com
WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram