People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child. Four short years after the release of ChatGPT, many of us now take it for granted that we can converse naturally with our phones or computers.
LLMs like Claude, DeepSeek, and OpenAI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language.
Frank, a cognitive scientist at Stanford University, says of LLMs. And it raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: How is it that kids can still outperform the most linguistically sophisticated machines ever built? Finding answers has stakes for both AI research and cognitive science.
For the past decade, language models have mostly gotten better by getting bigger. Meta’s open-weight LLM Llama 3.1, released two years ago, chewed through 15 trillion tokens (word-like chunks of language) in pretraining—the main step of training a model that happens before it is fine-tuned for a specific task, like being a chatbot. Frontier models could be pretraining on 10 times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University.
But there’s only so much internet to train on, and eventually—perhaps as early as the 2030s—the well of easily available data could run dry. Kids show that it could be possible to learn more with less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words.
Add literacy to the mix and you can boost that word count to maybe 300 million words by age 20. The difference in scale is something that can only really be gestured at in analogy. If you were to print out on paper all the words used to train a modern LLM, you could make a stack that would reach past the International Space Station.
Extract — continue reading at the source.