AI Can't Outsmart Kids, but the Principles of Learning Remain a Mystery
AI requires 100,000 times more data to learn than a child. According to an MIT Technology Review report, children become fluent in their native languages with a tiny amount of data, whereas AI must train on text equivalent to 100,000 times the amount of data a child is exposed to in order to reach a similar level.
While recent generative AI (such as ChatGPT and Claude) has acquired human-like language capabilities, Stanford cognitive scientist Michael C. Frank pointed out an efficiency issue, stating, "Recent progress is amazing, but to reproduce the achievements that happen in a living room in a single year, we have to burn down the forest and scrape together all of human knowledge."
Meta’s open-weight model Llama 3.1, for instance, processed 15 trillion tokens during its pre-training phase. However, the number of words a person encounters by age 20, including those necessary for literacy, is only around 300 million. Georgetown cognitive scientist and linguist Yang (Seung-eun) Gottlieb Wilcox noted that cutting-edge models are pre-trained on more than ten times that amount of data.
As projections indicate that training data could be depleted online as early as the 2030s, improving the efficiency of AI learning has emerged as an urgent task.
