lunes, 31 de agosto de 2026

It’s not child’s play: Teaching a computer to use language requires an inhuman amount of data—and it still learns much more slowly than kids MIT Technology Review | August 31, 2026

https://geneticliteracyproject.org/2026/08/31/its-not-childs-play-teaching-a-computer-to-use-language-requires-an-inhuman-amount-of-data-and-it-still-learns-much-more-slowly-than-kids/ Chatting naturally with a computer now feels routine, but as MIT Technology Review’s Elise Cutts reports, the large language models behind that fluency are shockingly data-hungry compared to human toddlers. Meta’s Llama 3.1 trained on 15 trillion tokens; a preteen raised in a language-rich home hears something like 100 million words. “We still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year,” says Stanford cognitive scientist Michael Frank.

No hay comentarios: