Y’all gonna get mad if I said I told you so …
Gary Marcus(@GaryMarcus) · 人工智能
Y’all gonna get mad if I said I told you so … Superman @thesupermanmx Researchers argues OpenAI and Anthropic will never get us to AGI. 21 top researchers from Stanford, Oxford, DeepMind, CMU & Meta released a paper saying LLMs are a dead end. It’s called “Visual General Intelligence”. And it completely flips how we look at artificial intelligence. For years, the playbook has been simple: feed mountains of web text into a Transformer, scale up the parameters, and watch reasoning emerge. GPT proved language can take you far. But text is fundamentally limited. It’s a compressed, human-abstracted symbol system. It lacks physics. It lacks geometry. It lacks the raw, unadulterated reality of the physical world. The paper argues that true general intelligence cannot be built on words alone. It requires a vision-centered foundation. Instead of starting with language and translating pixels into text, the next generation of models must start with raw visual experience, images, spatial geometry, and continuous video. Think about how humans learn. A baby understands gravity, permanence, and spatial reasoning long before it ever learns to string a sentence together. Vision isn't just an input modality. It is the core operating system of physical reality. When models learn natively from visual streams and video dynamics, they don't just memorize text patterns. They learn physics. They learn cause and effect. They build a true internal model of the world. 🔗 View Quoted Tweet 💬 17 🔄 14 ❤️ 88 👀 6008 📊 24 ⚡ Powered by xgo.ing