World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say Atlas unifies the two things computer v...
a16z(@a16z) · 商业与创投
World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say Atlas unifies the two things computer vision has always kept apart: Justin: "Historically, reconstruction has been its own subfield in computer vision with its own specialized tasks and models. Generation is what all the text-to-video models are really good at." "Those are great for creative applications if I want to imagine something that's never been there before." "But now with Atlas, for the first time, we're putting these two different parts of visual intelligence together in one model. It can do both 3D reconstruction and generation together in one architecture." "We had to make it multimodal from the start. This thing natively works on text, images, videos, and camera poses as a native input to the model... It uses 3D as a native modality that it works on." Fei-Fei: "Computer vision has been around for more than half a century... Our field traditionally has multiple tracks. You go to a computer vision conference: you have the pixel generation track, you have some recognition track, and you have a 3D reconstruction track." "This is an elegant model that unifies the problem of reconstruction and generation by anchoring on viewpoints and viewpoint estimation. That's just incredibly powerful." @jcjohnss @drfeifei Your browser does not support the video tag. 🔗 View on Twitter a16z @a16z World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just thr