MTS

MTS

@mtslive · Twitter ·

World Labs co-founder @jcjohnss predicts a real-to-sim-to-real future where Atlas could train a robot for any new environment in minutes using just 5 photos from your phone: "One is this notion of fully in-context learning. Maybe I've got a robotics foundation model, then I can demonstrate the robot once how a task should be performed, and that's enough for the robot to figure out." "Another version is real to sim to real. Maybe I've got my pre-trained robotics foundation model, but I want to adapt it to this particular environment, like this studio, this table, moving this microphone around." "There's a version where you could come in here and take just a couple casual videos with your phone or a couple images with your phone, use Atlas to reconstruct the space, and now stage all kinds of robotics interactions in this studio space." "If you could lower the barrier to doing that, you could do it in five minutes to take five photos, upload it to Atlas, have it generate a simulation, maybe describe in natural language what kind of task you want the robot to do, have an agent go build that simulation for you, RL fine-tune your general purpose robotics foundation model, and now you've got a robot maybe in the span of a couple minutes that could come in and is perfectly adapted to this space." @theworldlabs

World Labs

World Labs

Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.