Skip to content
Table of contents5 sections · tap to jump
  1. The reconstruction table, row by row
  2. The model that takes the seventh row has a B in its name
  3. The camera comparison is not like for like, and the post says so
  4. The robotics claim is the one to watch
  5. What you can actually use today
World Labs' Atlas wins six of seven 3D benchmarks, and you cannot run it

Newsai5 min read

World Labs' Atlas wins six of seven 3D benchmarks, and you cannot run it

Ahmad JSep 2, 2026

Signalsolid1independent source

World Labs published Atlas on 1 September and calls it an omni world model: one network pretrained from scratch to operate on text, images, camera poses and 3D depth maps together, rather than a video generator with camera control bolted on afterwards.

The claim that matters is not the demo reel. It is that a general model beat five specialist reconstruction models at their own job. That claim holds on the average, and the table underneath it is more interesting than the sentence on top of it.

The reconstruction table, row by row#

World Labs evaluates sparse-view 3D reconstruction against five recent baselines: Pi3X (posed), pi-cubed, VGGT-Omega 1B, Depth Anything 3 and MapAnything. The metric is mean absolute-relative pointmap error scaled by ten to the minus three, and lower is better.

Averaged across seven datasets, Atlas scores 25.3 against 28.7 for the closest baseline. That is a real lead, and it is the number the announcement rests on.

The individual rows say something more specific:

  • Atlas wins outright on DTU (8.6), ETH3D (9.3), NRGBD (6.5), 7-Scenes (37.8) and ScanNet (12.4).
  • On KITTI it scores 60.0 against 60.2. That is a win by two tenths, on a row where the other models run out to 115.3. It is a tie in everything except the ordering.
  • On Tanks and Temples it does not win. Atlas scores 42.4, Pi3X (posed) also scores 42.4, and VGGT-Omega 1B scores 40.2.

Six of seven, then, with one of the six decided by 0.2 and the seventh going to a baseline. World Labs published every one of those rows. That is the part worth crediting: the losing row is in the chart, not buried in an appendix, and the company still described the result as outperforming the specialists, which on the average it does.

The model that takes the seventh row has a B in its name#

VGGT-Omega 1B is a one-billion-parameter model built for a single task. It beats a model pretrained from scratch on a large multimodal corpus, on one of the seven datasets both were measured on, and comes second to Atlas on three more (DTU, ETH3D and NRGBD).

This is not an upset. It is the ordinary shape of the trade: a general model buys breadth and pays for it at the peak of any one narrow task, and a specialist buys the reverse. It is the same pattern behind small models eating the easy work, and it is the reason an average across seven benchmarks is a weaker claim than it looks. If the job in front of you is the Tanks and Temples job, the average is not the number you care about.

The camera comparison is not like for like, and the post says so#

The second evaluation is camera-controlled generation, judged by third-party human raters choosing between Atlas and five video models: MiniMax H3, Gemini Omni Flash, Happy Horse 1.1, FLUX 3 and Seedance 2.5.

Atlas takes camera geometry as a native input type. The others do not, so World Labs described the camera path to them in a text prompt using standard cinematic terms, and says so directly: "Other models do not accept cameras as a native input format." The post adds that more sophisticated prompt engineering or creative multimodal prompts could improve camera following for some of them.

So the result is honestly reported and it is also structural. It measures a model with a camera input against models that must infer a camera path from prose. As a description of what you would hit using these tools today, that is fair. As evidence about the underlying generative quality of the other five, it is not. Both readings are true, the post supports both, and that is more than most leaderboard claims offer.

The robotics claim is the one to watch#

The part of this with the most weight behind it is Real-to-Sim. World Labs reports capturing two large environments with phone video, 24 frames each, reconstructing them, then generating the colour and depth data a simulated robot's body-mounted cameras would observe along a path. For manipulation it reports building a simulation from a few casual recordings that also captures how objects move and interact, then varying the objects, their positions, the robot's motion, the lighting and the background to produce training data at scale.

If that holds up outside the demos it changes the cost of the capture step, which today wants a rig and a technician. It does not by itself close the gap that actually bites, because the sim-to-real gap is not a bug you can patch. A simulator can render a room beautifully and still get contact dynamics, friction and sensor noise wrong, and those are what break a policy on real hardware. Simulation that looks better and simulation that transfers better are different achievements. Only the first one has been shown here.

What you can actually use today#

Nothing, unless you are invited. Atlas is entering early access with selected partners, the call to action is a request form, and there are no weights, no public API and no published terms.

That is the practical shape of the release, and the useful response is not to wait for an invitation. The five baselines Atlas measured itself against are described in the same post as specialist open-source reconstruction models, and they are the ones with published results on the same seven datasets. If you have a sparse-view reconstruction problem in front of you now, that list is the shortlist, and VGGT-Omega 1B is the specific one to try first, because it is small and it already beat Atlas on one of the seven.

Come back to Atlas when a number in that table has been reproduced by somebody outside World Labs. The architecture is the genuinely new thing here, the scaling evidence is the company's own, and neither becomes a tool you can hold until access opens.

Sources

  1. Atlas: A World Model for Spatial Intelligence, World Labsworldlabs.ai

Ask about this article

Answered only from this piece — the AI never invents.

React
ShareXLinkedInBluesky

More in aiMore in ai

Discussion