What Fei-Fei Li actually shipped.To understand why Atlas matters, you have to get one thing straight.Most of the AI we've known is fundamentally "look at a picture, describe it" or "look at a picture, make a new one." It gets better at looking realistic, but it stays flat.Atlas is different. It's the first "omni world model." It takes text, images, video, and 3D information and puts them into one unified space. Give it one or a few photos, and it rebuilds a 3D scene you can move a camera through.The point isn't how sharp the picture is.The point is that it remembers the spatial continuity. You walk from the front to the back, and the table is still the same table, the wall is still the same wall. It doesn't collapse halfway through.In plain English: the old AI hands you a postcard. Atlas builds you a set you can walk into.Then there's the Chinese teamNow, Yingsu.This company flies so far under the radar that I'd never heard of it. But its model, InSpatio-World, open-sourced in March, does something heavily overlapping with Atlas.You give it a video, and it turns that video into a dynamic 4D scene you can keep exploring.Notice that word: dynamic.Atlas starts from still images and rebuilds a static scene. InSpatio-World starts from a video and models time itself. You can not only change the angle, you can change the moment. Same event, walk to a position that was never filmed, and watch it again.One builds a world. The other records the real world into a world you can re-enter. Different routes, same direction.Also: Yingsu says an upgraded version is coming soon, and it'll be open-source, as always.The wild part is the leaderboard."Similar capability" is easy to say and easy to doubt. So here's the data.There's a benchmark called WorldArena 2.0, jointly run by Tsinghua, SJTU, HKU, and others. It tests whether an "embodied world model" is actually useful.It doesn't care how pretty your frames are. It checks whether the robotic arm moves on an accurate trajectory, whether depth is stable, and whether an object pushed keeps moving according to physics.In other words, it's an exam hall for robots.As of August 27, 77 models competed, and Yingsu's InSpatio-Curious took first place with 66.11 points.It ranked first on four metrics: Physics Adherence, Trajectory Accuracy, JEPA Similarity, and Depth Accuracy. Depth Accuracy hit 99.29. Trajectory Accuracy was 64.89, a full 3.02 points ahead of second place.But here's the detail that stuck with me, and it's counterintuitive.Its Image Quality, the visual sharpness score, was only 60.64. Almost 9 points lower than the runner-up.And it still took first overall.Sit with that. Everyone else polishes their pixels to farm points. This one has the worst-looking frames and wins anyway, on trajectory, depth, and physics. The "substance" metrics.This leaderboard doesn't measure who's prettiest. It measures who actually understands the rules of the world.Fig 1: The one metric it "fails" (orange) is the one it cares about leastWhat a world model even is.I need to pause here and talk about this phrase everyone keeps throwing around: world model. It's the thread running under all of this.Twenty-something years ago, a computational neuroscientist named Jeff Hawkins wrote a line: intelligence isn't just pattern recognition; it's building a model of the world.I read that line once, and it didn't land. Watching Atlas and Yingsu this week, I suddenly got it.Everything AI has done the past few years is recognition. Recognizing what you said, recognizing the cat in the photo, recognizing what the next token is.A world model is about understanding. Understanding that a dropped glass shatters. Understanding there's space behind the wall. Understanding what I'd see if I walked around.Fei-Fei Li put it something like this at Stanford: we built a machine that can write poems and paint, but it has no idea where it is. It doesn't know it lives in a three-dimensional world.The world model is the thing trying to teach it that.Data is the next battleground.There's one more thing that matters more than the model itself.On August 29, at the second China Spatial Intelligence Conference in Wuhan, Yingsu teamed up with more than 20 universities and institutions to launch SIDO, a large-scale spatial intelligence 3D data open program.That reads like a press release. Let me translate it.Spatial intelligence isn't short on algorithms right now. It's short on data. The kind of data with millions of 3D scenes that have physical properties and can be interacted with. Without that, a world model just plays around in its own little sandbox and falls apart the moment it touches reality.SIDO wants to put the people who produce data, the people who train models, and the people who use the data at the same table. Over the next two years, they aim to build a million-scale 3D/4D scene database.That move reminds me of ImageNet.In 2009, Fei-Fei Li built ImageNet and almost single-handedly lit the fuse on deep learning. Now, it's a Chinese team working on world models' "ImageNet moment."Notice the coincidence. Fei-Fei Li shipped Atlas. A Chinese team is building the data foundation. Two lines, heading for the same point.This isn't a "who copied whom" story.A side note: Yingsu isn't alone on this track in China.Just the other day I saw that Yingmu Hyper3D, a top 3D-generation player, released its own world-generation model, WorldGen, which turns a single image into an interactive, editable 3D scene. Its underlying CAST technique won the best paper award at SIGGRAPH 2025, the top graphics conference. The only other commercial companies to win that year were Google and Meta.From Yingsu to Yingmu, Chinese teams are further along on "the world" than a lot of people assume.So let me be clear: this isn't a story about copying or being earlier.Atlas and Fei-Fei Li represent the American route: top talent, top capital, moving from academia to product. Yingsu represents the Chinese route: open source, benchmarks, building a data foundation, moving forward as an ecosystem.I don't know who winsI don't really want to conclude this.Honestly, I don't know which route is faster. Maybe there's no winner to name at all. One builds the world, the other records it, and maybe the complete answer is the two of them stitched together.What I'm sure of is something else.AI's next battlefield just moved from chat to the world.An AI that actually understands space is the one that can drive a car, move boxes in a warehouse, help an old person up in a care home. That's not hype. That's the foundation on which physical AI either stands or falls.So a lot of people are filing today under "just another product launch." I think one sentence is the actual point:Intelligence isn't just pattern recognition. It's building a model of the world.That line was written over twenty years ago. Today, it finally stopped being a prediction.
Fei-Fei Li Just Released a Landmark Model - A Chinese Open-Source Team Did It Six Months Ago
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.