BEIJING, Aug. 14, 2026 /PRNewswire/ -- Over the past two years, the core proposition of the AI industry has been "generation" — generating text, images, video, and music. But on July 19, 2026, at the World Artificial Intelligence Conference (WAIC), Fang Han, Chairman and CEO of Kunlun Tech (Parent company of Skywork AI), formally declared 2026 as "the first year of world models," marking a pivotal transition in which AI evolves from content generation toward understanding and interacting with the physical world. This is not a marketing slogan. It signals that the supply-side revolution has entered its second phase: AI is no longer merely producing content — it is beginning to comprehend the physical world.
I. The Camera Analogy, Continued
This transition also echoes the camera analogy proposed by Jiao Juan, Chief Analyst for Internet and Media and Head of AI Ethics Research at Founder Securities. In Jiao's view, when photography was invented, it did not merely add a new medium to the arts — it first destroyed the economic foundation of realist painting, then drastically lowered the cost of visual representation, and finally, on the ruins of the old supply, gave birth to an entirely new art form: cinema. The history of technology suggests that genuine revolutions do not begin by creating abundance; they begin by destroying the existing supply structure. AI's supply-side revolution is following precisely this trajectory.
II. From "Storing Frames" to "Storing Space": The Deepening of the Supply-Side Revolution
At the WAIC forum, the technical breakthroughs announced by Skywork AI revealed the mechanics of this deepening. Consider the most stubborn problem in world models: long-term memory. Previous models "remembered" by storing entire frames — but when the camera moved and all object positions shifted, the model could no longer recognize them. Evaluation data showed that no mainstream world model achieved an object-reappearance score above 0.6 (out of 1.0). Skywork AI's newly released interactive world model, Matrix-Game 3.5 adopted a fundamentally different approach: instead of storing whole "photographs," it breaks each frame into countless small patches, each tagged with its precise three-dimensional coordinates — its distance from the camera, its position in space. This mechanism, called "Patch Memory," represents an upgrade from "storing frames" to "storing space." The model no longer remembers "what a certain frame looked like" but rather "what exists at a certain three-dimensional location." As the technical report states, "the problem of 'disappearing objects' that has plagued the industry for so long has been solved in a true sense for the first time."
This is not an incremental improvement — it is a paradigm shift in how AI relates to physical reality. Combined with the PRoPE geometric position encoding and Warped RoPE mechanism, the model's positional encoding evolved from merely recording "which second on the timeline" and "which row and column in the frame" into a geometrically meaningful world coordinate system. The result: a 5B-parameter model achieving real-time streaming video generation at 720P resolution and approximately 20 FPS on a single GPU, with one-minute memory capability. As the report puts it, "20 FPS is not 'playing a video' — it is a 'living world': you operate, it responds, no stuttering, no memory loss, stable visuals, fast reactions." The supply-side revolution has produced something qualitatively new: not a video generator, but an interactive, living world.
III. Data as the Foundation: Letting AI "Understand" Rather Than Merely "See"
Beneath the technical breakthroughs lies a more foundational insight. Training a world model requires not just video, but video endowed with physical information. As the technical team recognized, "letting AI understand the physical world requires giving it data that carries physical information." They built a multi-agent system that operates 24 hours a day, automatically surveying thousands of games and scoring them on multiple dimensions. More critically, they constructed an automated annotation pipeline that performs 3D reconstruction on every video, estimating camera pose, metric depth, and camera intrinsics for each frame. The pipeline's output: over 5 million high-quality video segments, more than 10,000 hours of effective training data, covering over 1,200 game scenes — each carrying precise camera pose, depth maps, and text descriptions. Jiao describes this as industrial-grade infrastructure for the supply-side revolution: an "oil refinery" that turns raw pixels into training fuel with physical scale.
IV. "Not Stacking Parameters, but Clever Architecture": The Logic of the Entry Business
Perhaps the most counterintuitive decision was the team's choice to "almost not add any new parameters at all." Beyond the character Reference Adapter, all of Matrix-Game 3.5's core interactive capabilities — PRoPE geometric encoding, Patch Memory long-term memory, camera control, and dynamic-static decoupling — were implemented on the existing model architecture without additional parameter burden. This reflects a deeper design philosophy: "modular, pluggable, and transferable." The team built the interactive framework as a lightweight "plugin system" — each module works independently and can be quickly adapted to different video foundation models. The practical benefit: "the native open-ended content generation capability of the base video model is maximally preserved." This "not stacking parameters, but relying on clever architecture" approach reveals the essential logic of the first layer of the AIGC framework. Foundational model companies are not competing on who can pile up the most parameters — they are competing on who can establish the entry position that the entire ecosystem must pass through. The real game is architectural ingenuity and ecosystem gravity, not brute-force scale.
The open-source strategy confirms this logic. The team behind DiT, NYU Assistant Professor Xie Saining, built Solaris — the world's first multi-person video world model — on the open-source foundation of Matrix-Game 2.0. NVIDIA and Zhejiang University's joint work, Light Interaction, was developed on Matrix-Game 3.0. Furthermore, NVIDIA's SANA-WM and Adobe's RELIC, among other international leading companies' world model projects, all use Matrix-Game-related models as comparison benchmarks. When the world's top researchers and companies choose your foundation as their base or benchmark, the gateway position is established. This is the "positioning" phase of the gateway business — the first stage of the iOS-like three-phase path: positioning, clearing, revenue sharing.
V. Three Layers of World Models and the Concept of "Learnware"
Academic depth complemented industrial breakthroughs. Academician Zhou Zhihua of the Chinese Academy of Sciences, Professor and Vice President of Nanjing University, divided world models into a perception layer, an understanding layer, and a decision layer in his keynote speech, emphasizing that the decision layer has long faced two theoretical bottlenecks: "squared amplification of multi-step inference compound error" and "excessively high sample complexity." He shared his team's latest breakthrough: distribution-matching technology can compress inference error from squared to linear, while an innovative reverse application of the Bellman equation reduces the required sample size to one over the square root of T. More strikingly, he proposed the concept of "learnware" — composed of a model together with AI-automatically-generated "specifications," allowing multiple heterogeneous models to be reused and assembled without disclosing their underlying data, opening a new path for world models to evolve from single-pipeline customization toward a growable, evolvable model ecosystem. This concept of learnware points toward a future where the supply-side revolution's infrastructure becomes composable and reusable — models as building blocks, not monoliths.
VI. "Pores Are Acting": The Inflection Point Signal
The most vivid signal that the supply-side revolution has crossed an inflection point came not from a technologist, but from an actor. Huang Xiaoming, Vice Chairman of the China Film Association, described seeing an AI-generated film: three months earlier, he could still discern AI traces, but upon the latest viewing, "even when told it was AI, I could no longer tell." The detail had reached the point where "the spots and pores on the face, even the micro-expressions — the pores are acting," which genuinely startled him. When a seasoned performer cannot distinguish AI-generated performance from the real thing, the supply-side revolution has crossed the threshold from "good enough for experiments" to "good enough to compete with traditional production." This is the moment the camera, having destroyed realist painting, begins to give rise to cinema.
VII. Conclusion: The Supply-Side Revolution Deepens
For Skywork AI, the declaration of "the first year of world models" is not an endpoint but a starting point. As Fang Han's three predictions suggest, world models will step out of the screen and into the physical world; the gaming industry will be the first to be transformed; and AI music and AI video tools will evolve from technical experiments into mass creation tools. The supply-side revolution is deepening: from generating content to understanding the physical world, from offline rendering to real-time interaction, from memoryless short clips to minute-level consistency.
Jiao's camera analogy still holds: AI first disrupted the traditional supply of content production, then drastically lowered costs, and is now — through world models that understand physical reality — gestating entirely new forms. "In the next three to five years, world models will become the infrastructure of the gaming industry, and the way content is produced will be completely refreshed." The supply-side revolution continues. What we are witnessing is not the end of one industry, but the beginning of the next.
Sumber: www.prnasia.com
Masuk tirto.id




























