AI Video · 30 Sep 2026 · 09:32 CEST
NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks
Publisher preview · OZZZER analysis pending editorial review.
PUBLISHER ARTICLE PREVIEW
From the original article
Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals.
Their framework, Physis-Lang, treats physical language as a shared, optimizable representation. The same text drives data curation, model training and inference. On the public Physics-IQ Verified leaderboard snapshot dated September 29, 2026, Physis-Lang on Cosmos3-Super ranks first at 48.2 ± 1.4. The Cosmos3-Nano version ranks second at 43.3 ± 1.5.
A video world model can make a convincing clip and still get the physics wrong.Our researchers just released Physis-Lang, an open self-evolving framework that adds physics reasoning to video captions. The captions explain why and how a scene unfolds. We use them to fine-tune… pic.twitter.com/agIHuIB6N2
Conventional captions describe what happens, not why. ‘Butter melts as the temperature rises’ says nothing about heat transfer or gravity. Physis-Lang adds a physics_reasoning field to each base caption. It spells out entities, causes, interactions, governing principles, temporal evolution and effects.
The pipeline also writes a scene-specific physics_negative_prompt. This text describes likely implausible outcomes, such as a stone floating on water. It acts as negative conditioning at inference time.
The
Source
MarkTechPost · 30 Sep 2026 · 09:32 CEST
Open the original at MarkTechPost ↗