Runway Debuts GWM Worlds 2 for Real-Time Simulation
Runway has released GWM Worlds 2, a research preview featuring a WorldPrompt control layer that allows developers to generate and steer interactive, real-time simulated environments.

Runway, a generative video startup recently valued at $5.3 billion following a $315 million funding round in February, has launched GWM Worlds 2. This research preview succeeds the original GWM Worlds released last December. The new model streams continuous 720p video at 24 frames per second alongside synchronized audio at 48,000 Hz. It introduces WorldPrompt, an input format that acts as a control layer for characters, cameras, and environments. Users can lock in specific elements, such as the initial frame, and trigger timestamped events or actions in real time.
To achieve real-time performance, Runway transforms a bidirectional diffusion model into an autoregressive one that generates content frame by frame. The company then applies distillation techniques, such as reducing denoising steps from approximately 50 down to four. This approach positions Runway against rival systems like Google DeepMind's Genie 3, which also runs at 720p and 24 frames per second but is limited to a few minutes of continuous interaction, as well as Odyssey-2 Pro and World Labs' RTFM.
For practitioners, GWM Worlds 2 shifts video generation from static clips to dynamic, promptable environments. However, developers must navigate several technical hurdles. Autoregressive generation suffers from error accumulation as minor mistakes compound over time. Managing GPU memory during infinite generation and maintaining long-term memory remain open research challenges. Additionally, the model must master counterfactual generation to ensure that alternative user actions produce equally realistic outcomes.
Beyond gaming, this technology offers significant utility for AI agent development and robotics. Practitioners can deploy thousands of simulated environments to test autonomous agents at scale without needing a structured state, as the agents interact purely through visual and auditory inputs. The platform also serves as a source of synthetic training data and can be paired with reasoning models to plan scenes before rendering pixels.
This is our own summary of reporting by Latent Space



