Reka Debuts Rho-1 to Merge Reasoning, Video, and Robotics
Reka has unveiled Rho-1, a 19-billion-parameter model that processes text, video, and robot actions in one checkpoint, eliminating the need to chain separate specialized AI services.

Reka's new research preview, Rho-1, is a 19-billion-parameter omni model trained from scratch on 320 Nvidia H100 GPUs over approximately three months. Unlike traditional systems that chain separate models for language, video generation, and motor control, Rho-1 uses a symmetric two-stream transformer with a shared attention state. One stream handles discrete tokens for text and reasoning, while the other uses a flow-matching head to denoise continuous representations like image latents and robot actions. This unified architecture allows generated media and control signals to inform subsequent steps in the same session.
The base model, which uses 99 denoising passes, achieves a median generation rate of 0.79x real time, with streaming output starting after about six seconds. A distilled version called Rho-1 Flash can render a 5.3-second video in roughly one second using an eight-step denoising process, and re-render one-second clips in 1.1 seconds. In robotics, Rho-1 was tested on a LIBERO simulation task, where it emitted seven action channels alongside a predicted wrist-camera view. To address data scarcity, Reka paired the model with an inverse-dynamics model to infer control signals from unlabeled video footage.
For developers, this architecture eliminates the need to serialize state and transfer data across incompatible representations. However, Reka notes several limitations that prevent immediate production use. Video rollouts are capped at a resolution of 672x384 pixels, and the model suffers from long-horizon structural drift during 30-second streams. Additionally, temporal grounding is weak, meaning bounding boxes do not reliably track objects in motion, and video editing remains brittle. Because Rho-1 is currently only available as a research preview on Reka Cloud without public weights or an API, independent verification of these capabilities is not yet possible.
This is our own summary of reporting by AlphaSignal


