Reka AI releases Rho-1, a 19B omni-model for text, images, video and robot control

Reka AI releases Rho-1, a 19B omni-model for text, images, video and robot control

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model. It processes and generates text, images, video and robot control actions inside a single neural network.

The design is the point. According to The Decoder, most AI systems route tasks to specialized models. Rho-1 does not. It runs every modality as tokens in one shared context window, with no tool calls and no external models. The model generates continuous video in real time and responds to new instructions on the fly without restarting.

The robotics side rests on a shared set of weights: the same weights that predict camera images also drive robot movements. Robot training data is scarce, so Reka AI built an inverse dynamics model that pulls control signals from ordinary internet videos. Rho-1 was trained on 320 H100 GPUs over about three months.

The company has a multimodal track record. In April 2024 it shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3 and Gemini Ultra on benchmarks. The Decoder frames the Rho-1 release as part of a broader push in AI research toward so-called world models.

Key facts

  • Rho-1 is a 19-billion-parameter omni-model from Reka AI, released as a research preview.
  • It handles text, images, video and robot control actions as tokens in one shared context window, with no tool calls or external models.
  • It generates continuous video in real time and accepts new instructions on the fly without restarting.
  • To get around scarce robot data, Reka AI built an inverse dynamics model that extracts control signals from ordinary internet videos.
  • Training took 320 H100 GPUs over about three months.

Why it matters

Rho-1 is an attempt to fold language, vision, video generation and robot control into one model instead of stitching specialized models together. The Decoder notes that most AI systems route tasks to specialized models, while Rho-1 keeps everything as tokens in one context. That the same weights predict camera images and drive robot movements is the core of the idea. The release also fits a wider push in AI research toward so-called world models. Reka AI has form here: its April 2024 Reka Core was a multimodal language model that competed with GPT-4, Claude 3 and Gemini Ultra on benchmarks.

Who it affects

The most direct audience is researchers working on multimodal models, world models and robotics, since Rho-1 puts all of those in one network. Teams that currently chain separate models for perception, generation and control are the ones for whom a single-model approach is relevant. The source does not name specific users or customers.

How to use it

Rho-1 is a research preview, so this is something to follow rather than build on today. The source gives no information on availability, licensing, weights release, pricing or access, so check Reka AI's own announcements before planning anything around it.

How solid is it

The account comes from a short news write-up by The Decoder describing the release, and its specifics are limited to the design and the training setup: 19 billion parameters, 320 H100 GPUs, about three months. No benchmark results or comparisons for Rho-1 are given, so the claims about real-time video, on-the-fly instruction following and shared weights for images and robot movement are descriptions, not measured results.

Risks and caveats

It is a research preview, and the source reports no benchmarks, so how well Rho-1 performs is unknown from this account. The source does not say how many robots, which robot hardware or which tasks were tested. It also gives no training data size, context window length or architecture details. The inverse dynamics approach of pulling control signals from internet videos is described only in one sentence, so its quality and limits cannot be judged from the source.