Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control
Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.
Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions within a single neural network, according to The Decoder's report on the release.
What Rho-1 does
Most multimodal systems route tasks to specialized models or call external tools. Rho-1 does neither, according to Reka. All modalities run as tokens in one shared context window, with no tool calls and no external models.
Reported capabilities:
- Continuous video generation in real time. The model generates video as a continuous stream.
- Live instruction updates. It responds to new instructions on the fly without restarting.
- Unified perception and control. The same weights that predict camera images also drive robot movements.
Training approach
Robot training data is scarce. To work around this, Reka built an inverse dynamics model that extracts control signals from ordinary internet videos. That lets the team use large volumes of unlabeled video as a source of action data.
Rho-1 was trained on 320 H100 GPUs over about three months. Assuming continuous use, that implies roughly 700,000 H100 GPU-hours. This is our estimate, not a figure Reka has published.
What has not been disclosed
The source material does not include:
- Context window size: not yet disclosed
- Pricing per 1M tokens: not yet disclosed
- Benchmark scores: none published in the source report
- Training data cutoff: not yet disclosed
- Weights, license, and API availability: not specified
The capability descriptions above, including real-time video generation and unified robot control, come from Reka. We have not seen independent evaluation.
Background
Reka AI is not new to multimodal work. In April 2024 it shipped Reka Core, a multimodal language model that, per the company at the time, competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. Rho-1 extends the company's focus from understanding multimodal inputs to generating video and actions.
The release fits a broader push in AI research toward so-called world models, which learn to predict and act within visual environments rather than only produce text.
What this means
The notable claim is architectural: one set of weights handling perception, video prediction, and motor control as tokens in a single context. If it holds up, it removes the orchestration layer that most robotics stacks need between a vision-language model and a low-level controller.
Two caveats apply. First, this is a research preview with no published benchmarks, so there is no way yet to compare Rho-1 against specialized robot policies or dedicated video generators. Second, a 19B-parameter model generating video in real time raises questions about latency, resolution, and hardware requirements that Reka has not answered publicly.
The inverse dynamics approach to mining internet video for control signals may prove the more durable contribution. It addresses the main bottleneck in robot learning, the shortage of action-labeled data, without large-scale teleoperation collection. Its quality depends on how accurately the inferred actions match real robot behavior, which independent testing will need to confirm.
Until Reka releases technical details, access terms, and evaluations, Rho-1 is best treated as a research signal rather than a deployable option.
Related Articles
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter
Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.
Reflection AI unveils Beam: 501B-parameter open-weight MoE with 1M-token context
Reflection AI has unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active, and a 1 million token context window. The company claims it matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. Weights and the full technical report are due later this month.
China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs
China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.
Comments
Loading...