Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation

NVIDIA and Hugging Face publish a full recipe for adapting the Cosmos Predict 2.5 world model to robot manipulation using LoRA and DoRA adapters on a single H100.

Monday May 18, 2026 Source: huggingface.co
TL;DR — Quick Answer

A May 18, 2026 guide from NVIDIA and Hugging Face shows how to fine-tune the 2B-parameter Cosmos Predict 2.5 world model with LoRA or DoRA adapters for robot video generation. Using 92 pick-and-place demonstration videos, 100 epochs of rank-32 LoRA training takes 17 hours on a single H100 or 2.5 hours on 8 GPUs — and fixes the base model's hallucinated hands, wrong-hand errors, and jitter. Both adapters raise physical plausibility and instruction-following scores, with rank 32 best for instruction following.

Key Takeaways

Fine-Tuning NVIDIA Cosmos Predict 2.5 with LoRA/DoRA for Robot Video Generation — AI news article illustration

NVIDIA and Hugging Face have published a step-by-step guide for fine-tuning Cosmos Predict 2.5, the 2B-parameter world model that generates physically plausible videos from text, images, or clips — using LoRA and DoRA adapters to specialize it for robot manipulation without the cost and risks of full fine-tuning.

Why Fine-Tune a World Model

Robot policies need demonstration data, but real-robot trajectories are slow and expensive to collect. A fine-tuned world model can synthesize them at scale. Full fine-tuning of a 2B model, however, is expensive and risks catastrophic forgetting of general knowledge. LoRA and DoRA inject small trainable adapters into a frozen base, keeping memory requirements low and adapter files portable enough to swap per domain at inference.

The Recipe

Evaluating the Results

Quality is measured with Temporal and Cross-view Sampson Error for geometric consistency, plus an LLM-as-a-judge pipeline: Cosmos Reason2 scores each generated video from 1 to 5 on physical plausibility and instruction following.

What the Fine-Tune Fixes

Before fine-tuning, the base model struggled in three ways: robot hands were out-of-distribution, causing it to hallucinate human hands in later frames; it did not reliably use the hand specified in the prompt; and the videos jittered. LoRA and DoRA address all three issues. Rank 32 notably boosts instruction following over rank 8 — the extra capacity helps the model learn precisely which hand to use and which objects to touch — while geometric and physical priors, largely captured in the frozen world model weights, improve at either rank.

What This Means

The guide lowers the barrier to domain-specific world models: one 80GB GPU, an afternoon of multi-GPU time, and a portable safetensors adapter file. For robotics teams, that translates into synthetic trajectories at scale — a practical path to sim-to-real data generation without renting a fleet.

Frequently Asked Questions

How long does it take to fine-tune Cosmos Predict 2.5 for robot video generation?

About 100 epochs, which takes roughly 17 hours on a single 80GB H100 GPU or 2.5 hours across 8 H100 GPUs, using the 92-video GR1-100 robot manipulation dataset from NVIDIA.

What is the difference between LoRA and DoRA for Cosmos fine-tuning?

LoRA injects low-rank updates into the frozen model, while DoRA additionally decomposes each weight into magnitude and direction before the low-rank update, which can stabilize training. In NVIDIA's tests both converge to similar quality, with DoRA a reasonable fallback when LoRA is unstable at low rank.

How is fine-tuned Cosmos video quality evaluated?

NVIDIA uses Temporal and Cross-view Sampson Error for geometric consistency, plus an LLM-as-a-judge pipeline where Cosmos Reason2 scores physical plausibility and instruction following from 1 to 5.

What hardware do I need to fine-tune Cosmos Predict 2.5?

At minimum one 80GB GPU for single-GPU training, with 8 H100s recommended for faster iteration. The stack requires Python 3.10+, PyTorch 2.5+ with CUDA, and the diffusers, accelerate, and peft libraries.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news