NVIDIA and Hugging Face have published a step-by-step guide for fine-tuning Cosmos Predict 2.5, the 2B-parameter world model that generates physically plausible videos from text, images, or clips — using LoRA and DoRA adapters to specialize it for robot manipulation without the cost and risks of full fine-tuning.
Why Fine-Tune a World Model
Robot policies need demonstration data, but real-robot trajectories are slow and expensive to collect. A fine-tuned world model can synthesize them at scale. Full fine-tuning of a 2B model, however, is expensive and risks catastrophic forgetting of general knowledge. LoRA and DoRA inject small trainable adapters into a frozen base, keeping memory requirements low and adapter files portable enough to swap per domain at inference.
The Recipe
- Data: 92 robot manipulation videos from the GR1-100 dataset with text prompts describing pick-and-place tasks, plus 50 prompt-image test pairs for evaluation
- Training: adapters injected into the DiT’s attention projections and feedforward layers while the VAE, text encoder, and base DiT weights stay frozen; rank 32 adds roughly 50M trainable parameters
- Loss: rectified flow — the model learns to predict the velocity that transports noise toward clean data, with the first two video frames used as conditioning
- Compute: 100 epochs takes 17 hours on a single H100, or 2.5 hours on 8 H100s, with bf16 mixed precision
Evaluating the Results
Quality is measured with Temporal and Cross-view Sampson Error for geometric consistency, plus an LLM-as-a-judge pipeline: Cosmos Reason2 scores each generated video from 1 to 5 on physical plausibility and instruction following.
What the Fine-Tune Fixes
Before fine-tuning, the base model struggled in three ways: robot hands were out-of-distribution, causing it to hallucinate human hands in later frames; it did not reliably use the hand specified in the prompt; and the videos jittered. LoRA and DoRA address all three issues. Rank 32 notably boosts instruction following over rank 8 — the extra capacity helps the model learn precisely which hand to use and which objects to touch — while geometric and physical priors, largely captured in the frozen world model weights, improve at either rank.
What This Means
The guide lowers the barrier to domain-specific world models: one 80GB GPU, an afternoon of multi-GPU time, and a portable safetensors adapter file. For robotics teams, that translates into synthetic trajectories at scale — a practical path to sim-to-real data generation without renting a fleet.