Cosmos3-Edge on LIBERO-10: a robot-arm policy

NVIDIA's Cosmos3-Edge (4B), post-trained into a robot-arm policy for the LIBERO-10 benchmark. In 500 closed-loop episodes it completes 407 tasks (81.4 %). The released checkpoint, run the same way, completes 0. The recipe is proposed upstream as NVIDIA/cosmos-framework PR #278.

The last frame of one episode for each of the ten LIBERO-10 tasks, all marked SUCCESS

One episode per task with this checkpoint: the final frame of each, instruction on top. Full results, videos and the training curve: results page.

What it is

  • LIBERO-10 (also called LIBERO-Long) is a standard simulated benchmark: ten long-horizon kitchen and tabletop tasks for a Franka Panda arm, each given as a sentence ("put both moka pots on the stove"). Most need two pick-and-place moves in one episode.
  • The model is the Cosmos3-Edge world model (4B parameters, Mixture-of-Transformers) with its LIBERO action input and output layers trained by this recipe. Those layers start fresh; everything else starts from the released Cosmos3-Edge.
  • Inputs: the instruction and two cameras, agent-view and wrist, at 256 px. Output: 10-D actions (frame-wise relative position, 6-D rotation, gripper) at 20 fps, sampled with 30 denoising steps.

Results

Closed loop in the LIBERO simulator: 50 trials per task, seed 0, success = the task's own completion check at episode end (520-step limit).

# Instruction Success
0 put both the alphabet soup and the tomato sauce in the basket 17 / 50
1 put both the cream cheese box and the butter in the basket 49 / 50
2 turn on the stove and put the moka pot on it 49 / 50
3 put the black bowl in the bottom drawer of the cabinet and close it 49 / 50
4 put the white mug on the left plate and put the yellow and white mug on the right plate 44 / 50
5 pick up the book and place it in the back compartment of the caddy 46 / 50
6 put the white mug on the plate and put the chocolate pudding to the right of the plate 39 / 50
7 put both the alphabet soup and the cream cheese box in the basket 46 / 50
8 put both moka pots on the stove 25 / 50
9 put the yellow and white mug in the microwave and close it 43 / 50
Total 407 / 500 = 81.4 %

Seven of the ten tasks are above 85 %. The two weak ones (0 and 8) both need two pick-and-place moves of similar objects in one episode.

Without this post-training: 0 / 500. The same evaluation on the released nvidia/Cosmos3-Edge checkpoint (same server, sampler, cameras and simulator settings) never completes a task. Its LIBERO action slot is still at its initial values, so this is a floor, not a measurement of zero-shot skill.

Same task and start state: the fine-tuned policy completes it (SUCCESS), the released checkpoint does not (FAIL)

Same task and start state for every task: left, this checkpoint; right, the released Cosmos3-Edge.

How it compares

Published LIBERO-10 success rates of other policies, each post-trained on LIBERO by its authors. Inputs, training data, sizes and trial counts differ, so this is context, not a controlled comparison.

Model Size LIBERO-10 Source
Cosmos Policy (NVIDIA) 2B 97.6 % model card
OpenVLA-OFT 7B 94.5 % Kim et al. 2025
GR00T N1.6 3.3B 94.4 % Isaac GR00T docs (200 episodes)
π0.5 ≈3B 92.4 % openpi
π0 3.3B 85.2 % via OpenVLA-OFT
Cosmos3-Edge, this checkpoint 4B 81.4 % this run, 500 episodes
DiT Policy 334M 63.8 % via OpenVLA-OFT
π0-FAST ≈3B 60.2 % via OpenVLA-OFT
OpenVLA 7B 53.7 % via OpenVLA-OFT
Octo 93M 51.1 % via OpenVLA-OFT
Diffusion Policy n/a 50.5 % via OpenVLA-OFT
Cosmos3-Edge, released (no LIBERO training) 4B 0.0 % measured here (0 / 500)

Scale to keep in mind: NVIDIA's Cosmos Policy trains on all four LIBERO suites, with a wrist camera and robot state, for 40,000 steps on 64 H100s, and reports 3 seeds of 500 trials. This checkpoint trains on LIBERO-10 only, for 2,000 iterations on 4 GPUs, evaluated with one seed of 500 trials.

Training

Base model nvidia/Cosmos3-Edge (4B, Mixture-of-Transformers)
Recipe cosmos-framework experiment action_policy_libero_edge (PR #278)
Data LIBERO-10 demonstrations, LeRobot v3 format
Iterations 2,000; global batch 2,048 (128 × 4 GPUs × gradient accumulation 4), about 4.1 million samples
Optimizer learning rate 5e-5, 500 warm-up steps
Precision bfloat16, FSDP shard 4
Hardware 4 × NVIDIA RTX PRO 6000 Blackwell (96 GB)
Wall time 3 days 4 hours, about 304 GPU-hours
Energy at most about 182 kWh (upper bound: 304 GPU-hours × 600 W board power; the real draw is lower)

The upstream recipe runs on two 8-GPU nodes (HSDP 2 × 8, no gradient accumulation). We kept its experiment, learning-rate schedule and global batch, and changed only the parallelism to fit four GPUs.

Use

Files: model/ is the iteration-2000 checkpoint in PyTorch distributed-checkpoint format (4 shards of 6.7 GB); config.yaml is the training config with local paths replaced by <local>.

Serve it with cosmos-framework's LIBERO policy server, then run the closed-loop evaluation:

# policy server (cosmos-framework, PR #278 recipe); --checkpoint-path is the folder that contains model/
python -m cosmos_framework.scripts.action_policy_server_libero \
  --checkpoint-path <this repo> --config-file <this repo>/config.yaml \
  --action-normalization quantile_rot \
  --action-stats-path <cosmos-framework>/.../libero_native_frame_wise_relative_rot6d.json \
  --raw-action-dim 10 --fps 20

# simulator client
python cosmos_framework/simulation/libero/closed_loop_eval.py \
  --server_url http://localhost:8000 --task_suite libero_10 \
  --num_trials_per_task 50 --num_envs 8 \
  --camera agentview,wrist --image_size 256 \
  --action_space frame_wise_relative --rotation_space 6d --action_dim 10

Setup note: the LIBERO environment pulls in egl-probe, which does not build with CMake 4. Setting CMAKE_POLICY_VERSION_MINIMUM=3.5 fixes it.

Limitations

  • Simulation only. LIBERO runs in simulation; this checkpoint has not been tried on a real arm.
  • One training run, one evaluation seed. There is no estimate of run-to-run spread yet.
  • The comparison table mixes setups. Read it as context.
  • Two weak tasks (17 / 50 and 25 / 50), both double pick-and-place of similar objects.

License

Derived from nvidia/Cosmos3-Edge; use is subject to the base model's license (OpenMDW 1.1). A simulation research checkpoint.

UB Robotics · team UBR Stack · built during the NVIDIA Open Models Codefest 2026.

Downloads last month
-
Video Preview
loading

Model tree for ubr-physical-ai/cosmos3-edge-libero10

Finetuned
(16)
this model

Space using ubr-physical-ai/cosmos3-edge-libero10 1

Paper for ubr-physical-ai/cosmos3-edge-libero10