Instructions to use ubr-physical-ai/cosmos3-edge-libero10 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Cosmos
How to use ubr-physical-ai/cosmos3-edge-libero10 with Cosmos:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
# No code snippets available yet for this library.
# To use this model, check the repository files and the library's documentation.
# Want to help? PRs adding snippets are welcome at:
# https://github.com/huggingface/huggingface.jsCosmos3-Edge on LIBERO-10: a robot-arm policy
NVIDIA's Cosmos3-Edge (4B), post-trained into a robot-arm policy for the LIBERO-10 benchmark. In 500 closed-loop episodes it completes 407 tasks (81.4 %). The released checkpoint, run the same way, completes 0. The recipe is proposed upstream as NVIDIA/cosmos-framework PR #278.
One episode per task with this checkpoint: the final frame of each, instruction on top. Full results, videos and the training curve: results page.
What it is
- LIBERO-10 (also called LIBERO-Long) is a standard simulated benchmark: ten long-horizon kitchen and tabletop tasks for a Franka Panda arm, each given as a sentence ("put both moka pots on the stove"). Most need two pick-and-place moves in one episode.
- The model is the Cosmos3-Edge world model (4B parameters, Mixture-of-Transformers) with its LIBERO action input and output layers trained by this recipe. Those layers start fresh; everything else starts from the released Cosmos3-Edge.
- Inputs: the instruction and two cameras, agent-view and wrist, at 256 px. Output: 10-D actions (frame-wise relative position, 6-D rotation, gripper) at 20 fps, sampled with 30 denoising steps.
Results
Closed loop in the LIBERO simulator: 50 trials per task, seed 0, success = the task's own completion check at episode end (520-step limit).
| # | Instruction | Success |
|---|---|---|
| 0 | put both the alphabet soup and the tomato sauce in the basket | 17 / 50 |
| 1 | put both the cream cheese box and the butter in the basket | 49 / 50 |
| 2 | turn on the stove and put the moka pot on it | 49 / 50 |
| 3 | put the black bowl in the bottom drawer of the cabinet and close it | 49 / 50 |
| 4 | put the white mug on the left plate and put the yellow and white mug on the right plate | 44 / 50 |
| 5 | pick up the book and place it in the back compartment of the caddy | 46 / 50 |
| 6 | put the white mug on the plate and put the chocolate pudding to the right of the plate | 39 / 50 |
| 7 | put both the alphabet soup and the cream cheese box in the basket | 46 / 50 |
| 8 | put both moka pots on the stove | 25 / 50 |
| 9 | put the yellow and white mug in the microwave and close it | 43 / 50 |
| Total | 407 / 500 = 81.4 % |
Seven of the ten tasks are above 85 %. The two weak ones (0 and 8) both need two pick-and-place moves of similar objects in one episode.
Without this post-training: 0 / 500. The same evaluation on the released nvidia/Cosmos3-Edge checkpoint
(same server, sampler, cameras and simulator settings) never completes a task. Its LIBERO action slot is still at its
initial values, so this is a floor, not a measurement of zero-shot skill.
Same task and start state for every task: left, this checkpoint; right, the released Cosmos3-Edge.
How it compares
Published LIBERO-10 success rates of other policies, each post-trained on LIBERO by its authors. Inputs, training data, sizes and trial counts differ, so this is context, not a controlled comparison.
| Model | Size | LIBERO-10 | Source |
|---|---|---|---|
| Cosmos Policy (NVIDIA) | 2B | 97.6 % | model card |
| OpenVLA-OFT | 7B | 94.5 % | Kim et al. 2025 |
| GR00T N1.6 | 3.3B | 94.4 % | Isaac GR00T docs (200 episodes) |
| π0.5 | ≈3B | 92.4 % | openpi |
| π0 | 3.3B | 85.2 % | via OpenVLA-OFT |
| Cosmos3-Edge, this checkpoint | 4B | 81.4 % | this run, 500 episodes |
| DiT Policy | 334M | 63.8 % | via OpenVLA-OFT |
| π0-FAST | ≈3B | 60.2 % | via OpenVLA-OFT |
| OpenVLA | 7B | 53.7 % | via OpenVLA-OFT |
| Octo | 93M | 51.1 % | via OpenVLA-OFT |
| Diffusion Policy | n/a | 50.5 % | via OpenVLA-OFT |
| Cosmos3-Edge, released (no LIBERO training) | 4B | 0.0 % | measured here (0 / 500) |
Scale to keep in mind: NVIDIA's Cosmos Policy trains on all four LIBERO suites, with a wrist camera and robot state, for 40,000 steps on 64 H100s, and reports 3 seeds of 500 trials. This checkpoint trains on LIBERO-10 only, for 2,000 iterations on 4 GPUs, evaluated with one seed of 500 trials.
Training
| Base model | nvidia/Cosmos3-Edge (4B, Mixture-of-Transformers) |
| Recipe | cosmos-framework experiment action_policy_libero_edge (PR #278) |
| Data | LIBERO-10 demonstrations, LeRobot v3 format |
| Iterations | 2,000; global batch 2,048 (128 × 4 GPUs × gradient accumulation 4), about 4.1 million samples |
| Optimizer | learning rate 5e-5, 500 warm-up steps |
| Precision | bfloat16, FSDP shard 4 |
| Hardware | 4 × NVIDIA RTX PRO 6000 Blackwell (96 GB) |
| Wall time | 3 days 4 hours, about 304 GPU-hours |
| Energy | at most about 182 kWh (upper bound: 304 GPU-hours × 600 W board power; the real draw is lower) |
The upstream recipe runs on two 8-GPU nodes (HSDP 2 × 8, no gradient accumulation). We kept its experiment, learning-rate schedule and global batch, and changed only the parallelism to fit four GPUs.
Use
Files: model/ is the iteration-2000 checkpoint in PyTorch distributed-checkpoint format (4 shards of 6.7 GB);
config.yaml is the training config with local paths replaced by <local>.
Serve it with cosmos-framework's LIBERO policy server, then run the closed-loop evaluation:
# policy server (cosmos-framework, PR #278 recipe); --checkpoint-path is the folder that contains model/
python -m cosmos_framework.scripts.action_policy_server_libero \
--checkpoint-path <this repo> --config-file <this repo>/config.yaml \
--action-normalization quantile_rot \
--action-stats-path <cosmos-framework>/.../libero_native_frame_wise_relative_rot6d.json \
--raw-action-dim 10 --fps 20
# simulator client
python cosmos_framework/simulation/libero/closed_loop_eval.py \
--server_url http://localhost:8000 --task_suite libero_10 \
--num_trials_per_task 50 --num_envs 8 \
--camera agentview,wrist --image_size 256 \
--action_space frame_wise_relative --rotation_space 6d --action_dim 10
Setup note: the LIBERO environment pulls in egl-probe, which does not build with CMake 4. Setting
CMAKE_POLICY_VERSION_MINIMUM=3.5 fixes it.
Limitations
- Simulation only. LIBERO runs in simulation; this checkpoint has not been tried on a real arm.
- One training run, one evaluation seed. There is no estimate of run-to-run spread yet.
- The comparison table mixes setups. Read it as context.
- Two weak tasks (17 / 50 and 25 / 50), both double pick-and-place of similar objects.
License
Derived from nvidia/Cosmos3-Edge; use is subject to the base model's license (OpenMDW 1.1). A simulation research
checkpoint.
UB Robotics · team UBR Stack · built during the NVIDIA Open Models Codefest 2026.
- Downloads last month
- -
Model tree for ubr-physical-ai/cosmos3-edge-libero10
Base model
nvidia/Cosmos3-Edge
