Instructions to use ubr-physical-ai/cosmos3-edge-libero10 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Cosmos
How to use ubr-physical-ai/cosmos3-edge-libero10 with Cosmos:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Model card: what it is, per-task results, video preview (replay.mp4), stills, comparison, training, use, limitations
Browse files- .gitattributes +3 -0
- README.md +129 -15
- images/finetuned_vs_vanilla.jpg +0 -0
- images/libero10_tasks.jpg +3 -0
- media/libero10_finetuned_vs_vanilla.mp4 +3 -0
- replay.mp4 +3 -0
.gitattributes
CHANGED
|
@@ -38,3 +38,6 @@ model/.metadata filter=lfs diff=lfs merge=lfs -text
|
|
| 38 |
model/__0_0.distcp filter=lfs diff=lfs merge=lfs -text
|
| 39 |
model/__1_0.distcp filter=lfs diff=lfs merge=lfs -text
|
| 40 |
model/__2_0.distcp filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
model/__0_0.distcp filter=lfs diff=lfs merge=lfs -text
|
| 39 |
model/__1_0.distcp filter=lfs diff=lfs merge=lfs -text
|
| 40 |
model/__2_0.distcp filter=lfs diff=lfs merge=lfs -text
|
| 41 |
+
images/libero10_tasks.jpg filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
media/libero10_finetuned_vs_vanilla.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
replay.mp4 filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -2,37 +2,151 @@
|
|
| 2 |
license: other
|
| 3 |
license_name: openmdw-1.1
|
| 4 |
base_model: nvidia/Cosmos3-Edge
|
|
|
|
| 5 |
tags:
|
| 6 |
- robotics
|
| 7 |
- libero
|
| 8 |
- cosmos
|
| 9 |
- action-policy
|
|
|
|
|
|
|
| 10 |
---
|
| 11 |
|
| 12 |
-
# Cosmos3-Edge
|
| 13 |
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
| | |
|
| 19 |
|---|---|
|
| 20 |
-
|
|
| 21 |
-
|
|
| 22 |
-
|
|
| 23 |
-
|
|
| 24 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
-
|
|
|
|
| 27 |
|
| 28 |
## Use
|
| 29 |
|
| 30 |
-
|
| 31 |
-
--action-normalization quantile_rot --action-stats-path .../libero_native_frame_wise_relative_rot6d.json
|
| 32 |
-
--raw-action-dim 10 --fps 20`) and run `cosmos_framework/simulation/libero/closed_loop_eval.py` with
|
| 33 |
-
`--camera agentview,wrist --image_size 256 --action_space frame_wise_relative --rotation_space 6d --action_dim 10`.
|
| 34 |
`config.yaml` is the training config with local paths replaced by `<local>`.
|
| 35 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
## License
|
| 37 |
|
| 38 |
-
Derived from nvidia/Cosmos3-Edge; use is subject to the base model's license (OpenMDW 1.1).
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
license: other
|
| 3 |
license_name: openmdw-1.1
|
| 4 |
base_model: nvidia/Cosmos3-Edge
|
| 5 |
+
pipeline_tag: robotics
|
| 6 |
tags:
|
| 7 |
- robotics
|
| 8 |
- libero
|
| 9 |
- cosmos
|
| 10 |
- action-policy
|
| 11 |
+
- franka
|
| 12 |
+
- vision-language-action
|
| 13 |
---
|
| 14 |
|
| 15 |
+
# Cosmos3-Edge on LIBERO-10: a robot-arm policy
|
| 16 |
|
| 17 |
+
NVIDIA's **Cosmos3-Edge** (4B), post-trained into a robot-arm policy for the **LIBERO-10** benchmark. In 500 closed-loop
|
| 18 |
+
episodes it completes **407 tasks (81.4 %)**. The released checkpoint, run the same way, completes 0. The recipe is
|
| 19 |
+
proposed upstream as [NVIDIA/cosmos-framework PR #278](https://github.com/NVIDIA/cosmos-framework/pull/278).
|
| 20 |
+
|
| 21 |
+

|
| 22 |
+
|
| 23 |
+
*One episode per task with this checkpoint: the final frame of each, instruction on top. Full results, videos and the
|
| 24 |
+
training curve: [results page](https://huggingface.co/spaces/ubr-physical-ai/cosmos3-edge-libero10).*
|
| 25 |
+
|
| 26 |
+
<video controls muted playsinline src="https://huggingface.co/ubr-physical-ai/cosmos3-edge-libero10/resolve/main/replay.mp4"></video>
|
| 27 |
+
|
| 28 |
+
*The ten tasks, one episode each, as the policy ran them (simulation, agent-view camera).*
|
| 29 |
+
|
| 30 |
+
## What it is
|
| 31 |
+
|
| 32 |
+
- **LIBERO-10** (also called LIBERO-Long) is a standard simulated benchmark: ten long-horizon kitchen and tabletop
|
| 33 |
+
tasks for a Franka Panda arm, each given as a sentence ("put both moka pots on the stove"). Most need two
|
| 34 |
+
pick-and-place moves in one episode.
|
| 35 |
+
- **The model** is the Cosmos3-Edge world model (4B parameters, Mixture-of-Transformers) with its LIBERO action input
|
| 36 |
+
and output layers trained by this recipe. Those layers start fresh; everything else starts from the released
|
| 37 |
+
Cosmos3-Edge.
|
| 38 |
+
- **Inputs:** the instruction and two cameras, agent-view and wrist, at 256 px. **Output:** 10-D actions
|
| 39 |
+
(frame-wise relative position, 6-D rotation, gripper) at 20 fps, sampled with 30 denoising steps.
|
| 40 |
+
|
| 41 |
+
## Results
|
| 42 |
+
|
| 43 |
+
Closed loop in the LIBERO simulator: 50 trials per task, seed 0, success = the task's own completion check at episode
|
| 44 |
+
end (520-step limit).
|
| 45 |
+
|
| 46 |
+
| # | Instruction | Success |
|
| 47 |
+
|---|---|---|
|
| 48 |
+
| 0 | put both the alphabet soup and the tomato sauce in the basket | 17 / 50 |
|
| 49 |
+
| 1 | put both the cream cheese box and the butter in the basket | 49 / 50 |
|
| 50 |
+
| 2 | turn on the stove and put the moka pot on it | 49 / 50 |
|
| 51 |
+
| 3 | put the black bowl in the bottom drawer of the cabinet and close it | 49 / 50 |
|
| 52 |
+
| 4 | put the white mug on the left plate and put the yellow and white mug on the right plate | 44 / 50 |
|
| 53 |
+
| 5 | pick up the book and place it in the back compartment of the caddy | 46 / 50 |
|
| 54 |
+
| 6 | put the white mug on the plate and put the chocolate pudding to the right of the plate | 39 / 50 |
|
| 55 |
+
| 7 | put both the alphabet soup and the cream cheese box in the basket | 46 / 50 |
|
| 56 |
+
| 8 | put both moka pots on the stove | 25 / 50 |
|
| 57 |
+
| 9 | put the yellow and white mug in the microwave and close it | 43 / 50 |
|
| 58 |
+
| | **Total** | **407 / 500 = 81.4 %** |
|
| 59 |
+
|
| 60 |
+
Seven of the ten tasks are above 85 %. The two weak ones (0 and 8) both need two pick-and-place moves of similar
|
| 61 |
+
objects in one episode.
|
| 62 |
+
|
| 63 |
+
**Without this post-training: 0 / 500.** The same evaluation on the released `nvidia/Cosmos3-Edge` checkpoint
|
| 64 |
+
(same server, sampler, cameras and simulator settings) never completes a task. Its LIBERO action slot is still at its
|
| 65 |
+
initial values, so this is a floor, not a measurement of zero-shot skill.
|
| 66 |
+
|
| 67 |
+

|
| 68 |
+
|
| 69 |
+
<video controls muted playsinline src="https://huggingface.co/ubr-physical-ai/cosmos3-edge-libero10/resolve/main/media/libero10_finetuned_vs_vanilla.mp4"></video>
|
| 70 |
+
|
| 71 |
+
*Same task and start state for every task: left, this checkpoint; right, the released Cosmos3-Edge.*
|
| 72 |
+
|
| 73 |
+
## How it compares
|
| 74 |
+
|
| 75 |
+
Published LIBERO-10 success rates of other policies, each post-trained on LIBERO by its authors. Inputs, training
|
| 76 |
+
data, sizes and trial counts differ, so this is **context, not a controlled comparison**.
|
| 77 |
+
|
| 78 |
+
| Model | Size | LIBERO-10 | Source |
|
| 79 |
+
|---|---|---|---|
|
| 80 |
+
| Cosmos Policy (NVIDIA) | 2B | 97.6 % | [model card](https://huggingface.co/nvidia/Cosmos-Policy-LIBERO-Predict2-2B) |
|
| 81 |
+
| OpenVLA-OFT | 7B | 94.5 % | [Kim et al. 2025](https://arxiv.org/abs/2502.19645) |
|
| 82 |
+
| GR00T N1.6 | 3.3B | 94.4 % | [Isaac GR00T docs](https://nvidia-isaac-gr00t.mintlify.app/examples/libero) (200 episodes) |
|
| 83 |
+
| π0.5 | ≈3B | 92.4 % | [openpi](https://github.com/Physical-Intelligence/openpi/blob/main/examples/libero/README.md) |
|
| 84 |
+
| π0 | 3.3B | 85.2 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
|
| 85 |
+
| **Cosmos3-Edge, this checkpoint** | **4B** | **81.4 %** | this run, 500 episodes |
|
| 86 |
+
| DiT Policy | 334M | 63.8 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
|
| 87 |
+
| π0-FAST | ≈3B | 60.2 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
|
| 88 |
+
| OpenVLA | 7B | 53.7 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
|
| 89 |
+
| Octo | 93M | 51.1 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
|
| 90 |
+
| Diffusion Policy | n/a | 50.5 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
|
| 91 |
+
| Cosmos3-Edge, released (no LIBERO training) | 4B | 0.0 % | measured here (0 / 500) |
|
| 92 |
+
|
| 93 |
+
Scale to keep in mind: NVIDIA's Cosmos Policy trains on all four LIBERO suites, with a wrist camera and robot state,
|
| 94 |
+
for 40,000 steps on 64 H100s, and reports 3 seeds of 500 trials. This checkpoint trains on LIBERO-10 only, for 2,000
|
| 95 |
+
iterations on 4 GPUs, evaluated with one seed of 500 trials.
|
| 96 |
+
|
| 97 |
+
## Training
|
| 98 |
|
| 99 |
| | |
|
| 100 |
|---|---|
|
| 101 |
+
| Base model | `nvidia/Cosmos3-Edge` (4B, Mixture-of-Transformers) |
|
| 102 |
+
| Recipe | cosmos-framework experiment `action_policy_libero_edge` ([PR #278](https://github.com/NVIDIA/cosmos-framework/pull/278)) |
|
| 103 |
+
| Data | LIBERO-10 demonstrations, LeRobot v3 format |
|
| 104 |
+
| Iterations | 2,000; global batch 2,048 (128 × 4 GPUs × gradient accumulation 4), about 4.1 million samples |
|
| 105 |
+
| Optimizer | learning rate 5e-5, 500 warm-up steps |
|
| 106 |
+
| Precision | bfloat16, FSDP shard 4 |
|
| 107 |
+
| Hardware | 4 × NVIDIA RTX PRO 6000 Blackwell (96 GB) |
|
| 108 |
+
| Wall time | 3 days 4 hours, about 304 GPU-hours |
|
| 109 |
+
| Energy | at most about 182 kWh (upper bound: 304 GPU-hours × 600 W board power; the real draw is lower) |
|
| 110 |
|
| 111 |
+
The upstream recipe runs on two 8-GPU nodes (HSDP 2 × 8, no gradient accumulation). We kept its experiment,
|
| 112 |
+
learning-rate schedule and global batch, and changed only the parallelism to fit four GPUs.
|
| 113 |
|
| 114 |
## Use
|
| 115 |
|
| 116 |
+
Files: `model/` is the iteration-2000 checkpoint in PyTorch distributed-checkpoint format (4 shards of 6.7 GB);
|
|
|
|
|
|
|
|
|
|
| 117 |
`config.yaml` is the training config with local paths replaced by `<local>`.
|
| 118 |
|
| 119 |
+
Serve it with cosmos-framework's LIBERO policy server, then run the closed-loop evaluation:
|
| 120 |
+
|
| 121 |
+
```bash
|
| 122 |
+
# policy server (cosmos-framework, PR #278 recipe)
|
| 123 |
+
python -m cosmos_framework.scripts.action_policy_server_libero \
|
| 124 |
+
--checkpoint-path <this repo>/model --config-file <this repo>/config.yaml \
|
| 125 |
+
--action-normalization quantile_rot \
|
| 126 |
+
--action-stats-path <cosmos-framework>/.../libero_native_frame_wise_relative_rot6d.json \
|
| 127 |
+
--raw-action-dim 10 --fps 20
|
| 128 |
+
|
| 129 |
+
# simulator client
|
| 130 |
+
python cosmos_framework/simulation/libero/closed_loop_eval.py \
|
| 131 |
+
--server_url http://localhost:8000 --task_suite libero_10 \
|
| 132 |
+
--num_trials_per_task 50 --num_envs 8 \
|
| 133 |
+
--camera agentview,wrist --image_size 256 \
|
| 134 |
+
--action_space frame_wise_relative --rotation_space 6d --action_dim 10
|
| 135 |
+
```
|
| 136 |
+
|
| 137 |
+
Setup note: the LIBERO environment pulls in `egl-probe`, which does not build with CMake 4. Setting
|
| 138 |
+
`CMAKE_POLICY_VERSION_MINIMUM=3.5` fixes it.
|
| 139 |
+
|
| 140 |
+
## Limitations
|
| 141 |
+
|
| 142 |
+
- **Simulation only.** LIBERO runs in simulation; this checkpoint has not been tried on a real arm.
|
| 143 |
+
- **One training run, one evaluation seed.** There is no estimate of run-to-run spread yet.
|
| 144 |
+
- **The comparison table mixes setups.** Read it as context.
|
| 145 |
+
- **Two weak tasks** (17 / 50 and 25 / 50), both double pick-and-place of similar objects.
|
| 146 |
+
|
| 147 |
## License
|
| 148 |
|
| 149 |
+
Derived from `nvidia/Cosmos3-Edge`; use is subject to the base model's license (OpenMDW 1.1). A simulation research
|
| 150 |
+
checkpoint.
|
| 151 |
+
|
| 152 |
+
UB Robotics · team UBR Stack · built during the NVIDIA Open Models Codefest 2026.
|
images/finetuned_vs_vanilla.jpg
ADDED
|
images/libero10_tasks.jpg
ADDED
|
Git LFS Details
|
media/libero10_finetuned_vs_vanilla.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e27849703314b14cc80701be3b4de0270ade0bf87dfb5fc09f8c99b9da32285d
|
| 3 |
+
size 5380399
|
replay.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a6a075f0d477f248af9234f53aaec72a417daa00db914f83c7b2814d5cff0ba2
|
| 3 |
+
size 6177051
|