filipemiguelmartins commited on
Commit
edf16f9
·
verified ·
1 Parent(s): 1599831

Model card: what it is, per-task results, video preview (replay.mp4), stills, comparison, training, use, limitations

Browse files
.gitattributes CHANGED
@@ -38,3 +38,6 @@ model/.metadata filter=lfs diff=lfs merge=lfs -text
38
  model/__0_0.distcp filter=lfs diff=lfs merge=lfs -text
39
  model/__1_0.distcp filter=lfs diff=lfs merge=lfs -text
40
  model/__2_0.distcp filter=lfs diff=lfs merge=lfs -text
 
 
 
 
38
  model/__0_0.distcp filter=lfs diff=lfs merge=lfs -text
39
  model/__1_0.distcp filter=lfs diff=lfs merge=lfs -text
40
  model/__2_0.distcp filter=lfs diff=lfs merge=lfs -text
41
+ images/libero10_tasks.jpg filter=lfs diff=lfs merge=lfs -text
42
+ media/libero10_finetuned_vs_vanilla.mp4 filter=lfs diff=lfs merge=lfs -text
43
+ replay.mp4 filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -2,37 +2,151 @@
2
  license: other
3
  license_name: openmdw-1.1
4
  base_model: nvidia/Cosmos3-Edge
 
5
  tags:
6
  - robotics
7
  - libero
8
  - cosmos
9
  - action-policy
 
 
10
  ---
11
 
12
- # Cosmos3-Edge post-trained on LIBERO-10 (action policy)
13
 
14
- `nvidia/Cosmos3-Edge` post-trained as a robot-arm policy on the LIBERO-10 benchmark with the recipe proposed in
15
- [NVIDIA/cosmos-framework PR #278](https://github.com/NVIDIA/cosmos-framework/pull/278) (experiment
16
- `action_policy_libero_edge`). Results page: https://huggingface.co/spaces/ubr-physical-ai/cosmos3-edge-libero10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
17
 
18
  | | |
19
  |---|---|
20
- | Closed-loop success, LIBERO-10 | **407 / 500 = 81.4 %** (50 trials per task, seed 0) |
21
- | Same eval, Cosmos3-Edge without this post-training | 0 / 500 (its LIBERO action slot is untrained) |
22
- | Checkpoint | `iter_000002000/model` (DCP), 2,000 iterations |
23
- | Training | 4 x RTX PRO 6000 Blackwell, global batch 2,048 (128 x 4 GPUs x grad-accum 4), lr 5e-5, 500 warm-up, bf16, ~3 days |
24
- | Data | LIBERO-10 in LeRobot v3 format |
 
 
 
 
25
 
26
- Per task (successes / 50): 17, 49, 49, 49, 44, 46, 39, 46, 25, 43.
 
27
 
28
  ## Use
29
 
30
- Serve with cosmos-framework's `action_policy_server_libero` (`--checkpoint-path iter_000002000 --config-file config.yaml
31
- --action-normalization quantile_rot --action-stats-path .../libero_native_frame_wise_relative_rot6d.json
32
- --raw-action-dim 10 --fps 20`) and run `cosmos_framework/simulation/libero/closed_loop_eval.py` with
33
- `--camera agentview,wrist --image_size 256 --action_space frame_wise_relative --rotation_space 6d --action_dim 10`.
34
  `config.yaml` is the training config with local paths replaced by `<local>`.
35
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  ## License
37
 
38
- Derived from nvidia/Cosmos3-Edge; use is subject to the base model's license (OpenMDW 1.1). Simulation research checkpoint.
 
 
 
 
2
  license: other
3
  license_name: openmdw-1.1
4
  base_model: nvidia/Cosmos3-Edge
5
+ pipeline_tag: robotics
6
  tags:
7
  - robotics
8
  - libero
9
  - cosmos
10
  - action-policy
11
+ - franka
12
+ - vision-language-action
13
  ---
14
 
15
+ # Cosmos3-Edge on LIBERO-10: a robot-arm policy
16
 
17
+ NVIDIA's **Cosmos3-Edge** (4B), post-trained into a robot-arm policy for the **LIBERO-10** benchmark. In 500 closed-loop
18
+ episodes it completes **407 tasks (81.4 %)**. The released checkpoint, run the same way, completes 0. The recipe is
19
+ proposed upstream as [NVIDIA/cosmos-framework PR #278](https://github.com/NVIDIA/cosmos-framework/pull/278).
20
+
21
+ ![The last frame of one episode for each of the ten LIBERO-10 tasks, all marked SUCCESS](images/libero10_tasks.jpg)
22
+
23
+ *One episode per task with this checkpoint: the final frame of each, instruction on top. Full results, videos and the
24
+ training curve: [results page](https://huggingface.co/spaces/ubr-physical-ai/cosmos3-edge-libero10).*
25
+
26
+ <video controls muted playsinline src="https://huggingface.co/ubr-physical-ai/cosmos3-edge-libero10/resolve/main/replay.mp4"></video>
27
+
28
+ *The ten tasks, one episode each, as the policy ran them (simulation, agent-view camera).*
29
+
30
+ ## What it is
31
+
32
+ - **LIBERO-10** (also called LIBERO-Long) is a standard simulated benchmark: ten long-horizon kitchen and tabletop
33
+ tasks for a Franka Panda arm, each given as a sentence ("put both moka pots on the stove"). Most need two
34
+ pick-and-place moves in one episode.
35
+ - **The model** is the Cosmos3-Edge world model (4B parameters, Mixture-of-Transformers) with its LIBERO action input
36
+ and output layers trained by this recipe. Those layers start fresh; everything else starts from the released
37
+ Cosmos3-Edge.
38
+ - **Inputs:** the instruction and two cameras, agent-view and wrist, at 256 px. **Output:** 10-D actions
39
+ (frame-wise relative position, 6-D rotation, gripper) at 20 fps, sampled with 30 denoising steps.
40
+
41
+ ## Results
42
+
43
+ Closed loop in the LIBERO simulator: 50 trials per task, seed 0, success = the task's own completion check at episode
44
+ end (520-step limit).
45
+
46
+ | # | Instruction | Success |
47
+ |---|---|---|
48
+ | 0 | put both the alphabet soup and the tomato sauce in the basket | 17 / 50 |
49
+ | 1 | put both the cream cheese box and the butter in the basket | 49 / 50 |
50
+ | 2 | turn on the stove and put the moka pot on it | 49 / 50 |
51
+ | 3 | put the black bowl in the bottom drawer of the cabinet and close it | 49 / 50 |
52
+ | 4 | put the white mug on the left plate and put the yellow and white mug on the right plate | 44 / 50 |
53
+ | 5 | pick up the book and place it in the back compartment of the caddy | 46 / 50 |
54
+ | 6 | put the white mug on the plate and put the chocolate pudding to the right of the plate | 39 / 50 |
55
+ | 7 | put both the alphabet soup and the cream cheese box in the basket | 46 / 50 |
56
+ | 8 | put both moka pots on the stove | 25 / 50 |
57
+ | 9 | put the yellow and white mug in the microwave and close it | 43 / 50 |
58
+ | | **Total** | **407 / 500 = 81.4 %** |
59
+
60
+ Seven of the ten tasks are above 85 %. The two weak ones (0 and 8) both need two pick-and-place moves of similar
61
+ objects in one episode.
62
+
63
+ **Without this post-training: 0 / 500.** The same evaluation on the released `nvidia/Cosmos3-Edge` checkpoint
64
+ (same server, sampler, cameras and simulator settings) never completes a task. Its LIBERO action slot is still at its
65
+ initial values, so this is a floor, not a measurement of zero-shot skill.
66
+
67
+ ![Same task and start state: the fine-tuned policy completes it (SUCCESS), the released checkpoint does not (FAIL)](images/finetuned_vs_vanilla.jpg)
68
+
69
+ <video controls muted playsinline src="https://huggingface.co/ubr-physical-ai/cosmos3-edge-libero10/resolve/main/media/libero10_finetuned_vs_vanilla.mp4"></video>
70
+
71
+ *Same task and start state for every task: left, this checkpoint; right, the released Cosmos3-Edge.*
72
+
73
+ ## How it compares
74
+
75
+ Published LIBERO-10 success rates of other policies, each post-trained on LIBERO by its authors. Inputs, training
76
+ data, sizes and trial counts differ, so this is **context, not a controlled comparison**.
77
+
78
+ | Model | Size | LIBERO-10 | Source |
79
+ |---|---|---|---|
80
+ | Cosmos Policy (NVIDIA) | 2B | 97.6 % | [model card](https://huggingface.co/nvidia/Cosmos-Policy-LIBERO-Predict2-2B) |
81
+ | OpenVLA-OFT | 7B | 94.5 % | [Kim et al. 2025](https://arxiv.org/abs/2502.19645) |
82
+ | GR00T N1.6 | 3.3B | 94.4 % | [Isaac GR00T docs](https://nvidia-isaac-gr00t.mintlify.app/examples/libero) (200 episodes) |
83
+ | π0.5 | ≈3B | 92.4 % | [openpi](https://github.com/Physical-Intelligence/openpi/blob/main/examples/libero/README.md) |
84
+ | π0 | 3.3B | 85.2 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
85
+ | **Cosmos3-Edge, this checkpoint** | **4B** | **81.4 %** | this run, 500 episodes |
86
+ | DiT Policy | 334M | 63.8 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
87
+ | π0-FAST | ≈3B | 60.2 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
88
+ | OpenVLA | 7B | 53.7 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
89
+ | Octo | 93M | 51.1 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
90
+ | Diffusion Policy | n/a | 50.5 % | via [OpenVLA-OFT](https://arxiv.org/abs/2502.19645) |
91
+ | Cosmos3-Edge, released (no LIBERO training) | 4B | 0.0 % | measured here (0 / 500) |
92
+
93
+ Scale to keep in mind: NVIDIA's Cosmos Policy trains on all four LIBERO suites, with a wrist camera and robot state,
94
+ for 40,000 steps on 64 H100s, and reports 3 seeds of 500 trials. This checkpoint trains on LIBERO-10 only, for 2,000
95
+ iterations on 4 GPUs, evaluated with one seed of 500 trials.
96
+
97
+ ## Training
98
 
99
  | | |
100
  |---|---|
101
+ | Base model | `nvidia/Cosmos3-Edge` (4B, Mixture-of-Transformers) |
102
+ | Recipe | cosmos-framework experiment `action_policy_libero_edge` ([PR #278](https://github.com/NVIDIA/cosmos-framework/pull/278)) |
103
+ | Data | LIBERO-10 demonstrations, LeRobot v3 format |
104
+ | Iterations | 2,000; global batch 2,048 (128 × 4 GPUs × gradient accumulation 4), about 4.1 million samples |
105
+ | Optimizer | learning rate 5e-5, 500 warm-up steps |
106
+ | Precision | bfloat16, FSDP shard 4 |
107
+ | Hardware | 4 × NVIDIA RTX PRO 6000 Blackwell (96 GB) |
108
+ | Wall time | 3 days 4 hours, about 304 GPU-hours |
109
+ | Energy | at most about 182 kWh (upper bound: 304 GPU-hours × 600 W board power; the real draw is lower) |
110
 
111
+ The upstream recipe runs on two 8-GPU nodes (HSDP 2 × 8, no gradient accumulation). We kept its experiment,
112
+ learning-rate schedule and global batch, and changed only the parallelism to fit four GPUs.
113
 
114
  ## Use
115
 
116
+ Files: `model/` is the iteration-2000 checkpoint in PyTorch distributed-checkpoint format (4 shards of 6.7 GB);
 
 
 
117
  `config.yaml` is the training config with local paths replaced by `<local>`.
118
 
119
+ Serve it with cosmos-framework's LIBERO policy server, then run the closed-loop evaluation:
120
+
121
+ ```bash
122
+ # policy server (cosmos-framework, PR #278 recipe)
123
+ python -m cosmos_framework.scripts.action_policy_server_libero \
124
+ --checkpoint-path <this repo>/model --config-file <this repo>/config.yaml \
125
+ --action-normalization quantile_rot \
126
+ --action-stats-path <cosmos-framework>/.../libero_native_frame_wise_relative_rot6d.json \
127
+ --raw-action-dim 10 --fps 20
128
+
129
+ # simulator client
130
+ python cosmos_framework/simulation/libero/closed_loop_eval.py \
131
+ --server_url http://localhost:8000 --task_suite libero_10 \
132
+ --num_trials_per_task 50 --num_envs 8 \
133
+ --camera agentview,wrist --image_size 256 \
134
+ --action_space frame_wise_relative --rotation_space 6d --action_dim 10
135
+ ```
136
+
137
+ Setup note: the LIBERO environment pulls in `egl-probe`, which does not build with CMake 4. Setting
138
+ `CMAKE_POLICY_VERSION_MINIMUM=3.5` fixes it.
139
+
140
+ ## Limitations
141
+
142
+ - **Simulation only.** LIBERO runs in simulation; this checkpoint has not been tried on a real arm.
143
+ - **One training run, one evaluation seed.** There is no estimate of run-to-run spread yet.
144
+ - **The comparison table mixes setups.** Read it as context.
145
+ - **Two weak tasks** (17 / 50 and 25 / 50), both double pick-and-place of similar objects.
146
+
147
  ## License
148
 
149
+ Derived from `nvidia/Cosmos3-Edge`; use is subject to the base model's license (OpenMDW 1.1). A simulation research
150
+ checkpoint.
151
+
152
+ UB Robotics · team UBR Stack · built during the NVIDIA Open Models Codefest 2026.
images/finetuned_vs_vanilla.jpg ADDED
images/libero10_tasks.jpg ADDED

Git LFS Details

  • SHA256: 9b162b00b494e09fc52650fbda9115a5f1072a201015dfc62fc57e72acbcd519
  • Pointer size: 131 Bytes
  • Size of remote file: 175 kB
media/libero10_finetuned_vs_vanilla.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e27849703314b14cc80701be3b4de0270ade0bf87dfb5fc09f8c99b9da32285d
3
+ size 5380399
replay.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a6a075f0d477f248af9234f53aaec72a417daa00db914f83c7b2814d5cff0ba2
3
+ size 6177051