Image-to-3D
Transformers
Safetensors
shape-opt

Add pipeline tag and improve model card

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +27 -11
README.md CHANGED
@@ -1,22 +1,38 @@
1
  ---
2
- library_name: transformers
3
- license: cc-by-sa-4.0
4
- datasets:
5
- - zx1239856/3d-front-ar-packed
6
  base_model:
7
  - zx1239856/edgerunner
 
 
 
 
 
8
  ---
9
 
10
- # Model Card for Model ID
11
 
12
- PixARMesh based on EdgeRunner
13
 
14
  ## Model Details
15
 
16
- PixARMesh is a mesh-native autoregressive framework for single-view 3D scene reconstruction.
17
- Instead of reconstructing via intermediate volumetric or implicit representations, PixARMesh directly models instances with native mesh representation. Object poses and meshes are predicted in a unified autoregressive sequence.
18
-
19
- ### Model Sources
20
 
21
  - **Repository:** https://github.com/mlpc-ucsd/PixARMesh
22
- - **Paper:** https://arxiv.org/pdf/2603.05888
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
 
 
 
 
2
  base_model:
3
  - zx1239856/edgerunner
4
+ datasets:
5
+ - zx1239856/3d-front-ar-packed
6
+ library_name: transformers
7
+ license: cc-by-sa-4.0
8
+ pipeline_tag: image-to-3d
9
  ---
10
 
11
+ # PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
12
 
13
+ PixARMesh is a mesh-native autoregressive framework for single-view 3D scene reconstruction. Instead of reconstructing via intermediate volumetric or implicit representations, PixARMesh directly models instances with native mesh representation. Object poses and meshes are predicted in a unified autoregressive sequence.
14
 
15
  ## Model Details
16
 
17
+ Reconstructing complete 3D indoor scenes from a single RGB image is a complex task. PixARMesh jointly predicts object layout and geometry within a unified model, producing coherent and artist-ready meshes in a single forward pass. By augmenting a point-cloud encoder with pixel-aligned image features and global scene context via cross-attention, the model enables accurate spatial reasoning. Scenes are generated autoregressively from a unified token stream containing context, pose, and mesh, yielding compact meshes with high-fidelity geometry.
 
 
 
18
 
19
  - **Repository:** https://github.com/mlpc-ucsd/PixARMesh
20
+ - **Paper:** [arXiv:2603.05888](https://arxiv.org/abs/2603.05888)
21
+ - **Project Page:** https://mlpc-ucsd.github.io/PixARMesh/
22
+
23
+ ## Citation
24
+
25
+ If you find PixARMesh useful in your research, please consider citing:
26
+
27
+ ```bibtex
28
+ @article{zhang2026pixarmesh,
29
+ title={PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction},
30
+ author={Zhang, Xiang and Yoo, Sohyun and Wu, Hongrui and Li, Chuan and Xie, Jianwen and Tu, Zhuowen},
31
+ journal={arXiv preprint arXiv:2603.05888},
32
+ year={2026}
33
+ }
34
+ ```
35
+
36
+ ## Acknowledgements
37
+
38
+ PixARMesh builds upon several excellent open-source projects including [Grounded-Segment-Anything](https://github.com/IDEA-Research/Grounded-Segment-Anything), [Depth Pro](https://github.com/apple/ml-depth-pro), [DINOv2](https://github.com/facebookresearch/dinov2), and weights from [EdgeRunner](https://github.com/NVlabs/EdgeRunner) and [BPT](https://github.com/Tencent-Hunyuan/bpt). We also use physically-based renderings from the [3D-FRONT](https://tianchi.aliyun.com/specials/promotion/alibaba-3d-scene-dataset) scenes.