XPENG AI
AI & ML interests
LLM, VLM, Omni Model, Agent, VLA
Recent Activity
Papers
VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
XPENG AI
Open-source models and research for intelligent mobility.
Official Website · GitHub · GroundingPI · GroundAnything · X-AuT · OmniGUI
About
XPENG AI shares selected models, research, and practical tools with the open-source community. Our work explores efficient AI systems for intelligent mobility, speech, and multimodal interaction.
Featured Projects
GroundingPI
A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
GroundingPI is a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. It sets a new state of the art across 34 grounding benchmarks (averaging 73.68%, above GPT-6 Astra) and improves robot manipulation and autonomous driving when used as the visual backbone.
| Resource | Link |
|---|---|
| Model weights | GroundingPI/GroundingPI |
| Online demo | GroundingPI/GroundingPI Space |
| Project website | groundingpi.github.io |
| Paper | arXiv:2609.39601 |
| Source code | groundingpi/GroundingPI |
GroundAnything
Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
GroundAnything combines parallel decoding with precise visual grounding for flash-speed inference. It is released in two variants: a diffusion model (GroundAnything) and an autoregressive VLM (GroundAnything-VLM), with an interactive demo Space.
| Resource | Link |
|---|---|
| Model weights (VLM) | GroundingPI/GroundAnything-VLM |
| Model weights (DLM) | GroundingPI/GroundAnything |
| Project website | groundingpi.github.io/groundanything |
| Paper | arXiv:2609.39600 |
| Source code | groundingpi/GroundAnything |
X-AuT
Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
X-AuT is a compact automatic speech recognition model based on Qwen3-ASR-0.6B. It reduces the audio encoder from 18 to 14 Transformer blocks through progressive compression and cross-scale knowledge transfer.
The release includes full model weights, standalone inference, and a compact LoRA finetuning example.
| Resource | Link |
|---|---|
| Model weights | XPENG-AI/X-AuT |
| Project website | xpeng-ai.github.io/x-aut |
| Paper | arXiv:2609.11412 |
| Source code | XPENG-AI/X-AuT |
OmniGUI
Benchmarking GUI Agents in Omni-Modal Smartphone Environments
OmniGUI is a step-level benchmark for GUI agents operating in smartphone environments with interleaved screenshots, audio, video, and action history. It evaluates localization, semantic understanding, cross-modal discrimination, temporal reasoning, and instant response across 29 applications.
| Resource | Link |
|---|---|
| Project website | omni-gui.github.io |
| Paper | arXiv:2605.18758 |
| Source code | omni-gui/OmniGUI |
| Dataset | OmniGUI/OmniGUI |
More to Come
New models and research projects will be added here as they become publicly available.