Recent News

  • [2026.04] 🎉🎉 One paper is accepted by ACMMM 2026. See you in Rio de Janeiro, Brazil.
  • [2026.04] 🥳🥳 One paper is accepted by TPAMI!
  • [2026.02] 🎉🎉 One paper is accepted by CVPR 2026!
  • [2025.11] 🥳🥳 One paper is accepted by AAAI 2026. See you in Singapore!
  • [2025.03] 🎉🎉 One paper is accepted by ICME 2025!

Research Interests

0
Publications & Preprints
0
Citations
0
1st Author

Services

  • Program Committee Member: ICME 2025, CVPR 2026, etc.

Publications & Preprints

Full publication and BibTeX details are also available on Google Scholar.

Technical Report
The first cross-embodied foundation model to integrate autonomous driving and embodied AI, setting new records across 17 embodied and 12 driving benchmarks.
MiMo-Embodied

arXiv 2026
Combines outcome rewards with process-level guidance for critic-free policy optimization, improving multi-step LLM reasoning and credit assignment.
PRPO

ACM MM 2026
A unified driving-safety benchmark covering 10 external-risk and in-cabin categories, with 98K instances for evaluating and improving VLM safety.
DSBench

Preprint 2026
A survey of VLA evaluation that identifies key benchmarking bottlenecks and systematically examines how existing protocols measure vision-language-action models.
The Evaluation Bottleneck of Vision-Language-Action Models

TPAMI 2026
Studies category-level articulated-object pose perception and introduces a rotation-decoupled strategy for accurate, efficient 6D pose estimation and tracking.
Probing Effective and Efficient Category-Level Articulated Object Pose Perception

CVPR 2026
Models category-level articulated-object pose estimation in discrete state spaces, improving pose accuracy and robustness for articulated objects.
DICArt

ICME 2025
Estimates category-level articulation pose with a conditional diffusion model, capturing uncertainty and complex articulation distributions for 6D pose estimation.
Diff-Art

arXiv 2026
Generates physically plausible, text-driven indoor scenes in non-Manhattan rooms using statistical layout priors and hierarchical object placement.
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

arXiv 2026
A training-free Sim(3) alignment framework with coarse-to-fine matching and hallucination filtering for grounding generative 3D priors in partial monocular observations.
Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations

AAAI 2026
Formulates articulated-object tracking on SE(3) manifolds and uses SE(3)-invariant pose-voting parameters for robust category-level pose tracking.
Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds

Contact



This page is maintained by Xianhui Meng. Last update: 2026.08.29.