Skip to content

Navigation Menu

Sign in
Sign up

H-EmbodVis

Embodied Vision Projects from Huazhong University of Science and Technology

H-EmbodVis

Embodied Vision · World Models · Autonomous Driving · 3D Scene Understanding

H-EmbodVis (Huazhong University of Science and Technology Embodied Vision Projects) is a research initiative. We primarily focus on Embodied AI, while also exploring Autonomous Driving and Generative Models.


🔬 Research Areas

We focus on building intelligent systems that can perceive, understand, and interact with the physical world. Key directions include:

  • Embodied AI & Agents: Integrating vision, language, and action planning.
  • World Models for Autonomous Driving: Developing end-to-end driving frameworks and simulators.
  • 3D Vision & Point Cloud Analysis: Efficient architectures for 3D representation learning.
  • Multimodal Foundation Models: Large-scale models for diverse data modalities.

🌟 Featured Projects

Autonomous Driving & World Models

  • HERMES (ICCV 2025) A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.
  • Orion (ICCV 2025) Holistic End-to-End Autonomous Driving via Vision-Language Instructed Action Generation.
  • Awesome-World-Model Curated collection of papers on World Models for Autonomous Driving and Robotics.

3D Vision & Efficient Computing

  • PointMamba (NeurIPS 2024) State Space Models (Mamba) applied to Point Cloud Analysis.
  • UniSeg3D (NeurIPS 2024) A Unified Framework for 3D Scene Understanding.
  • PointGST (IEEE TPAMI) Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning.
  • EasyCache Training-Free Video Diffusion Acceleration.

Multimodal & Embodied Agents

  • NAUTILUS (NeurIPS 2025) A Large Multimodal Model for Underwater Scene Understanding.
  • GRANT (AAAI 2026 Oral) Teaching Embodied Agents for Parallel Task Execution.
  • MERGE (NeurIPS 2025) Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models.

Collaboration

We are always looking for passionate collaborators and students.

  • Connect: Reach out via email (dkliang@hust.edu.cn).
  • Reuse: Creating impactful open-source software is a core value. Please cite our papers if you use our code.

🌐 Website | 🎓 Google Scholar | 📂 Repositories

Pinned Loading

  1. VEGA-3D VEGA-3D Public

    [ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    Python 420 23

  2. TurboVLA TurboVLA Public

    TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    Python 479 61

  3. DOMINO DOMINO Public

    [ECCV 2026] Towards Generalizable Robotic Manipulation in Dynamic Environments

    Python 228 10

  4. HyDRA HyDRA Public

    Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

    Python 276 14

  5. Orion Orion Public

    Forked from xiaomi-mlab/Orion

    [ICCV 2025] Official code of "ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation"

    Python

  6. PointGST PointGST Public

    Forked from jerryfeng2003/PointGST

    [IEEE TPAMI] Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning

    Python

Repositories

Loading
Type
Select type
Language
Select language
Sort
Select order
Showing 10 of 24 repositories

Top languages

Loading...

Most used topics

Loading...

AltStyle によって変換されたページ (->オリジナル) /