English

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Computer Vision and Pattern Recognition 2023-09-04 v1 Artificial Intelligence Computation and Language Machine Learning Multimedia

Abstract

We introduce Point-Bind, a 3D multi-modality model aligning point clouds with 2D image, language, audio, and video. Guided by ImageBind, we construct a joint embedding space between 3D and multi-modalities, enabling many promising applications, e.g., any-to-3D generation, 3D embedding arithmetic, and 3D open-world understanding. On top of this, we further present Point-LLM, the first 3D large language model (LLM) following 3D multi-modal instructions. By parameter-efficient fine-tuning techniques, Point-LLM injects the semantics of Point-Bind into pre-trained LLMs, e.g., LLaMA, which requires no 3D instruction data, but exhibits superior 3D and multi-modal question-answering capacity. We hope our work may cast a light on the community for extending 3D point clouds to multi-modality applications. Code is available at https://github.com/ZiyuGuo99/Point-Bind_Point-LLM.

Keywords

Cite

@article{arxiv.2309.00615,
  title  = {Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following},
  author = {Ziyu Guo and Renrui Zhang and Xiangyang Zhu and Yiwen Tang and Xianzheng Ma and Jiaming Han and Kexin Chen and Peng Gao and Xianzhi Li and Hongsheng Li and Pheng-Ann Heng},
  journal= {arXiv preprint arXiv:2309.00615},
  year   = {2023}
}

Comments

Work in progress. Code is available at https://github.com/ZiyuGuo99/Point-Bind_Point-LLM

R2 v1 2026-06-28T12:10:38.030Z