English
Related papers

Related papers: MMPart: Harnessing Multi-Modal Large Language Mode…

200 papers

The creation of 3D assets with explicit, editable part structures is crucial for advancing interactive applications, yet most generative methods produce only monolithic shapes, limiting their utility. We introduce OmniPart, a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yunhan Yang , Yufan Zhou , Yuan-Chen Guo , Zi-Xin Zou , Yukun Huang , Ying-Tian Liu , Hao Xu , Ding Liang , Yan-Pei Cao , Xihui Liu

We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, our model…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Chunshi Wang , Junliang Ye , Yunhan Yang , Yang Li , Zizhuo Lin , Jun Zhu , Zhuo Chen , Yawei Luo , Chunchao Guo

Recent advances in 3D generation have transitioned from multi-view 2D rendering approaches to 3D-native latent diffusion frameworks that exploit geometric priors in ground truth data. Despite progress, three key limitations persist: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Shaocong Dong , Lihe Ding , Xiao Chen , Yaokun Li , Yuxin Wang , Yucheng Wang , Qi Wang , Jaehyeok Kim , Chenjian Gao , Zhanpeng Huang , Zibin Wang , Tianfan Xue , Dan Xu

Generating 3D shapes at part level is pivotal for downstream applications such as mesh retopology, UV mapping, and 3D printing. However, existing part-based generation methods often lack sufficient controllability and suffer from poor…

Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Noor Ahmed , Cameron Braunstein , Steffen Eger , Eddy Ilg

Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted behaviors, yet most generative 3D models produce either monolithic meshes or arbitrary part…

Impressive progress in generative models and implicit representations gave rise to methods that can generate 3D shapes of high quality. However, being able to locally control and edit shapes is another essential property that can unlock…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Konstantinos Tertikas , Despoina Paschalidou , Boxiao Pan , Jeong Joon Park , Mikaela Angelina Uy , Ioannis Emiris , Yannis Avrithis , Leonidas Guibas

Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures. These assets typically consist of a single, fused representation, like an implicit neural field, a Gaussian mixture, or a mesh,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minghao Chen , Roman Shapovalov , Iro Laina , Tom Monnier , Jianyuan Wang , David Novotny , Andrea Vedaldi

Part-level 3D generation is essential for applications requiring decomposable and structured 3D synthesis. However, existing methods either rely on implicit part segmentation with limited granularity control or depend on strong external…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Xufan He , Yushuang Wu , Xiaoyang Guo , Chongjie Ye , Jiaqing Zhou , Tianlei Hu , Xiaoguang Han , Dong Du

3D part segmentation is a crucial and challenging task in 3D perception, playing a vital role in applications such as robotics, 3D generation, and 3D editing. Recent methods harness the powerful Vision Language Models (VLMs) for 2D-to-3D…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yunhan Yang , Yukun Huang , Yuan-Chen Guo , Liangjun Lu , Xiaoyang Wu , Edmund Y. Lam , Yan-Pei Cao , Xihui Liu

Recently, 3D generation methods have shown their powerful ability to automate 3D model creation. However, most 3D generation methods only rely on an input image or a text prompt to generate a 3D model, which lacks the control of each…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Peng Li , Suizhi Ma , Jialiang Chen , Yuan Liu , Congyi Zhang , Wei Xue , Wenhan Luo , Alla Sheffer , Wenping Wang , Yike Guo

Advances in 3D generative AI have enabled the creation of physical objects from text prompts, but challenges remain in creating objects involving multiple component types. We present a pipeline that integrates 3D generative AI with…

We introduce PartCrafter, the first structured 3D generative model that jointly synthesizes multiple semantically meaningful and geometrically distinct 3D meshes from a single RGB image. Unlike existing methods that either produce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Yuchen Lin , Chenguo Lin , Panwang Pan , Honglei Yan , Yiqiang Feng , Yadong Mu , Katerina Fragkiadaki

Generating stylized 3D human motion from speech signals presents substantial challenges, primarily due to the intricate and fine-grained relationships among speech signals, individual styles, and the corresponding body movements. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Lianlian Liu , YongKang He , Zhaojie Chu , Xiaofen Xing , Xiangmin Xu

Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of human appearance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Zehuan Huang , Hongxing Fan , Lipeng Wang , Lu Sheng

Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most text-to-3D methods overlook the semantic and functional structure of parts. While recent part-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Tianjiao Yu , Xinzhuo Li , Muntasir Wahed , Jerry Xiong , Yifan Shen , Ying Shen , Ismini Lourentzou

Generating articulated objects, such as laptops and microwaves, is a crucial yet challenging task with extensive applications in Embodied AI and AR/VR. Current image-to-3D methods primarily focus on surface geometry and texture, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Ruijie Lu , Yu Liu , Jiaxiang Tang , Junfeng Ni , Yuxiang Wang , Diwen Wan , Gang Zeng , Yixin Chen , Siyuan Huang

3D part amodal segmentation--decomposing a 3D shape into complete, semantically meaningful parts, even when occluded--is a challenging but crucial task for 3D content creation and understanding. Existing 3D part segmentation methods only…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Yunhan Yang , Yuan-Chen Guo , Yukun Huang , Zi-Xin Zou , Zhipeng Yu , Yangguang Li , Yan-Pei Cao , Xihui Liu

Methods that use neural networks for synthesizing 3D shapes in the form of a part-based representation have been introduced over the last few years. These methods represent shapes as a graph or hierarchy of parts and enable a variety of…

Graphics · Computer Science 2024-09-20 Yanran Guan , Oliver van Kaick

Recently, the emergence of diffusion models has opened up new opportunities for single-view reconstruction. However, all the existing methods represent the target object as a closed mesh devoid of any structural information, thus neglecting…

Graphics · Computer Science 2024-05-28 Anran Liu , Cheng Lin , Yuan Liu , Xiaoxiao Long , Zhiyang Dou , Hao-Xiang Guo , Ping Luo , Wenping Wang
‹ Prev 1 2 3 10 Next ›