English
Related papers

Related papers: Collaborative Multi-Modal Coding for High-Quality …

200 papers

Score Distillation Sampling (SDS) leverages pretrained 2D diffusion models to advance text-to-3D generation but neglects multi-view correlations, being prone to geometric inconsistencies and multi-face artifacts in the generated 3D content.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Feng Yang , Wenliang Qian , Wangmeng Zuo , Hui Li

We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learned cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Muhammad Abdullah Jamal , Omid Mohareri

In perception, multiple sensory information is integrated to map visual information from 2D views onto 3D objects, which is beneficial for understanding in 3D environments. But in terms of a single 2D view rendered from different angles,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Hai-Tao Yu , Mofei Song

Recent research has demonstrated that Large Language Models (LLMs) are not limited to text-only tasks but can also function as multimodal models across various modalities, including audio, images, and videos. In particular, research on 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Keonwoo Kim , Yeongjae Cho , Taebaek Hwang , Minsoo Jo , Sangdo Han

We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, our model…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Chunshi Wang , Junliang Ye , Yunhan Yang , Yang Li , Zizhuo Lin , Jun Zhu , Zhuo Chen , Yawei Luo , Chunchao Guo

Text-to-motion generation, a rapidly evolving field in computer vision, aims to produce realistic and text-aligned motion sequences. Current methods primarily focus on spatial-temporal modeling or independent frequency domain analysis,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Yiyang Cao , Yunze Deng , Ziyu Lin , Bin Feng , Xinggang Wang , Wenyu Liu , Dandan Zheng , Jingdong Chen

Adversarial attacks pose a significant threat to learning-based 3D point cloud models, critically undermining their reliability in security-sensitive applications. Existing defense methods often suffer from (1) high computational overhead…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Xiang Gu , Liming Lu , Xu Zheng , Anan Du , Yongbin Zhou , Shuchao Pang

Generative models for structure-based molecular design hold significant promise for drug discovery, with the potential to speed up the hit-to-lead development cycle, while improving the quality of drug candidates and reducing costs. Data…

Machine Learning · Statistics 2022-04-25 Lucian Chan , Rajendra Kumar , Marcel Verdonk , Carl Poelking

We present LTM3D, a Latent Token space Modeling framework for conditional 3D shape generation that integrates the strengths of diffusion and auto-regressive (AR) models. While diffusion-based methods effectively model continuous latent…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Xin Kang , Zihan Zheng , Lei Chu , Yue Gao , Jiahao Li , Hao Pan , Xuejin Chen , Yan Lu

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

Latent diffusion models (LDMs) dominate high-quality image generation, yet integrating representation learning with generative modeling remains a challenge. We introduce a novel generative image modeling framework that seamlessly bridges…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Theodoros Kouzelis , Efstathios Karypidis , Ioannis Kakogeorgiou , Spyros Gidaris , Nikos Komodakis

In modern e-commerce, item content features in various modalities offer accurate yet comprehensive information to recommender systems. The majority of previous work either focuses on learning effective item representation during modelling…

Information Retrieval · Computer Science 2024-08-15 Hao Wu , Alejandro Ariza-Casabona , Bartłomiej Twardowski , Tri Kurniawan Wijaya

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

Recent 3D human generative models have achieved remarkable progress by learning 3D-aware GANs from 2D images. However, existing 3D human generative methods model humans in a compact 1D latent space, ignoring the articulated structure and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Tao Hu , Fangzhou Hong , Ziwei Liu

The rising importance of 3D understanding, pivotal in computer vision, autonomous driving, and robotics, is evident. However, a prevailing trend, which straightforwardly resorted to transferring 2D alignment strategies to the 3D domain,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Jiayi Ji , Haowei Wang , Changli Wu , Yiwei Ma , Xiaoshuai Sun , Rongrong Ji

Semantic-driven 3D shape generation aims to generate 3D objects conditioned on text. Previous works face problems with single-category generation, low-frequency 3D details, and requiring a large number of paired datasets for training. To…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Bo Han , Yitong Fu , Yixuan Shen

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI…

Computer Vision and Pattern Recognition · Computer Science 2023-05-22 Zineng Tang , Ziyi Yang , Chenguang Zhu , Michael Zeng , Mohit Bansal

Following rapid advancements in text and image generation, research has increasingly shifted towards 3D generation. Unlike the well-established pixel-based representation in images, 3D representations remain diverse and fragmented,…

Recently, multi-view diffusion-based 3D generation methods have gained significant attention. However, these methods often suffer from shape and texture misalignment across generated multi-view images, leading to low-quality 3D generation…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zhuojiang Cai , Yiheng Zhang , Meitong Guo , Mingdao Wang , Yuwang Wang

Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xue-Feng Zhu , Tianyang Xu , Yifan Pan , Jinjie Gu , Xi Li , Jiwen Lu , Xiao-Jun Wu , Josef Kittler