English
Related papers

Related papers: Collaborative Multi-Modal Coding for High-Quality …

200 papers

Traditional and neural video codecs commonly encounter limitations in controllability and generality under ultra-low-bitrate coding scenarios. To overcome these challenges, we propose M3-CVC, a controllable video compression framework…

Image and Video Processing · Electrical Eng. & Systems 2024-12-30 Rui Wan , Qi Zheng , Yibo Fan

It is counter-intuitive that multi-modality methods based on point cloud and images perform only marginally better or sometimes worse than approaches that solely use point cloud. This paper investigates the reason behind this phenomenon.…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Wenwei Zhang , Zhe Wang , Chen Change Loy

Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environments. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Zongyan Han , Mohamed El Amine Boudjoghra , Jiahua Dong , Jinhong Wang , Rao Muhammad Anwer

Recent advances in 3D AIGC have shown promise in directly creating 3D objects from text and images, offering significant cost savings in animation and product design. However, detailed edit and customization of 3D assets remains a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Zhangyang Qi , Yunhan Yang , Mengchen Zhang , Long Xing , Xiaoyang Wu , Tong Wu , Dahua Lin , Xihui Liu , Jiaqi Wang , Hengshuang Zhao

Multi-modal vehicle Re-Identification (ReID) aims to leverage complementary information from RGB, Near Infrared (NIR), and Thermal Infrared (TIR) modalities to retrieve the same vehicle. The challenges of multi-modal vehicle ReID arise from…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Aihua Zheng , Ya Gao , Shihao Li , Chenglong Li , Jin Tang

Casting neural networks in generative frameworks is a highly sought-after endeavor these days. Contemporary methods, such as Generative Adversarial Networks, capture some of the generative capabilities, but not all. In particular, they lack…

Machine Learning · Computer Science 2018-03-28 Or Sharir , Ronen Tamari , Nadav Cohen , Amnon Shashua

We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Kotaro Kikuchi , Ukyo Honda , Naoto Inoue , Mayu Otani , Edgar Simo-Serra , Kota Yamaguchi

AI-driven design problems, such as DNA/protein sequence design, are commonly tackled from two angles: generative modeling, which efficiently captures the feasible design space (e.g., natural images or biological sequences), and model-based…

Developing robust multi-modal feature representations is crucial for enhancing object tracking performance. In pursuit of this objective, a novel X Modality Assisting Network (X-Net) is introduced, which explores the impact of the fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Zhaisheng Ding , Haiyan Li , Ruichao Hou , Yanyu Liu , Shidong Xie

Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level embedding vector. In this work we present a self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2021-06-11 Alexander H. Liu , SouYoung Jin , Cheng-I Jeff Lai , Andrew Rouditchenko , Aude Oliva , James Glass

Autonomous driving sensors generate an enormous amount of data. In this paper, we explore learned multimodal compression for autonomous driving, specifically targeted at 3D object detection. We focus on camera and LiDAR modalities and…

Image and Video Processing · Electrical Eng. & Systems 2024-08-16 Hadi Hadizadeh , Ivan V. Bajić

RGB cloth generation has been deeply studied in the related literature, however, 3D garment generation remains an open problem. In this paper, we build a conditional variational autoencoder for 3D garment generation and draping. We propose…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Hunor Laczkó , Meysam Madadi , Sergio Escalera , Jordi Gonzalez

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

Generation of graphs is a major challenge for real-world tasks that require understanding the complex nature of their non-Euclidean structures. Although diffusion models have achieved notable success in graph generation recently, they are…

Machine Learning · Computer Science 2024-06-04 Jaehyeong Jo , Dongki Kim , Sung Ju Hwang

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Bingyu Li , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Modeling the 3D world from sensor data for simulation is a scalable way of developing testing and validation environments for robotic learning problems such as autonomous driving. However, manually creating or re-creating real-world-like…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Bokui Shen , Xinchen Yan , Charles R. Qi , Mahyar Najibi , Boyang Deng , Leonidas Guibas , Yin Zhou , Dragomir Anguelov

Generative 3D modeling has advanced rapidly, driven by applications in VR/AR, metaverse, and robotics. However, most methods represent the target object as a closed mesh devoid of any structural information, limiting editing, animation, and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Omid Bonakdar , Nasser Mozayani

Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Junyoung Seo , Susung Hong , Wooseok Jang , Inès Hyeonsu Kim , Minseop Kwak , Doyup Lee , Seungryong Kim

In this paper, we study an under-explored but important factor of diffusion generative models, i.e., the combinatorial complexity. Data samples are generally high-dimensional, and for various structured generation tasks, additional…

Machine Learning · Computer Science 2026-04-30 Rui Xu , Jiepeng Wang , Hao Pan , Yang Liu , Xin Tong , Shiqing Xin , Changhe Tu , Taku Komura , Wenping Wang

The goal of multi-modal learning is to use complimentary information on the relevant task provided by the multiple modalities to achieve reliable and robust performance. Recently, deep learning has led significant improvement in multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Jaekyum Kim , Junho Koh , Yecheol Kim , Jaehyung Choi , Youngbae Hwang , Jun Won Choi