English
Related papers

Related papers: Delivering Arbitrary-Modal Semantic Segmentation

200 papers

Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete…

Machine Learning · Computer Science 2026-05-13 Camillo Maria Caruso , Valerio Guarrasi , Paolo Soda

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Pengcheng Zhao , Jinxing Zhou , Yang Zhao , Dan Guo , Yanxiang Chen

This work introduces RGBX-DiffusionDet, an object detection framework extending the DiffusionDet model to fuse the heterogeneous 2D data (X) with RGB imagery via an adaptive multimodal encoder. To enable cross-modal interaction, we design…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Eliraz Orfaig , Inna Stainvas , Igal Bilik

The large-scale pretrained model CLIP, trained on 400 million image-text pairs, offers a promising paradigm for tackling vision tasks, albeit at the image level. Later works, such as DenseCLIP and LSeg, extend this paradigm to dense…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Ke Jin , Wankou Yang

Segmentation of drivable roads and negative obstacles is critical to the safe driving of autonomous vehicles. Currently, many multi-modal fusion methods have been proposed to improve segmentation accuracy, such as fusing RGB and depth…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Zhen Feng , Yuchao Feng , Yanning Guo , Yuxiang Sun

Segmentation models based on deep neural networks demonstrate strong generalization for medical image segmentation. However, they often exhibit overconfidence or underconfidence, leading to unreliable confidence scores for segmentation…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Qiuyu Tian , Haoliang Sun , Yunshan Wang , Yinghuan Shi , Yilong Yin

Sensor fusion is a fundamental process in robotic systems as it extends the perceptual range and increases robustness in real-world operations. Current multi-sensor deep learning based semantic segmentation approaches do not provide…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Hermann Blum , Abel Gawel , Roland Siegwart , Cesar Cadena

LiDAR semantic segmentation is crucial for autonomous vehicles and mobile robots, requiring high accuracy and real-time processing, especially on resource-constrained embedded systems. Previous state-of-the-art methods often face a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Samir Abou Haidar , Alexandre Chariot , Mehdi Darouich , Cyril Joly , Jean-Emmanuel Deschaud

Multimodal deep sensor fusion has the potential to enable autonomous vehicles to visually understand their surrounding environments in all weather conditions. However, existing deep sensor fusion methods usually employ convoluted…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Sri Aditya Deevi , Connor Lee , Lu Gan , Sushruth Nagesh , Gaurav Pandey , Soon-Jo Chung

Evolving data streams induce joint nonstationarity in continual semantic segmentation, where semantic classes, input distributions, and supervision availability change simultaneously over time. This setting reflects practical structured…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Prashant Pandey , Himanshu Kumar , Devineni Sri Venkatraya Chowdary , Brejesh Lall

Website fingerprinting (WF) attacks infer the websites visited by users from encrypted traffic in anonymous networks such as Tor. Existing deep learning methods achieve high accuracy under the single-tab assumption but degrade substantially…

Cryptography and Security · Computer Science 2026-04-20 Yali Yuan , Yaosheng Liu , Qianqi Niu , Guang Cheng

Multimodal Emotion Recognition in Conversations (MERC) identifies emotional states across text, audio and video, which is essential for intelligent dialogue systems and opinion analysis. Existing methods emphasize heterogeneous modal fusion…

Machine Learning · Computer Science 2025-04-01 Jiagen Li , Rui Yu , Huihao Huang , Huaicheng Yan

Despite significant progress in 3D object detection, point clouds remain challenging due to sparse data, incomplete structures, and limited semantic information. Capturing contextual relationships between distant objects presents additional…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Md Sohag Mia , Md Nahid Hasan , Muhammad Abdullah Adnan

Multispectral pedestrian detection is capable of adapting to insufficient illumination conditions by leveraging color-thermal modalities. On the other hand, it is still lacking of in-depth insights on how to fuse the two modalities…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Kailai Zhou , Linsen Chen , Xun Cao

Multimodal learning leverages complementary information derived from different modalities, thereby enhancing performance in medical image segmentation. However, prevailing multimodal learning methods heavily rely on extensive well-annotated…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Xiaogen Zhou , Yiyou Sun , Min Deng , Winnie Chiu Wing Chu , Qi Dou

While numerous 3D detection works leverage the complementary relationship between RGB images and point clouds, developments in the broader framework of semi-supervised object recognition remain uninfluenced by multi-modal fusion. Current…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Jinhyung Park , Chenfeng Xu , Yiyang Zhou , Masayoshi Tomizuka , Wei Zhan

Classification using multimodal data arises in many machine learning applications. It is crucial not only to model cross-modal relationship effectively but also to ensure robustness against loss of part of data or modalities. In this paper,…

Machine Learning · Computer Science 2019-04-22 Jun-Ho Choi , Jong-Seok Lee

In the data era, the integration of multiple data types, known as multimodality, has become a key area of interest in the research community. This interest is driven by the goal to develop cutting edge multimodal models capable of serving…

Machine Learning · Computer Science 2025-03-12 Wei Zhang , Xinyue Wang , Lan Yu , Shi Li

The success of multimodal data fusion in deep learning appears to be attributed to the use of complementary in-formation between multiple input data. Compared to their predictive performance, relatively less attention has been devoted to…

Computer Vision and Pattern Recognition · Computer Science 2020-05-25 Youngjoon Yu , Hong Joo Lee , Byeong Cheon Kim , Jung Uk Kim , Yong Man Ro

Most existing federated learning (FL) methods for medical image analysis only considered intramodal heterogeneity, limiting their applicability to multimodal imaging applications. In practice, some FL participants may possess only a subset…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Hong Liu , Dong Wei , Qian Dai , Xian Wu , Yefeng Zheng , Liansheng Wang