中文
相关论文

相关论文: GeminiFusion: Efficient Pixel-wise Multimodal Fusi…

200 篇论文

As a concrete application of multi-view learning, multi-view classification improves the traditional classification methods significantly by integrating various views optimally. Although most of the previous efforts have been demonstrated…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Jinglin Xu , Wenbin Li , Jiantao Shen , Xinwang Liu , Peicheng Zhou , Xiangsen Zhang , Xiwen Yao , Junwei Han

Detecting user interface (UI) controls from software screenshots is a critical task for automated testing, accessibility, and software analytics, yet it remains challenging due to visual ambiguities, design variability, and the lack of…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Milad Moradi , Ke Yan , David Colwell , Matthias Samwald , Rhona Asgari

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an innovative feature fusion framework based on cross-modal…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Jifeng Shen , Haibo Zhan , Xin Zuo , Heng Fan , Xiaohui Yuan , Jun Li , Wankou Yang

There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech translation (ST),…

计算与语言 · 计算机科学 2026-04-02 Sai Koneru , Matthias Huck , Jan Niehues

The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages--cameras provide rich texture information and LiDAR offers precise…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Zitian Wang , Zehao Huang , Yulu Gao , Naiyan Wang , Si Liu

The video topic segmentation (VTS) task segments videos into intelligible, non-overlapping topics, facilitating efficient comprehension of video content and quick access to specific content. VTS is also critical to various downstream video…

人工智能 · 计算机科学 2024-12-31 Hai Yu , Chong Deng , Qinglin Zhang , Jiaqing Liu , Qian Chen , Wen Wang

Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Chunyang Cheng , Tianyang Xu , Xiao-Jun Wu , Tao Zhou , Hui Li , Zhangyong Tang , Josef Kittler

The end-to-end image fusion framework has achieved promising performance, with dedicated convolutional networks aggregating the multi-modal local appearance. However, long-range dependencies are directly neglected in existing CNN fusion…

计算机视觉与模式识别 · 计算机科学 2022-02-07 Dongyu Rao , Xiao-Jun Wu , Tianyang Xu

Multi-focus image fusion aims to generate an all-in-focus image from a sequence of partially focused input images. Existing fusion algorithms generally assume that, for every spatial location in the scene, there is at least one input image…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Xinzhe Xie , Buyu Guo , Bolin Li , Shuangyan He , Yanzhen Gu , Qingyan Jiang , Peiliang Li

Multimodal sentiment analysis aims to extract and integrate semantic information collected from multiple modalities to recognize the expressed emotions and sentiment in multimodal data. This research area's major concern lies in developing…

人工智能 · 计算机科学 2021-08-31 Wei Han , Hui Chen , Alexander Gelbukh , Amir Zadeh , Louis-philippe Morency , Soujanya Poria

Multi-modal approaches employ data from multiple input streams such as textual and visual domains. Deep neural networks have been successfully employed for these approaches. In this paper, we present a novel multi-modal approach that fuses…

计算机视觉与模式识别 · 计算机科学 2018-10-05 Ignazio Gallo , Alessandro Calefati , Shah Nawaz , Muhammad Kamran Janjua

Deep multimodal fusion by using multiple sources of data for classification or regression has exhibited a clear advantage over the unimodal counterpart on various applications. Yet, current methods including aggregation-based and…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Yikai Wang , Wenbing Huang , Fuchun Sun , Tingyang Xu , Yu Rong , Junzhou Huang

Multimodal human action understanding is a significant problem in computer vision, with the central challenge being the effective utilization of the complementarity among diverse modalities while maintaining model efficiency. However, most…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hongsong Wang , Heng Fei , Bingxuan Dai , Jie Gui

Most existing speech disfluency detection techniques only rely upon acoustic data. In this work, we present a practical multimodal disfluency detection approach that leverages available video data together with audio. We curate an…

计算与语言 · 计算机科学 2024-06-12 Payal Mohapatra , Shamika Likhite , Subrata Biswas , Bashima Islam , Qi Zhu

Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Yiming Sun , Bing Cao , Pengfei Zhu , Qinghua Hu

Improving hyperspectral image (HSI) semantic segmentation by exploiting complementary information from a supplementary data type (referred to X-modality) is promising but challenging due to differences in imaging sensors, image content, and…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Xuming Zhang , Xingfa Gu , Qingjiu Tian , Lorenzo Bruzzone

The introduction of the neural implicit representation has notably propelled the advancement of online dense reconstruction techniques. Compared to traditional explicit representations, such as TSDF, it improves the mapping completeness and…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Yuqing Lan , Chenyang Zhu , Shuaifeng Zhi , Jiazhao Zhang , Zhoufeng Wang , Renjiao Yi , Yijie Wang , Kai Xu

We introduce TransformerFusion, a transformer-based 3D scene reconstruction approach. From an input monocular RGB video, the video frames are processed by a transformer network that fuses the observations into a volumetric feature grid…

计算机视觉与模式识别 · 计算机科学 2021-07-07 Aljaž Božič , Pablo Palafox , Justus Thies , Angela Dai , Matthias Nießner

Multimodal molecular models often suffer from 3D conformer unreliability and modality collapse, limiting their robustness and generalization. We propose MuMo, a structured multimodal fusion framework that addresses these challenges in…

机器学习 · 计算机科学 2025-10-29 Zihao Jing , Yan Sun , Yan Yi Li , Sugitha Janarthanan , Alana Deng , Pingzhao Hu

Existing multimodal tracking studies focus on bi-modal scenarios such as RGB-Thermal, RGB-Event, and RGB-Language. Although promising tracking performance is achieved through leveraging complementary cues from different sources, it remains…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Andong Lu , Mai Wen , Jinhu Wang , Yuanzhi Guo , Chenglong Li , Jin Tang , Bin Luo
‹ 上一页 1 8 9 10 下一页 ›