中文
相关论文

相关论文: Transformer-Driven Multimodal Fusion for Explainab…

200 篇论文

Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucination risks due to weak visual foundations. To address this,…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Weilin Zhou , Zonghao Ying , Chunlei Meng , Jiahui Liu , Hengyang Zhou , Quanchen Zou , Deyue Zhang , Dongdong Yang , Xiangzheng Zhang

Intelligent surveillance systems often handle perceptual tasks such as object detection, facial recognition, and emotion analysis independently, but they lack a unified, adaptive runtime scheduler that dynamically allocates computational…

Object detection in unmanned aerial vehicle (UAV) remote sensing images poses significant challenges due to unstable image quality, small object sizes, complex backgrounds, and environmental occlusions. Small objects, in particular, occupy…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Xudong Wang , Yaxin Peng , Chaomin Shen

Object recognition from live video streams comes with numerous challenges such as the variation in illumination conditions and poses. Convolutional neural networks (CNNs) have been widely used to perform intelligent visual object…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Muhammad Usman Yaseen , Ashiq Anjum , Giancarlo Fortino , Antonio Liotta , Amir Hussain

Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To facilitate the use of these improvements for road safety, this survey systematically…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Yaoqi Huang , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic scenes. While many approaches have crafted task-specific…

计算机视觉与模式识别 · 计算机科学 2023-09-18 Junwen Xiong , Peng Zhang , Chuanyue Li , Wei Huang , Yufei Zha , Tao You

Accurate recognition of sign language in healthcare communication poses a significant challenge, requiring frameworks that can accurately interpret complex multimodal gestures. To deal with this, we propose FusionEnsemble-Net, a novel…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Md. Milon Islam , Md Rezwanul Haque , S M Taslim Uddin Raju , Fakhri Karray

Robust object detection for Unmanned Surface Vehicles (USVs) in complex water environments is essential for reliable navigation and operation. Specifically, water surface object detection faces challenges from blurred edges and diverse…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Huilin Yin , Pengyu Wang , Senmao Li , Jun Yan , Daniel Watzenig

Due to their ability to offer more comprehensive information than data from a single view, multi-view (multi-source, multi-modal, multi-perspective, etc.) data are being used more frequently in remote sensing tasks. However, as the number…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Kun Zhao , Qian Gao , Siyuan Hao , Jie Sun , Lijian Zhou

A key technical challenge in performing 6D object pose estimation from RGB-D image is to fully leverage the two complementary data sources. Prior works either extract information from the RGB image and depth separately or use costly…

计算机视觉与模式识别 · 计算机科学 2019-01-16 Chen Wang , Danfei Xu , Yuke Zhu , Roberto Martín-Martín , Cewu Lu , Li Fei-Fei , Silvio Savarese

Deepfake is a generative deep learning algorithm that creates or changes facial features in a very realistic way making it hard to differentiate the real from the fake features It can be used to make movies look better as well as to spread…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Nadeem Jabbar CH , Aqib Saghir , Ayaz Ahmad Meer , Salman Ahmad Sahi , Bilal Hassan , Siddiqui Muhammad Yasir

With every advancement in generative AI models, forensics is under increasing pressure. The constant emergence of new generation techniques makes it impossible to collect data for each manipulation to train a deepfake detection model. Thus,…

人工智能 · 计算机科学 2026-05-20 Aritra Marik , Marcel Klemt , Anna Rohrbach

Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest challenges to this task is the one-to-many mapping from…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Jiangnan Tang , Jingya Wang , Kaiyang Ji , Lan Xu , Jingyi Yu , Ye Shi

In this study, we present a comprehensive public dataset for driver drowsiness detection, integrating multimodal signals of facial, behavioral, and biometric indicators. Our dataset includes 3D facial video using a depth camera, IR camera…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Morteza Bodaghi , Majid Hosseini , Raju Gottumukkala , Ravi Teja Bhupatiraju , Iftikhar Ahmad , Moncef Gabbouj

Most existing video anomaly detectors rely solely on RGB frames, which lack the temporal resolution needed to capture abrupt or transient motion cues, key indicators of anomalous events. To address this limitation, we propose Image-Event…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Sungheon Jeong , Jihong Park , Mohsen Imani

To ensure safe clinical integration, deep learning models must provide more than just high accuracy; they require dependable uncertainty quantification. While current Medical Vision Transformers perform well, they frequently struggle with…

图像与视频处理 · 电气工程与系统科学 2026-04-13 Mohammed Maaz Sibhai , Abedalrhman Alkhateeb , Saad B. Ahmed

The automated real-time recognition of unexpected situations plays a crucial role in the safety of autonomous vehicles, especially in unsupported and unpredictable scenarios. This paper evaluates different Bayesian uncertainty…

机器学习 · 计算机科学 2025-02-14 Ruben Grewal , Paolo Tonella , Andrea Stocco

We introduce ForeSight, a novel joint detection and forecasting framework for vision-based 3D perception in autonomous vehicles. Traditional approaches treat detection and forecasting as separate sequential tasks, limiting their ability to…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Sandro Papais , Letian Wang , Brian Cheong , Steven L. Waslander

We propose a novel convolutional neural network approach to address the fine-grained recognition problem of multi-view dynamic facial action unit detection. We leverage recent gains in large-scale object recognition by formulating the task…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Andres Romero , Juan Leon , Pablo Arbelaez

We present a novel unsupervised deep learning framework for anomalous event detection in complex video scenes. While most existing works merely use hand-crafted appearance and motion features, we propose Appearance and Motion DeepNet (AMDN)…

计算机视觉与模式识别 · 计算机科学 2015-10-07 Dan Xu , Elisa Ricci , Yan Yan , Jingkuan Song , Nicu Sebe