中文
相关论文

相关论文: Deep Multimodal Fusion for Surgical Feedback Class…

200 篇论文

In order to perform multimodal fusion of heterogeneous signals, we need to understand their interactions: how each modality individually provides information useful for a task and how this information changes in the presence of other…

机器学习 · 计算机科学 2023-11-01 Paul Pu Liang , Yun Cheng , Ruslan Salakhutdinov , Louis-Philippe Morency

Auditory and visual signals usually present together and correlate with each other, not only in natural environments but also in clinical settings. However, the audio-visual modelling in the latter case can be more challenging, due to the…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Jianbo Jiao , Mohammad Alsharid , Lior Drukker , Aris T. Papageorghiou , Andrew Zisserman , J. Alison Noble

While biological vision systems rely heavily on feedback connections to iteratively refine perception, most artificial neural networks remain purely feedforward, processing input in a single static pass. In this work, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-07-15 David Calhas , Arlindo L. Oliveira

In this paper, we present methods in deep multimodal learning for fusing speech and visual modalities for Audio-Visual Automatic Speech Recognition (AV-ASR). First, we study an approach where uni-modal deep networks are trained separately…

计算与语言 · 计算机科学 2015-01-23 Youssef Mroueh , Etienne Marcheret , Vaibhava Goel

Learning from multiple modalities, such as audio and video, offers opportunities for leveraging complementary information, enhancing robustness, and improving contextual understanding and performance. However, combining such modalities…

多媒体 · 计算机科学 2024-10-15 Konstantinos Kontras , Christos Chatzichristos , Matthew Blaschko , Maarten De Vos

In the last decade, video blogs (vlogs) have become an extremely popular method through which people express sentiment. The ubiquitousness of these videos has increased the importance of multimodal fusion models, which incorporate video and…

计算机视觉与模式识别 · 计算机科学 2018-07-04 Nathaniel Blanchard , Daniel Moreira , Aparna Bharati , Walter J. Scheirer

The current cancer treatment practice collects multimodal data, such as radiology images, histopathology slides, genomics and clinical data. The importance of these data sources taken individually has fostered the recent raise of radiomics…

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

图像与视频处理 · 电气工程与系统科学 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Haytham M. Fayek , Anurag Kumar

Laparoscopic surgery constrains surgeons spatial awareness because procedures are performed through a monocular, two-dimensional (2D) endoscopic view. Conventional training methods using dry-lab models or recorded videos provide limited…

人机交互 · 计算机科学 2025-11-05 Songyang Liu , Yunpeng Tan , Shuai Li

Medical conversations between patients and medical professionals have implicit functional sections, such as "history taking", "summarization", "education", and "care plan." In this work, we are interested in learning to automatically…

计算与语言 · 计算机科学 2022-10-10 Mengqian Wang , Ilya Valmianski , Xavier Amatriain , Anitha Kannan

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

声音 · 计算机科学 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li

Creating a meaningful representation by fusing single modalities (e.g., text, images, or audio) is the core concept of multimodal learning. Although several techniques for building multimodal representations have been proven successful,…

机器学习 · 计算机科学 2025-08-08 Maciej Pawłowski , Anna Wróblewska , Sylwia Sysko-Romańczuk

Summaries generated from medical conversations can improve recall and understanding of care plans for patients and reduce documentation burden for doctors. Recent advancements in automatic speech recognition (ASR) and natural language…

计算与语言 · 计算机科学 2020-07-28 Benjamin Schloss , Sandeep Konam

Despite renewed awareness of the importance of articulation, it remains a challenge for instructors to handle the pronunciation needs of language learners. There are relatively scarce pedagogical tools for pronunciation teaching and…

计算机视觉与模式识别 · 计算机科学 2020-05-15 M. Hamed Mozaffari , Won-Sook Lee

This paper presents a system for detecting fake audio-visual content (i.e., video deepfake), developed for Track 2 of the DDL Challenge. The proposed system employs a two-stage framework, comprising unimodal detection and multimodal score…

多媒体 · 计算机科学 2026-02-03 Qingcao Li , Miao He , Liang Yi , Qing Wen , Yitao Zhang , Hongshuo Jin , Peng Cheng , Zhongjie Ba , Li Lu , Kui Ren

Background: AI-driven prediction algorithms have the potential to enhance emergency medicine by enabling rapid and accurate decision-making regarding patient status and potential deterioration. However, the integration of multimodal data,…

机器学习 · 计算机科学 2025-05-02 Juan Miguel Lopez Alcaraz , Hjalmar Bouma , Nils Strodthoff

Evaluating AI generated dubbed content is inherently multi-dimensional, shaped by synchronization, intelligibility, speaker consistency, emotional alignment, and semantic context. Human Mean Opinion Scores (MOS) remain the gold standard but…

音频与语音处理 · 电气工程与系统科学 2026-04-27 Ashwini Dasare , Nirmesh Shah , Ashishkumar Gudmalwar , Pankaj Wasnik

Multimodal pathological images are usually in clinical diagnosis, but computer vision-based multimodal image-assisted diagnosis faces challenges with modality fusion, especially in the absence of expert-annotated data. To achieve the…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Qinghua Lin , Guang-Hai Liu , Zuoyong Li , Yang Li , Yuting Jiang , Xiang Wu

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

计算机视觉与模式识别 · 计算机科学 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang