English
Related papers

Related papers: Integrating Audio Narrations to Strengthen Domain …

200 papers

First person action recognition is an increasingly researched topic because of the growing popularity of wearable cameras. This is bringing to light cross-domain issues that are yet to be addressed in this context. Indeed, the information…

Computer Vision and Pattern Recognition · Computer Science 2021-06-04 Mirco Planamente , Chiara Plizzari , Emanuele Alberti , Barbara Caputo

This paper strives for activity recognition under domain shift, for example caused by change of scenery or camera viewpoint. The leading approaches reduce the shift in activity appearance by adversarial training and self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Yunhua Zhang , Hazel Doughty , Ling Shao , Cees G. M. Snoek

First person action recognition is becoming an increasingly researched area thanks to the rising popularity of wearable cameras. This is bringing to light cross-domain issues that are yet to be addressed in this context. Indeed, the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Mirco Planamente , Chiara Plizzari , Emanuele Alberti , Barbara Caputo

Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for integrating such multimodal information often stumble, leading…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jun Yu , Zerui Zhang , Zhihong Wei , Gongpeng Zhao , Zhongpeng Cai , Yongqi Wang , Guochen Xie , Jichao Zhu , Wangyuan Zhu

Adapting speaker recognition systems to new environments is a widely-used technique to improve a well-performing model learned from large-scale data towards a task-specific small-scale data scenarios. However, previous studies focus on…

Sound · Computer Science 2022-11-21 Zhenyu Wang , John H. L. Hansen

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

Sound · Computer Science 2022-04-29 Dan Oneata , Horia Cucu

Automatic video activity recognition is crucial across numerous domains like surveillance, healthcare, and robotics. However, recognizing human activities from video data becomes challenging when training and test data stem from diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Partho Ghosh , Raisa Bentay Hossain , Mohammad Zunaed , Taufiq Hasan

Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine…

For models to generalize under unseen domains (a.k.a domain generalization), it is crucial to learn feature representations that are domain-agnostic and capture the underlying semantics that makes up an object category. Recent advances…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Puneet Mangla , Shivam Chandhok , Milan Aggarwal , Vineeth N Balasubramanian , Balaji Krishnamurthy

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker…

Machine Learning · Computer Science 2025-06-02 Sean Foley , Hong Nguyen , Jihwan Lee , Sudarsana Reddy Kadiri , Dani Byrd , Louis Goldstein , Shrikanth Narayanan

In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities of foundation models, towards more comprehensive…

In real-world scenarios, achieving domain adaptation and generalization poses significant challenges, as models must adapt to or generalize across unknown target distributions. Extending these capabilities to unseen multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Hao Dong , Moru Liu , Kaiyang Zhou , Eleni Chatzi , Juho Kannala , Cyrill Stachniss , Olga Fink

Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, such as remote sensing and medical imaging, remains limited.…

Computation and Language · Computer Science 2026-02-05 Qinglong Cao , Yuntian Chen , Chao Ma , Xiaokang Yang

Speech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute. Most studies focus on tackling distant speech recognition as a separate problem,…

Computation and Language · Computer Science 2018-06-14 Hao Tang , Wei-Ning Hsu , Francois Grondin , James Glass

Person re-identification plays a significant role in realistic scenarios due to its various applications in public security and video surveillance. Recently, leveraging the supervised or semi-unsupervised learning paradigms, which benefits…

Computer Vision and Pattern Recognition · Computer Science 2023-01-02 Suncheng Xiang , Hao Chen , Wei Ran , Zefang Yu , Ting Liu , Dahong Qian , Yuzhuo Fu

Multimodal machine learning has gained significant attention in recent years due to its potential for integrating information from multiple modalities to enhance learning and decision-making processes. However, it is commonly observed that…

Machine Learning · Computer Science 2025-09-12 Sahiti Yerramilli , Jayant Sravan Tamarapalli , Jonathan Francis , Eric Nyberg

In line with the human capacity to perceive the world by simultaneously processing and integrating high-dimensional inputs from multiple modalities like vision and audio, we propose a novel model, MAiVAR-T (Multimodal Audio-Image to Video…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Muhammad Bilal Shaikh , Douglas Chai , Syed Mohammed Shamsul Islam , Naveed Akhtar

Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual signals have been shown to be useful for recovering…

Computation and Language · Computer Science 2020-10-07 Tejas Srinivasan , Ramon Sanabria , Florian Metze , Desmond Elliott

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

Sound · Computer Science 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

Machine Learning · Computer Science 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li
‹ Prev 1 2 3 10 Next ›