English
Related papers

Related papers: AUD-TGN: Advancing Action Unit Detection with Temp…

200 papers

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

This study introduces a novel method that transforms multimodal physiological signalsphotoplethysmography (PPG), galvanic skin response (GSR), and acceleration (ACC) into 2D image matrices to enhance stress detection using convolutional…

Machine Learning · Computer Science 2025-09-18 Yasin Hasanpoor , Bahram Tarvirdizadeh , Khalil Alipour , Mohammad Ghamari

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

Multimedia · Computer Science 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

With the rise in manipulated media, deepfake detection has become an imperative task for preserving the authenticity of digital content. In this paper, we present a novel multi-modal audio-video framework designed to concurrently process…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Aaditya Kharel , Manas Paranjape , Aniket Bera

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to modeling multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-05-05 Yanbei Chen , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set of applications. However, the exploration of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Kaiwen Zheng , Xuri Ge , Junchen Fu , Jun Peng , Joemon M. Jose

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson

In recent years, researchers combine both audio and video signals to deal with challenges where actions are not well represented or captured by visual cues. However, how to effectively leverage the two modalities is still under development.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Wentao Zhu

Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Jun Yu , Lingsi Zhu , Yanjun Chi , Yunxiang Zhang , Yang Zheng , Yongqi Wang , Xilong Lu

Sentiment analysis, mostly based on text, has been rapidly developing in the last decade and has attracted widespread attention in both academia and industry. However, the information in the real world usually comes from multiple…

Computation and Language · Computer Science 2019-12-12 Feiyang Chen , Ziqian Luo , Yanyan Xu , Dengfeng Ke

In this paper, we propose a solution for the semi-supervised learning track (MER-SEMI) in MER2024. First, in order to enhance the performance of the feature extractor on sentiment classification tasks,we fine-tuned video and text feature…

Sound · Computer Science 2024-09-10 Pujin Shi , Fei Gao

Understanding animal vocalizations through multi-source data fusion is crucial for assessing emotional states and enhancing animal welfare in precision livestock farming. This study aims to decode dairy cow contact calls by employing…

Sound · Computer Science 2024-11-04 Bubacarr Jobarteh , Madalina Mincu , Gavojdian Dinu , Suresh Neethirajan

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Amit Eliav , Sharon Gannot

In this paper, we propose a solution for improving the quality of temporal sound localization. We employ a multimodal fusion approach to combine visual and audio features. High-quality visual features are extracted using a state-of-the-art…

Sound · Computer Science 2024-07-03 Yurui Huang , Yang Yang , Shou Chen , Xiangyu Wu , Qingguo Chen , Jianfeng Lu

This paper introduces a novel approach for multimodal sentiment analysis on social media, particularly in the context of natural disasters, where understanding public sentiment is crucial for effective crisis management. Unlike conventional…

Machine Learning · Computer Science 2025-08-20 Meriem Zerkouk , Miloud Mihoubi , Belkacem Chikhaoui

Temporal modelling is the key for efficient video action recognition. While understanding temporal information can improve recognition accuracy for dynamic actions, removing temporal redundancy and reusing past features can significantly…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Yue Meng , Rameswar Panda , Chung-Ching Lin , Prasanna Sattigeri , Leonid Karlinsky , Kate Saenko , Aude Oliva , Rogerio Feris

Person or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion rely on score-level or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 R. Gnana Praveen , Jahangir Alam

Emotion recognition has a wide range of applications in human-computer interaction, marketing, healthcare, and other fields. In recent years, the development of deep learning technology has provided new methods for emotion recognition.…

Computation and Language · Computer Science 2025-01-28 Junwei Feng , Xueyan Fan

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha
‹ Prev 1 3 4 5 6 7 10 Next ›