English
Related papers

Related papers: Multimodal Machine Learning Can Predict Videoconfe…

200 papers

The analysis of speech in individuals with amyotrophic lateral sclerosis is a powerful tool to support clinicians in the assessment of bulbar dysfunction. However, current methods used in clinical practice consist of subjective evaluations…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Francesco Pierotti , Andrea Bandini

Multimodal affective computing analyzes user-generated social media content to predict emotional states. However, a critical gap remains in understanding how visual content shapes cognitive interpretations and elicits specific affective…

Artificial Intelligence · Computer Science 2026-04-28 Nastaran Dab , Raziyeh Zall , Mohammadreza Kangavari

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers that jointly learn…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Dota Tianai Dong , Mariya Toneva

Why do some speakers capture a room almost instantly while others fail to connect? The real-time architecture of audience engagement remains largely a black box. Here, we used motion-captured animations to present the pure nonverbal…

Human-Computer Interaction · Computer Science 2026-03-02 Ralf Schmälzle , Yuetong Du , Sue Lim , Gary Bente

Anticipating human actions is an important task that needs to be addressed for the development of reliable intelligent agents, such as self-driving cars or robot assistants. While the ability to make future predictions with high accuracy is…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Olga Zatsarynna , Yazan Abu Farha , Juergen Gall

Studies on emotion recognition (ER) show that combining lexical and acoustic information results in more robust and accurate models. The majority of the studies focus on settings where both modalities are available in training and…

Computation and Language · Computer Science 2019-06-26 Gustavo Aguilar , Viktor Rozgić , Weiran Wang , Chao Wang

Emotion recognition is a core research area at the intersection of artificial intelligence and human communication analysis. It is a significant technical challenge since humans display their emotions through complex idiosyncratic…

Human-Computer Interaction · Computer Science 2018-09-14 Paul Pu Liang , Amir Zadeh , Louis-Philippe Morency

Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most existing research assume that all modalities are available during both training and testing,…

Sound · Computer Science 2026-04-21 Weide Liu , Huijing Zhan

Automated deception detection systems can enhance health, justice, and security in society by helping humans detect deceivers in high-stakes situations across medical and legal domains, among others. This paper presents a novel analysis of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Leena Mathur , Maja J Matarić

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we extend this paradigm by leveraging motion and audio that…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Rui Qian , Yeqing Li , Zheng Xu , Ming-Hsuan Yang , Serge Belongie , Yin Cui

Multimodal deep learning, especially vision-language models, have gained significant traction in recent years, greatly improving performance on many downstream tasks, including content moderation and violence detection. However, standard…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Zhuokai Zhao , Harish Palani , Tianyi Liu , Lena Evans , Ruth Toner

Current state-of-the-art video quality models, such as VMAF, give excellent prediction results by comparing the degraded video with its reference video. However, they do not consider temporal distortions (e.g., frame freezes or skips) that…

Image and Video Processing · Electrical Eng. & Systems 2023-03-23 Gabriel Mittag , Babak Naderi , Vishak Gopal , Ross Cutler

Barriers to accessing mental health assessments including cost and stigma continues to be an impediment in mental health diagnosis and treatment. Machine learning approaches based on speech samples could help in this direction. In this…

Computation and Language · Computer Science 2023-12-27 Prabhat Agarwal , Akshat Jindal , Shreya Singh

It has been shown that learning audiovisual features can lead to improved speech recognition performance over audio-only features, especially for noisy speech. However, in many common applications, the visual features are partially or…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-20 Oscar Chang , Otavio Braga , Hank Liao , Dmitriy Serdyuk , Olivier Siohan

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to modeling multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-05-05 Yanbei Chen , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

Artificial Intelligence · Computer Science 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda

Emotion Recognition in Conversations (ERC) is crucial in developing sympathetic human-machine interaction. In conversational videos, emotion can be present in multiple modalities, i.e., audio, video, and transcript. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Vishal Chudasama , Purbayan Kar , Ashish Gudmalwar , Nirmesh Shah , Pankaj Wasnik , Naoyuki Onoe

Learners' use of video controls in educational videos provides implicit signals of cognitive processing and instructional design quality, yet the lack of scalable and explainable predictive models limits instructors' ability to anticipate…

Artificial Intelligence · Computer Science 2026-04-07 Dominik Glandorf , Fares Fawzi , Tanja Käser

In recent years, there have been numerous developments towards solving multimodal tasks, aiming to learn a stronger representation than through a single modality. Certain aspects of the data can be particularly useful in this case - for…

Machine Learning · Statistics 2023-09-06 Cătălina Cangea , Petar Veličković , Pietro Liò

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other…

Multimedia · Computer Science 2026-05-05 Mayesha Maliha R. Mithila , Mylene C. Q. Farias