中文
相关论文

相关论文: Detecting expressions with multimodal transformers

200 篇论文

We used two multimodal models for continuous valence-arousal recognition using visual, audio, and linguistic information. The first model is the same as we used in ABAW2 and ABAW3, which employs the leader-follower attention. The second…

多媒体 · 计算机科学 2023-04-18 Su Zhang , Ziyuan Zhao , Cuntai Guan

Automatic human affect recognition is a key step towards more natural human-computer interaction. Recent trends include recognition in the wild using a fusion of audiovisual and physiological sensors, a challenging setting for conventional…

机器学习 · 计算机科学 2019-01-11 Philipp V. Rouast , Marc T. P. Adam , Raymond Chiong

Automated co-located human-human interaction analysis has been addressed by the use of nonverbal communication as measurable evidence of social and psychological phenomena. We survey the computing studies (since 2010) detecting phenomena…

人机交互 · 计算机科学 2023-10-05 Cigdem Beyan , Alessandro Vinciarelli , Alessio Del Bue

The goal of our research is to automatically retrieve the satisfaction and the frustration in real-life call-center conversations. This study focuses an industrial application in which the customer satisfaction is continuously tracked down…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Manon Macary , Marie Tahon , Yannick Estève , Daniel Luzzati

In a voice-controlled smart-home, a controller must respond not only to user's requests but also according to the interaction context. This paper describes Arcades, a system which uses deep reinforcement learning to extract context from a…

机器学习 · 计算机科学 2018-07-19 Alexis Brenon , François Portet , Michel Vacher

Natural human interactions for Mixed Reality Applications are overwhelmingly multimodal: humans communicate intent and instructions via a combination of visual, aural and gestural cues. However, supporting low-latency and accurate…

Audio deepfake is so sophisticated that the lack of effective detection methods is fatal. While most detection systems primarily rely on low-level acoustic features or pretrained speech representations, they frequently neglect high-level…

声音 · 计算机科学 2025-09-16 Xiaokang Li , Yicheng Gong , Dinghao Zou , Xin Cao , Sunbowen Lee

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

图像与视频处理 · 电气工程与系统科学 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

Dynamic Facial Expression Recognition(DFER) is a rapidly evolving field of research that focuses on the recognition of time-series facial expressions. While previous research on DFER has concentrated on feature learning from a deep learning…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Feng Liu , Lingna Gu , Chen Shi , Xiaolan Fu

Artificial Intelligence algorithms have now become pervasive in multiple high-stakes domains. However, their internal logic can be obscure to humans. Explainable Artificial Intelligence aims to design tools and techniques to illustrate the…

人机交互 · 计算机科学 2024-04-29 Eleonora Cappuccio , Daniele Fadda , Rosa Lanzilotti , Salvatore Rinzivillo

It has already been observed that audio-visual embedding is more robust than uni-modality embedding for person verification. Here, we proposed a novel audio-visual strategy that considers aggregators from a fusion perspective. First, we…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Peiwen Sun , Shanshan Zhang , Zishan Liu , Yougen Yuan , Taotao Zhang , Honggang Zhang , Pengfei Hu

This paper aims to bring a new lightweight yet powerful solution for the task of Emotion Recognition and Sentiment Analysis. Our motivation is to propose two architectures based on Transformers and modulation that combine the linguistic and…

计算与语言 · 计算机科学 2020-10-06 Jean-Benoit Delbrouck , Noé Tits , Stéphane Dupont

Much of the work on automatic facial expression recognition relies on databases containing a certain number of emotion classes and their exaggerated facial configurations (generally six prototypical facial expressions), based on Ekman's…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Wenjing Yan , Shan Li , Chengtao Que , JiQuan Pei , Weihong Deng

Affect understanding capability is essential for social robots to autonomously interact with a group of users in an intuitive and reciprocal way. However, the challenge of multi-person affect understanding comes from not only the accurate…

计算机视觉与模式识别 · 计算机科学 2023-01-02 Yubin Kim , Huili Chen , Sharifa Alghowinem , Cynthia Breazeal , Hae Won Park

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

计算机与社会 · 计算机科学 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

Transformers have seen an unprecedented rise in Natural Language Processing and Computer Vision tasks. However, in audio tasks, they are either infeasible to train due to extremely large sequence length of audio waveforms or incur a…

机器学习 · 计算机科学 2022-02-02 Surya Kant Sahu , Sai Mitheran , Juhi Kamdar , Meet Gandhi

The recent proliferation of hyper-realistic deepfake videos has drawn attention to the threat of audio and visual forgeries. Most previous studies on detecting artificial intelligence-generated fake videos only utilize visual modality or…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Ammarah Hashmi , Sahibzada Adil Shahzad , Chia-Wen Lin , Yu Tsao , Hsin-Min Wang

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Davide Berghi , Philip J. B. Jackson

With the recent advancements in Artificial Intelligence (AI), Intelligent Virtual Assistants (IVA) such as Alexa, Google Home, etc., have become a ubiquitous part of many homes. Currently, such IVAs are mostly audio-based, but going…

多媒体 · 计算机科学 2019-12-27 Shachi H Kumar , Eda Okur , Saurav Sahay , Jonathan Huang , Lama Nachman

In social robotics, endowing humanoid robots with the ability to generate bodily expressions of affect can improve human-robot interaction and collaboration, since humans attribute, and perhaps subconsciously anticipate, such traces to…

机器人学 · 计算机科学 2022-05-03 Mina Marmpena , Fernando Garcia , Angelica Lim , Nikolas Hemion , Thomas Wennekers
‹ 上一页 1 8 9 10 下一页 ›