English
Related papers

Related papers: Sentiment-oriented Transformer-based Variational A…

200 papers

In Emotion Recognition in Conversations (ERC), the emotions of target utterances are closely dependent on their context. Therefore, existing works train the model to generate the response of the target utterance, which aims to recognise…

Computation and Language · Computer Science 2023-05-30 Kailai Yang , Tianlin Zhang , Sophia Ananiadou

Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xuecheng Wu , Heli Sun , Junxiao Xue , Jiayu Nie , Xiangyan Kong , Ruofan Zhai , Liang He

Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective…

Sound · Computer Science 2020-07-29 Siddique Latif , Rajib Rana , Junaid Qadir , Julien Epps

In recent years, Variational Autoencoders (VAEs) have been shown to be highly effective in both standard collaborative filtering applications and extensions such as incorporation of implicit feedback. We extend VAEs to collaborative…

Machine Learning · Statistics 2018-10-09 Giannis Karamanolakis , Kevin Raji Cherian , Ananth Ravi Narayan , Jie Yuan , Da Tang , Tony Jebara

This paper presents a novel spatiotemporal transformer network that introduces several original components to detect actions in untrimmed videos. First, the multi-feature selective semantic attention model calculates the correlations…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Matthew Korban , Peter Youngs , Scott T. Acton

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng

With the increasing integration of intelligent driving functions into serial-produced vehicles, ensuring their functionality and robustness poses greater challenges. Compared to traditional road testing, scenario-based virtual testing…

Robotics · Computer Science 2025-10-29 Li Li , Tobias Brinkmann , Till Temmen , Markus Eisenbarth , Jakob Andert

Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Jing Li , Di Kang , Wenjie Pei , Xuefei Zhe , Ying Zhang , Zhenyu He , Linchao Bao

Generating a vivid, novel, and diverse essay with only several given topic words is a challenging task of natural language generation. In previous work, there are two problems left unsolved: neglect of sentiment beneath the text and…

Computation and Language · Computer Science 2020-10-13 Lin Qiao , Jianhao Yan , Fandong Meng , Zhendong Yang , Jie Zhou

Recent years have seen remarkable progress in speech emotion recognition (SER), thanks to advances in deep learning techniques. However, the limited availability of labeled data remains a significant challenge in the field. Self-supervised…

Sound · Computer Science 2023-04-24 Samir Sadok , Simon Leglaive , Renaud Séguier

Live commenting on video streams has surged in popularity on platforms like Twitch, enhancing viewer engagement through dynamic interactions. However, automatically generating contextually appropriate comments remains a challenging and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Anam Fatima , Yi Yu , Janak Kapuriya , Julien Lalanne , Jainendra Shukla

Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most existing research assume that all modalities are available during both training and testing,…

Sound · Computer Science 2026-04-21 Weide Liu , Huijing Zhan

Causal Emotion Entailment (CEE) aims to discover the potential causes behind an emotion in a conversational utterance. Previous works formalize CEE as independent utterance pair classification problems, with emotion and speaker information…

Computation and Language · Computer Science 2022-09-23 Duzhen Zhang , Zhen Yang , Fandong Meng , Xiuyi Chen , Jie Zhou

Images shared online strongly influence emotions and public well-being. Understanding the emotions an image elicits is therefore vital for fostering healthier and more sustainable digital communities, especially during public crises. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Fanhang Man , Xiaoyue Chen , Huandong Wang , Baining Zhao , Han Li , Xinlei Chen

Syntactic information contains structures and rules about how text sentences are arranged. Incorporating syntax into text modeling methods can potentially benefit both representation learning and generation. Variational autoencoders (VAEs)…

Computation and Language · Computer Science 2019-08-28 Yijun Xiao , William Yang Wang

To generate proper captions for videos, the inference needs to identify relevant concepts and pay attention to the spatial relationships between them as well as to the temporal development in the clip. Our end-to-end encoder-decoder video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-22 Zohreh Ghaderi , Leonard Salewski , Hendrik P. A. Lensch

In today's day and age when almost every industry has an online presence with users interacting in online marketplaces, personalized recommendations have become quite important. Traditionally, the problem of collaborative filtering has been…

Information Retrieval · Computer Science 2018-09-25 Kilol Gupta , Mukund Yelahanka Raghuprasad , Pankhuri Kumar

To synthesize a realistic action sequence based on a single human image, it is crucial to model both motion patterns and diversity in the action video. This paper proposes an Action Conditional Temporal Variational AutoEncoder (ACT-VAE) to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-13 Xiaogang Xu , Yi Wang , Liwei Wang , Bei Yu , Jiaya Jia

Understanding Affect from video segments has brought researchers from the language, audio and video domains together. Most of the current multimodal research in this area deals with various techniques to fuse the modalities, and mostly…

Computation and Language · Computer Science 2018-06-11 Saurav Sahay , Shachi H Kumar , Rui Xia , Jonathan Huang , Lama Nachman

Extracorporeal membrane oxygenation (ECMO) is an essential life-supporting modality for COVID-19 patients who are refractory to conventional therapies. However, the proper treatment decision has been the subject of significant debate and it…

Machine Learning · Computer Science 2023-07-10 Bing Xue , Ahmed Sameh Said , Ziqi Xu , Hanyang Liu , Neel Shah , Hanqing Yang , Philip Payne , Chenyang Lu