English
Related papers

Related papers: Aether Weaver: Multimodal Affective Narrative Co-G…

200 papers

Traditional techniques for emotion recognition have focused on the facial expression analysis only, thus providing limited ability to encode context that comprehensively represents the emotional responses. We present deep networks for…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Jiyoung Lee , Seungryong Kim , Sunok Kim , Jungin Park , Kwanghoon Sohn

Design processes involve exploration, iteration, and movement across interconnected stages such as persona creation, problem framing, solution ideation, and prototyping. However, time and resource constraints often hinder designers from…

Human-Computer Interaction · Computer Science 2025-08-18 Sangho Suh , Michael Lai , Kevin Pu , Steven P. Dow , Tovi Grossman

Affective Computing (AC) is essential for advancing Artificial General Intelligence (AGI), with emotion recognition serving as a key component. However, human emotions are inherently dynamic, influenced not only by an individual's…

Computation and Language · Computer Science 2025-03-31 Yupei Li , Qiyang Sun , Sunil Munthumoduku Krishna Murthy , Emran Alturki , Björn W. Schuller

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Huaize Liu , Wenzhang Sun , Donglin Di , Shibo Sun , Jiahui Yang , Changqing Zou , Hujun Bao

This report presents a comprehensive framework for generating high-quality 3D shapes and textures from diverse input prompts, including single images, multi-view images, and text descriptions. The framework consists of 3D shape generation…

Multimodal emotion recognition (MER) is crucial for enabling emotionally intelligent systems that perceive and respond to human emotions. However, existing methods suffer from limited cross-modal interaction and imbalanced contributions…

Multimedia · Computer Science 2025-07-30 Zeyu Deng , Yanhui Lu , Jiashu Liao , Shuang Wu , Chongfeng Wei

Text-to-Avatar generation has recently made significant strides due to advancements in diffusion models. However, most existing work remains constrained by limited diversity, producing avatars with subtle differences in appearance for a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Weijing Tao , Biwen Lei , Kunhao Liu , Shijian Lu , Miaomiao Cui , Xuansong Xie , Chunyan Miao

Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual information of the current sentence to enhance the prosody of…

Multimedia · Computer Science 2024-09-05 Yuan Zhao , Zhenqi Jia , Rui Liu , De Hu , Feilong Bao , Guanglai Gao

Multimodal emotion recognition in conversation (MERC) seeks to identify the speakers' emotions expressed in each utterance, offering significant potential across diverse fields. The challenge of MERC lies in balancing speaker modeling and…

Multimedia · Computer Science 2025-07-25 Zijian Yi , Ziming Zhao , Zhishu Shen , Tiehua Zhang

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Authoring realistic haptic textures typically requires low-level parameter tuning and repeated trial-and-error, limiting speed, transparency, and creative reach. We present a language-driven authoring system that turns natural-language…

Human-Computer Interaction · Computer Science 2026-04-09 Wanli Qian , Aiden Chang , Shihan Lu , Michael Gu , Heather Culbertson

Emotions play a key role in human communication and public presentations. Human emotions are usually expressed through multiple modalities. Therefore, exploring multimodal emotions and their coherence is of great value for understanding…

Computer Vision and Pattern Recognition · Computer Science 2019-10-10 Haipeng Zeng , Xingbo Wang , Aoyu Wu , Yong Wang , Quan Li , Alex Endert , Huamin Qu

Attention-based encoder-decoder framework is widely used in the scene text recognition task. However, for the current state-of-the-art(SOTA) methods, there is room for improvement in terms of the efficient usage of local visual and global…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Mengmeng Cui , Wei Wang , Jinjin Zhang , Liang Wang

Story continuation focuses on generating the next image in a narrative sequence so that it remains coherent with both the ongoing text description and the previously observed images. A central challenge in this setting lies in utilizing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Seyed Mohammad Mousavi , Morteza Analoui

Sentiment analysis and emotion recognition in videos are challenging tasks, given the diversity and complexity of the information conveyed in different modalities. Developing a highly competent framework that effectively addresses the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Prasad Chaudhari , Aman Kumar , Chandravardhan Singh Raghaw , Mohammad Zia Ur Rehman , Nagendra Kumar

Emotion recognition from EEG signals is essential for affective computing and has been widely explored using deep learning. While recent deep learning approaches have achieved strong performance on single EEG emotion datasets, their…

Machine Learning · Computer Science 2025-11-17 Yuning Chen , Sha Zhao , Shijian Li , Gang Pan

An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data. In this paper, we propose the VQ-MAE-AV model, a self-supervised multimodal model that leverages masked autoencoders to learn…

Sound · Computer Science 2025-05-12 Samir Sadok , Simon Leglaive , Renaud Séguier

The dyadic reaction generation task involves synthesizing responsive facial reactions that align closely with the behaviors of a conversational partner, enhancing the naturalness and effectiveness of human-like interaction simulations. This…

Machine Learning · Computer Science 2025-05-14 Minh-Duc Nguyen , Hyung-Jeong Yang , Soo-Hyung Kim , Ji-Eun Shin , Seung-Won Kim

Emotion recognition based on Electroencephalography (EEG) has gained significant attention and diversified development in fields such as neural signal processing and affective computing. However, the unique brain anatomy of individuals…

Signal Processing · Electrical Eng. & Systems 2024-05-31 Yihang Dong , Xuhang Chen , Yanyan Shen , Michael Kwok-Po Ng , Tao Qian , Shuqiang Wang

Synthesizing realistic data samples is of great value for both academic and industrial communities. Deep generative models have become an emerging topic in various research areas like computer vision and signal processing. Affective…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Noushin Hajarolasvadi , Miguel Arjona Ramírez , Hasan Demirel