中文
相关论文

相关论文: Aether Weaver: Multimodal Affective Narrative Co-G…

200 篇论文

Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on carefully designed fusion strategies that either use…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Zelin Zhang , Kedi Li , Huiqi Liang , Tao Zhang , Chuanzhi Xu

Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length…

多智能体系统 · 计算机科学 2025-10-03 Haoyuan Shi , Yunxin Li , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph…

多媒体 · 计算机科学 2020-05-15 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli

Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the UniEmo, a unified framework that seamlessly integrates…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yijie Zhu , Lingsen Zhang , Zitong Yu , Rui Shao , Tao Tan , Liqiang Nie

Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient…

Due to the complex nature of human emotions and the diversity of emotion representation methods in humans, emotion recognition is a challenging field. In this research, three input modalities, namely text, audio (speech), and video, are…

人工智能 · 计算机科学 2024-02-13 Minoo Shayaninasab , Bagher Babaali

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Farzaneh Jafari , Stefano Berretti , Anup Basu

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video…

多媒体 · 计算机科学 2024-12-31 Dingyi Yang , Chunru Zhan , Ziheng Wang , Biao Wang , Tiezheng Ge , Bo Zheng , Qin Jin

Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trained text-to-image model into an auto-regressive manner.…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Ming Tao , Bing-Kun Bao , Hao Tang , Yaowei Wang , Changsheng Xu

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Guanwen Feng , Haoran Cheng , Yunan Li , Zhiyuan Ma , Chaoneng Li , Zhihao Qian , Qiguang Miao , Chi-Man Pun

Visual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Xiaoyue Mi , Fan Tang , Juan Cao , Qiang Sheng , Ziyao Huang , Peng Li , Yang Liu , Tong-Yee Lee

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grained emotional modeling, existing systems remain limited to…

计算与语言 · 计算机科学 2026-04-30 Shuhao Xu , Yifan Hu , Jingjing Wu , Zhihao Du , Zheng Lian , Rui Liu

Earnings calls represent a uniquely rich and semi-structured source of financial communication, blending scripted managerial commentary with unscripted analyst dialogue. Although recent advances in financial sentiment analysis have…

计算与语言 · 计算机科学 2025-09-05 Alejandro Álvarez Castro , Joaquín Ordieres-Meré

Recent advances in AI-driven storytelling have enhanced video generation and story visualization. However, translating dialogue-centric scripts into coherent storyboards remains a significant challenge due to limited script detail,…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Min Zhang , Zilin Wang , Liyan Chen , Kunhong Liu , Juncong Lin

Visual Emotion Analysis (VEA) aims at finding out how people feel emotionally towards different visual stimuli, which has attracted great attention recently with the prevalence of sharing images on social networks. Since human emotion…

计算机视觉与模式识别 · 计算机科学 2021-10-26 Jingyuan Yang , Xinbo Gao , Leida Li , Xiumei Wang , Jinshan Ding

Multimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategies, they often lack…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Guanyu Hu , Dimitrios Kollias , Xinyu Yang

Stylized Text-to-Image Generation (STIG) aims to generate images from text prompts and style reference images. In this paper, we present ArtWeaver, a novel framework that leverages pretrained Stable Diffusion (SD) to address challenges such…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Chengming Xu , Kai Hu , Qilin Wang , Donghao Luo , Jiangning Zhang , Xiaobin Hu , Yanwei Fu , Chengjie Wang

Human emotions are essentially molded by lived experiences, from which we construct personalised meaning. The engagement in such meaning-making process has been practiced as an intervention in various psychotherapies to promote wellness.…

人机交互 · 计算机科学 2024-03-06 Qian Wan , Xin Feng , Yining Bei , Zhiqi Gao , Zhicong Lu

Texture map production is an important part of 3D modeling and determines the rendering quality. Recently, diffusion-based methods have opened a new way for texture generation. However, restricted control flexibility and limited prompt…

图形学 · 计算机科学 2025-06-04 Dongyu Yan , Leyi Wu , Jiantao Lin , Luozhou Wang , Tianshuo Xu , Zhifei Chen , Zhen Yang , Lie Xu , Shunsi Zhang , Yingcong Chen