中文
相关论文

相关论文: Aether Weaver: Multimodal Affective Narrative Co-G…

200 篇论文

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Zhifei Yang , Keyang Lu , Chao Zhang , Jiaxing Qi , Hanqi Jiang , Ruifei Ma , Shenglin Yin , Yifan Xu , Mingzhe Xing , Zhen Xiao , Jieyi Long , Guangyao Zhai

Achieving empathy is a crucial step toward humanized dialogue systems. Current approaches for empathetic dialogue generation mainly perceive an emotional label to generate an empathetic response conditioned on it, which simply treat…

计算与语言 · 计算机科学 2023-11-28 Fengyi Fu , Lei Zhang , Quan Wang , Zhendong Mao

While diffusion models generate high-fidelity video clips, transforming them into coherent storytelling engines remains challenging. Current agentic pipelines automate this via chained modules but suffer from semantic drift and cascading…

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

多媒体 · 计算机科学 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

Storytelling is an open-ended task that entails creative thinking and requires a constant flow of ideas. Natural language generation (NLG) for storytelling is especially challenging because it requires the generated text to follow an…

计算与语言 · 计算机科学 2021-09-17 Eden Bensaid , Mauro Martino , Benjamin Hoover , Hendrik Strobelt

Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and roughly concatenate them as a story, which leads to the…

计算与语言 · 计算机科学 2020-11-02 Ruize Wang , Zhongyu Wei , Ying Cheng , Piji Li , Haijun Shan , Ji Zhang , Qi Zhang , Xuanjing Huang

Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Existing methods struggle to generalize to multi-humanoid…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Xirui Hu , Yanbo Ding , Jiahao Wang , Tingting Shi , Yali Wang , Guo Zhi Zhi , Weizhan Zhang

Emotional Video Captioning is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectively perceive subtle and ambiguous visual emotional cues during…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Cheng Ye , Weidong Chen , Jingyu Li , Lei Zhang , Zhendong Mao

Multimodal Emotion Recognition in Conversation (ERC) plays an influential role in the field of human-computer interaction and conversational robotics since it can motivate machines to provide empathetic services. Multimodal data modeling is…

多媒体 · 计算机科学 2023-11-23 Jiang Li , Xiaoping Wang , Guoqing Lv , Zhigang Zeng

Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Sreehari Rajan , Kunal Bhosikar , Charu Sharma

This paper tackles \textbf{open-ended deep research (OEDR)}, a complex challenge where AI agents must synthesize vast web-scale information into insightful reports. Current approaches are plagued by dual-fold limitations: static research…

计算与语言 · 计算机科学 2025-10-08 Zijian Li , Xin Guan , Bo Zhang , Shen Huang , Houquan Zhou , Shaopeng Lai , Ming Yan , Yong Jiang , Pengjun Xie , Fei Huang , Jun Zhang , Jingren Zhou

The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and world modeling. While panoramic images have proven effective…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Felix Wimbauer , Fabian Manhardt , Michael Oechsle , Nikolai Kalischek , Christian Rupprecht , Daniel Cremers , Federico Tombari

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modalities via multimodal…

声音 · 计算机科学 2024-12-10 Feng Li , Jiusong Luo , Wanjun Xia

Emotional Image Content Generation (EICG) aims to generate semantically clear and emotionally faithful images based on given emotion categories, with broad application prospects. While recent text-to-image diffusion models excel at…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Kaishen Yuan , Yuting Zhang , Shang Gao , Yijie Zhu , Wenshuo Chen , Yutao Yue

Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capabilities that existing approaches struggle to unify. We…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Haoxuan Li , Mengyan Li , Junjun Zheng

Multimodal emotion understanding requires effective integration of text, audio, and visual modalities for both discrete emotion recognition and continuous sentiment analysis. We present EGMF, a unified framework combining expert-guided…

计算与语言 · 计算机科学 2026-01-13 Jiaqi Qiao , Xiujuan Xu , Xinran Li , Yu Liu

Recent unified models have made unprecedented progress in both understanding and generation. However, while most of them accept multi-modal inputs, they typically produce only single-modality outputs. This challenge of producing interleaved…

Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Qinglin Zeng , Kaitong Cai , Ruiqi Chen , Qinhan Lv , Keze Wang

Multimodal Variational Autoencoders have emerged as a popular tool to extract effective representations from rich multimodal data. However, such models rely on fusion strategies in latent space that destroy the joint statistical structure…

机器学习 · 计算机科学 2026-03-03 Federico Caretti , Guido Sanguinetti

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Haotian Wang , Yuzhe Weng , Yueyan Li , Zilu Guo , Jun Du , Shutong Niu , Jiefeng Ma , Shan He , Xiaoyan Wu , Qiming Hu , Bing Yin , Cong Liu , Qingfeng Liu