中文
相关论文

相关论文: MUSE: Manipulating Unified Framework for Synthesiz…

200 篇论文

Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media, ranging from movies to social media posts. Machine learning models that can synthesize music are…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Sanjoy Chowdhury , Sayan Nag , K J Joseph , Balaji Vasan Srinivasan , Dinesh Manocha

In recent years, emotional Text-to-Speech (TTS) synthesis and emphasis-controllable speech synthesis have advanced significantly. However, their interaction remains underexplored. We propose Emphasis Meets Emotion TTS (EME-TTS), a novel…

声音 · 计算机科学 2025-07-17 Haoxun Li , Leyuan Qu , Jiaxi Hu , Taihao Li

Cognitive reappraisal is a key strategy in emotion regulation, involving reinterpretation of emotionally charged stimuli to alter affective responses. Despite its central role in clinical and cognitive science, real-world reappraisal…

机器学习 · 计算机科学 2025-07-16 Edoardo Pinzuti , Oliver Tüscher , André Ferreira Castro

Synthesizing realistic data samples is of great value for both academic and industrial communities. Deep generative models have become an emerging topic in various research areas like computer vision and signal processing. Affective…

计算机视觉与模式识别 · 计算机科学 2020-11-10 Noushin Hajarolasvadi , Miguel Arjona Ramírez , Hasan Demirel

This paper examines potential biases and inconsistencies in emotional evocation of images produced by generative artificial intelligence (AI) models and their potential bias toward negative emotions. In particular, we assess this bias by…

计算机与社会 · 计算机科学 2024-12-17 Maneet Mehta , Cody Buntain

Generating long-form audio-visual stories from a short user prompt remains challenging due to an intent-execution gap, where high-level narrative intent must be preserved across coherent, shot-level multimodal generation over long horizons.…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Wenzhang Sun , Zhenyu Wang , Zhangchi Hu , Chunfeng Wang , Hao Li , Wei Chen

Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yujiang Pu , Zhanbo Huang , Vishnu Boddeti , Yu Kong

We present a novel information-theoretic approach to introduce dependency among features of a deep convolutional neural network (CNN). The core idea of our proposed method, called MUSE, is to combine MUtual information and SElf-information…

计算机视觉与模式识别 · 计算机科学 2021-10-26 Yu Gong , Ye Yu , Gaurav Mittal , Greg Mori , Mei Chen

We propose a multi-singer emotional singing voice synthesizer, Muse-SVS, that expresses emotion at various intensity levels by controlling subtle changes in pitch, energy, and phoneme duration while accurately following the score. To…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Sungjae Kim , Yewon Kim , Jewoo Jun , Injung Kim

This paper presents a new synthesis-based approach for batch image processing. Unlike existing tools that can only apply global edits to the entire image, our method can apply fine-grained edits to individual objects within the image. For…

编程语言 · 计算机科学 2023-06-16 Celeste Barnaby , Qiaochu Chen , Roopsha Samanta , Isil Dillig

The emergence of diffusion models has significantly advanced image synthesis. The recent studies of model interaction and self-corrective reasoning approach in large language models offer new insights for enhancing text-to-image models.…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Zhongjie Duan , Qianyi Zhao , Cen Chen , Daoyuan Chen , Wenmeng Zhou , Yaliang Li , Yingda Chen

Affective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotional responses. This task is inherently complex due to its twofold objective: significantly evoking the intended emotion, while preserving the…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Jingyuan Yang , Jiawei Feng , Weibin Luo , Dani Lischinski , Daniel Cohen-Or , Hui Huang

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings…

计算与语言 · 计算机科学 2022-11-22 Guimin Hu , Ting-En Lin , Yi Zhao , Guangming Lu , Yuchuan Wu , Yongbin Li

Text-to-image synthesis models require the ability to generate diverse images while maintaining stability. To overcome this challenge, a number of methods have been proposed, including the collection of prompt-image datasets and the…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Keunwoo Park , Jihye Chae , Joong Ho Ahn , Jihoon Kweon

The capability to automatically detect human stress can benefit artificial intelligent agents involved in affective computing and human-computer interaction. Stress and emotion are both human affective states, and stress has proven to have…

计算与语言 · 计算机科学 2021-05-19 Yiqun Yao , Michalis Papakostas , Mihai Burzo , Mohamed Abouelenien , Rada Mihalcea

Emotion serves as an essential component in daily human interactions. Existing human motion generation frameworks do not consider the impact of emotions, which reduces naturalness and limits their application in interactive tasks, such as…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Chen Zhu , Buzhen Huang , Zijing Wu , Binghui Zuo , Yangang Wang

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

In this paper, we conduct a study on the state-of-the-art methods for text-to-image synthesis and propose a framework to evaluate these methods. We consider syntheses where an image contains a single or multiple objects. Our study outlines…

计算机视觉与模式识别 · 计算机科学 2022-07-20 Tan M. Dinh , Rang Nguyen , Binh-Son Hua

In computational pathology, few-shot whole slide image classification is primarily driven by the extreme scarcity of expert-labeled slides. Recent vision-language methods incorporate textual semantics generated by large language models, but…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Jiahao Xu , Sheng Huang , Xin Zhang , Zhixiong Nan , Jiajun Dong , Nankun Mu

Large language models (LLMs) have recently advanced text-driven 3D generation, yet Text-to-CAD remains far from supporting industrial product design. Existing benchmarks focus primarily on generating single-part CAD models and evaluate them…

人工智能 · 计算机科学 2026-05-28 Xiaoyu Dong , Zhi Li , Xiao-Ming Wu