中文
相关论文

相关论文: VividListener: Expressive and Controllable Listene…

200 篇论文

We introduce a video framework for modeling the association between verbal and non-verbal communication during dyadic conversation. Given the input speech of a speaker, our approach retrieves a video of a listener, who has facial…

计算机视觉与模式识别 · 计算机科学 2023-01-27 Scott Geng , Revant Teotia , Purva Tendulkar , Sachit Menon , Carl Vondrick

Earables, such as True Wireless Stereo earphones and VR/AR headsets, are increasingly popular, yet their compact design poses challenges for robust voice-related applications like telecommunication and voice assistant interactions in noisy…

声音 · 计算机科学 2025-12-03 Lixing He , Yunqi Guo , Haozheng Hou , Zhenyu Yan

Purpose: Emotion is a fundamental component of human communication, shaping understanding, trust, and engagement across domains such as education, healthcare, and mental health. While large language models (LLMs) exhibit strong reasoning…

计算与语言 · 计算机科学 2025-10-15 Yurui Dong , Luozhijie Jin , Yao Yang , Bingjie Lu , Jiaxi Yang , Zhi Liu

The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR),…

图像与视频处理 · 电气工程与系统科学 2025-02-11 Nicholas Evans , Stephen Baker , Miles Reed

Despite significant advancements in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), current models still face substantial challenges in handling complex, multi-turn, and visually-grounded tasks that demand deep…

计算与语言 · 计算机科学 2025-08-22 Seungmin Han , Haeun Kwon , Ji-jun Park , Taeyang Yoon

Current audio-driven 3D head generation methods mainly focus on single-speaker scenarios, lacking natural, bidirectional listen-and-speak interaction. Achieving seamless conversational behavior, where speaking and listening states…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Lei Zhu , Lijian Lin , Ye Zhu , Jiahao Wu , Xuehan Hou , Yu Li , Yunfei Liu , Jie Chen

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription…

音频与语音处理 · 电气工程与系统科学 2025-01-09 Xinyu Wang , Haotian Jiang , Haolin Huang , Yu Fang , Mengjie Xu , Qian Wang

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional…

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme…

图形学 · 计算机科学 2023-01-18 Linchao Bao , Haoxian Zhang , Yue Qian , Tangli Xue , Changhai Chen , Xuefei Zhe , Di Kang

Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and…

音频与语音处理 · 电气工程与系统科学 2026-02-16 Mandip Goswami

We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate in long and…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Mohan Zhou , Yalong Bai , Wei Zhang , Ting Yao , Tiejun Zhao

Non verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepresence, in form of an avatar, it is important to model these…

计算机视觉与模式识别 · 计算机科学 2019-10-08 Chaitanya Ahuja , Shugao Ma , Louis-Philippe Morency , Yaser Sheikh

Room Impulse Responses (RIRs) enable realistic acoustic simulation, with applications ranging from multimedia production to speech data augmentation. However, acquiring high-quality real-world RIRs is labor-intensive, and data scarcity…

音频与语音处理 · 电气工程与系统科学 2026-05-14 Kirak Kim , Sungyoung Kim

Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Phyo Thet Yee , Dimitrios Kollias , Sudeepta Mishra , Abhinav Dhall

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a joint audio-text…

计算机视觉与模式识别 · 计算机科学 2021-12-08 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

Understanding the intrinsic mechanisms of social platforms is an urgent demand to maintain social stability. The rise of large language models provides significant potential for social network simulations to capture attitude dynamics and…

物理与社会 · 物理学 2025-12-16 Yanhui Sun , Wu Liu , Wentao Wang , Hantao Yao , Jiebo Luo , Yongdong Zhang

Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectional conversational…

音频与语音处理 · 电气工程与系统科学 2025-11-25 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Multimodal emotion analysis is shifting from static classification to generative reasoning. Beyond simple label prediction, robust affective reasoning must synthesize fine-grained signals such as facial micro-expressions and prosodic which…

多媒体 · 计算机科学 2026-02-05 Zhixian Zhao , Wenjie Tian , Lei Xie

Bootstrapping from pre-trained language models has been proven to be an efficient approach for building vision-language models (VLM) for tasks such as image captioning or visual question answering. However, outputs of these models rarely…

机器学习 · 计算机科学 2023-06-01 Manuel Brack , Patrick Schramowski , Björn Deiseroth , Kristian Kersting

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world.…

声音 · 计算机科学 2026-02-04 Chengyuan Ma , Jiawei Jin , Ruijie Xiong , Chunxiang Jin , Canxiang Yan , Wenming Yang