中文
相关论文

相关论文: The Right to Talk: An Audio-Visual Transformer App…

200 篇论文

Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed.…

音频与语音处理 · 电气工程与系统科学 2025-10-02 Ishan D. Biyani , Nirmesh J. Shah , Ashishkumar P. Gudmalwar , Pankaj Wasnik , Rajiv R. Shah

The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the temporal relationship of adjacent audio clips is highly…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Zhihua Xu , Tianshui Chen , Zhijing Yang , Siyuan Peng , Keze Wang , Liang Lin

Dynamically synthesizing talking speech that actively responds to a listening head is critical during the face-to-face interaction. For example, the speaker could take advantage of the listener's facial expression to adjust the tones,…

音频与语音处理 · 电气工程与系统科学 2023-06-22 Mohan Zhou , Yalong Bai , Wei Zhang , Ting Yao , Tiejun Zhao , Tao Mei

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

图像与视频处理 · 电气工程与系统科学 2022-09-27 Rahul Sharma , Shrikanth Narayanan

Speaker identity plays a significant role in human communication and is being increasingly used in societal applications, many through advances in machine learning. Speaker identity perception is an essential cognitive phenomenon that can…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Gasser Elbanna

Sound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on utilizing deep…

Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first introduce speaker…

声音 · 计算机科学 2023-12-19 Peng Shen , Xugang Lu , Hisashi Kawai

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Zhe Kong , Feng Gao , Yong Zhang , Zhuoliang Kang , Xiaoming Wei , Xunliang Cai , Guanying Chen , Wenhan Luo

In this paper, we present a novel multi-modal attention guidance method designed to address the challenges of turn-taking dynamics in meetings and enhance group conversations within virtual reality (VR) environments. Recognizing the…

人机交互 · 计算机科学 2024-06-21 Geonsun Lee , Dae Yeol Lee , Guan-Ming Su , Dinesh Manocha

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

Turn-taking management is crucial for any social interaction. Still, it is challenging to model human-machine interaction due to the complexity of the social context and its multimodal nature. Unlike conventional systems based on silence…

计算与语言 · 计算机科学 2025-06-05 Takeshi Saga , Catherine Pelachaud

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate…

计算机视觉与模式识别 · 计算机科学 2018-02-13 Aviv Gabbay , Ariel Ephrat , Tavi Halperin , Shmuel Peleg

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefined conditions,…

声音 · 计算机科学 2020-12-01 Peng Zhang , Jiaming Xu , Jing shi , Yunzhe Hao , Bo Xu

Conversational AI is constrained in many real-world settings where only one side of a dialogue can be recorded, such as telemedicine, call centers, and smart glasses. We formalize this as the one-sided conversation problem (1SC): inferring…

计算与语言 · 计算机科学 2026-04-20 Victoria Ebert , Rishabh Singh , Tuochao Chen , Noah A. Smith , Shyamnath Gollakota

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Junhyeok Kim , Min Soo Kim , Jiwan Chung , Jungbin Cho , Jisoo Kim , Sungwoong Kim , Gyeongbo Sim , Youngjae Yu

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Triantafyllos Afouras , Andrew Owens , Joon Son Chung , Andrew Zisserman

Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With…

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Hongyu Wang , Hui Li , Bo Li

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not…

计算与语言 · 计算机科学 2025-06-27 Anne Wu , Laurent Mazaré , Neil Zeghidour , Alexandre Défossez

Determining the head orientation of a talker is not only beneficial for various speech signal processing applications, such as source localization or speech enhancement, but also facilitates intuitive voice control and interaction with…

音频与语音处理 · 电气工程与系统科学 2026-02-10 Kaspar Müller , Bilgesu Çakmak , Paul Didier , Simon Doclo , Jan Østergaard , Tobias Wolff