English
Related papers

Related papers: UniTalk: Towards Universal Active Speaker Detectio…

200 papers

This study considers the problem of detecting and locating an active talker's horizontal position from multichannel audio captured by a microphone array. We refer to this as active speaker detection and localization (ASDL). Our goal was to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-28 Davide Berghi , Philip J. B. Jackson

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Speaker verification is a task of confirming an individual's identity through the analysis of their voice. Whispered speech differs from phonated speech in acoustic characteristics, which degrades the performance of speaker verification…

Sound · Computer Science 2026-05-08 Magdalena Gołębiowska , Piotr Syga

Task-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and real-world spoken…

Computation and Language · Computer Science 2025-06-25 Shuzheng Si , Wentao Ma , Haoyu Gao , Yuchuan Wu , Ting-En Lin , Yinpei Dai , Hangyu Li , Rui Yan , Fei Huang , Yongbin Li

Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, instruction-aligned speech-text data. This bottleneck is…

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Ruohan Gao , Kristen Grauman

Speech-driven 3D talking heads generation has emerged as a significant area of interest among researchers, presenting numerous challenges. Existing methods are constrained by animating faces with fixed topologies, wherein point-wise…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Federico Nocentini , Thomas Besnier , Claudio Ferrari , Sylvain Arguillere , Stefano Berretti , Mohamed Daoudi

Speech signals are subjected to more acoustic interference and emotional factors than other signals. Noisy emotion-riddled speech data is a challenge for real-time speech processing applications. It is essential to find an effective way to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Shibani Hamsa , Ismail Shahin , Youssef Iraqi , Ernesto Damiani , Naoufel Werghi

We propose a Perceiver-based sequence classifier to detect abnormalities in speech reflective of several neurological disorders. We combine this classifier with a Universal Speech Model (USM) that is trained (unsupervised) on 12 million…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-23 Hagen Soltau , Izhak Shafran , Alex Ottenwess , Joseph R. JR Duffy , Rene L. Utianski , Leland R. Barnard , John L. Stricker , Daniela Wiepert , David T. Jones , Hugo Botha

PAD and FFD are proposed to protect face data from physical media-based Presentation Attacks and digital editing-based DeepFakes, respectively. However, isolated training of these two models significantly increases vulnerability towards…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Ajian Liu , Haocheng Yuan , Xiao Guo , Hui Ma , Wanyi Zhuang , Changtao Miao , Yan Hong , Chuanbiao Song , Jun Lan , Qi Chu , Tao Gong , Yanyan Liang , Weiqiang Wang , Jun Wan , Xiaoming Liu , Zhen Lei

We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline video understanding or text-prompted streaming QA,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Xudong Lu , Xueying Li , Annan Wang , Yang Bo , Jinpeng Chen , Zengliang Li , Nianzu Yang , Rui Liu , Xue Yang , Jingwen Hou , Hongsheng Li

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is the movement of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Ander Arriandiaga , Giovanni Morrone , Luca Pasa , Leonardo Badino , Chiara Bartolozzi

With the rapid advancement of smart glasses, voice interaction has been widely adopted due to its naturalness and convenience. However, its practical deployment is often undermined by vulnerability to spoofing attacks, while no public…

Human-Computer Interaction · Computer Science 2026-05-12 Weiye Xu , Zhang Jiang , Siqi Zheng , Xiyuxing Zhang , Changhao Zhang , Jian Liu , Weiqiang Wang , Yuntao Wang

Speech technologies are transforming interactions across various sectors, from healthcare to call centers and robots, yet their performance on African-accented conversations remains underexplored. We introduce Afrispeech-Dialog, a benchmark…

Self-supervised learning approaches have lately achieved great success on a broad spectrum of machine learning problems. In the field of speech processing, one of the most successful recent self-supervised models is wav2vec 2.0. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-10 Marie Kunešová , Zbyněk Zajíc

Current state-of-the-art open-vocabulary segmentation methods typically rely on image-mask-text triplet annotations for supervision. However, acquiring such detailed annotations is labour-intensive and poses scalability challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Zhaoqing Wang , Xiaobo Xia , Ziye Chen , Xiao He , Yandong Guo , Mingming Gong , Tongliang Liu

Large vision and language models show strong performance in tasks like image captioning, visual question answering, and retrieval. However, challenges remain in integrating speech, text, and vision into a unified model, especially for…

Multimedia · Computer Science 2025-07-08 Ngoc Dung Huynh , Mohamed Reda Bouadjenek , Imran Razzak , Hakim Hacid , Sunil Aryal

Voice activity detection (VAD) makes a distinction between speech and non-speech and its performance is of crucial importance for speech based services. Recently, deep neural network (DNN)-based VADs have achieved better performance than…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-14 Zhenpeng Zheng , Jianzong Wang , Ning Cheng , Jian Luo , Jing Xiao

Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely…

Sound · Computer Science 2024-03-18 René Groh , Nina Goes , Andreas M. Kist

Voice activity detection (VAD) improves the performance of speaker verification (SV) by preserving speech segments and attenuating the effects of non-speech. However, this scheme is not ideal: (1) it fails in noisy environments or…

Sound · Computer Science 2023-06-01 Zuheng Kang , Jianzong Wang , Junqing Peng , Jing Xiao
‹ Prev 1 8 9 10 Next ›