中文
相关论文

相关论文: LOCATA challenge: speaker localization with a plan…

200 篇论文

Conventional approaches to sound localization and separation are based on microphone arrays in artificial systems. Inspired by the selective perception of human auditory system, we design a multi-source listening system which can separate…

声音 · 计算机科学 2019-11-11 Xuecong Sun , Han Jia , Zhe Zhang , Yuzhen Yang , Zhaoyong Sun , Jun Yang

Robustly and accurately localizing objects in real-world environments can be challenging due to noisy data, hardware limitations, and the inherent randomness of physical systems. To account for these factors, existing works estimate the…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Moussa Kassem Sbeyti , Michelle Karg , Christian Wirth , Azarm Nowzad , Sahin Albayrak

Phrase localization is a task that studies the mapping from textual phrases to regions of an image. Given difficulties in annotating phrase-to-object datasets at scale, we develop a Multimodal Alignment Framework (MAF) to leverage more…

计算与语言 · 计算机科学 2020-10-13 Qinxin Wang , Hao Tan , Sheng Shen , Michael W. Mahoney , Zhewei Yao

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without…

计算机视觉与模式识别 · 计算机科学 2021-12-23 Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , Ji-Rong Wen

This paper introduces the third DIHARD challenge, the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions, and conversational…

音频与语音处理 · 电气工程与系统科学 2020-12-04 Neville Ryant , Kenneth Church , Christopher Cieri , Jun Du , Sriram Ganapathy , Mark Liberman

Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolution with an enhanced accuracy, a multi-scale approach is a way…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Tae Jin Park , Nithin Rao Koluguri , Jagadeesh Balam , Boris Ginsburg

Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA)…

声音 · 计算机科学 2026-01-27 Zexia Fan , Yu Chen , Qiquan Zhang , Kainan Chen , Xinyuan Qian

Direction-of-arrival estimation of multiple speakers in a room is an important task for a wide range of applications. In particular, challenging environments with moving speakers, reverberation and noise, lead to significant performance…

音频与语音处理 · 电气工程与系统科学 2024-09-24 Daniel A. Mitchell , Boaz Rafaely , Anurag Kumar , Vladimir Tourbabin

Re-localizing a camera from a single image in a previously mapped area is vital for many computer vision applications in robotics and augmented/virtual reality. In this work, we address the problem of estimating the 6 DoF camera pose…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Mohammad Altillawi , Zador Pataki , Shile Li , Ziyuan Liu

In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a…

计算机视觉与模式识别 · 计算机科学 2023-03-29 Fiona Ryan , Hao Jiang , Abhinav Shukla , James M. Rehg , Vamsi Krishna Ithapu

The objective of this work is to develop a speaker recognition model to be used in diverse scenarios. We hypothesise that two components should be adequately configured to build such a model. First, adequate architecture would be required.…

Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from other students talking…

音频与语音处理 · 电气工程与系统科学 2022-07-05 Antonio Gomez

Robust robot localization is an important prerequisite for navigation, but it becomes challenging when the map and robot measurements are obtained from different sensors. Prior methods are often tailored to specific environments, relying on…

机器人学 · 计算机科学 2026-04-03 Evgenii Kruzhkov , Raphael Memmesheimer , Sven Behnke

When human listeners try to guess the spatial position of a speech source, they are influenced by the speaker's production level, regardless of the intensity level reaching their ears. Because the perception of distance is a very difficult…

机器人学 · 计算机科学 2023-05-23 Ambre Davat , Véronique Aubergé , Gang Feng

Attractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the case where the number…

音频与语音处理 · 电气工程与系统科学 2021-09-24 Shota Horiguchi , Shinji Watanabe , Paola Garcia , Yawen Xue , Yuki Takashima , Yohei Kawaguchi

Nowadays, users are encouraged to activate across multiple online social networks simultaneously. Anchor link prediction, which aims to reveal the correspondence among different accounts of the same user across networks, has been regarded…

社会与信息网络 · 计算机科学 2021-08-26 Jiangli Shao , Yongqing Wang , Hao Gao , Huawei Shen , Yangyang Li , Xueqi Cheng

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research…

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Tinglong Zhu , Rongjie Huang , Chong Deng , Qian Chen , Shiliang Zhang , Wen Wang , Xihao Li

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Joseph Heyward , Nikhil Parthasarathy , Tyler Zhu , Aravindh Mahendran , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

The INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Xiaoyi Qin , Ming Li , Hui Bu , Wei Rao , Rohan Kumar Das , Shrikanth Narayanan , Haizhou Li