中文
相关论文

相关论文: VoxSRC 2021: The Third VoxCeleb Speaker Recognitio…

200 篇论文

In the field of speaker diarization, the development of technology is constrained by two problems: insufficient data resources and poor generalization ability of deep learning models. To address these two problems, firstly, we propose an…

音频与语音处理 · 电气工程与系统科学 2025-07-01 Shilong Wu

Challenges based on Computational Paralinguistics in the INTERSPEECH Conference have always had a good reception among the attendees owing to its competitive academic and research demands. This year, the INTERSPEECH 2020 Computational…

音频与语音处理 · 电气工程与系统科学 2020-08-25 José Vicente Egas-López

Recent development of speech processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most…

声音 · 计算机科学 2022-02-28 Fan Yu , Shiliang Zhang , Yihui Fu , Lei Xie , Siqi Zheng , Zhihao Du , Weilong Huang , Pengcheng Guo , Zhijie Yan , Bin Ma , Xin Xu , Hui Bu

We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of…

Speech emotion recognition (SER), particularly for naturally expressed emotions, remains a challenging computational task. Key challenges include the inherent subjectivity in emotion annotation and the imbalanced distribution of emotion…

声音 · 计算机科学 2025-06-03 Tiantian Feng , Thanathai Lertpetchpun , Dani Byrd , Shrikanth Narayanan

We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select…

音频与语音处理 · 电气工程与系统科学 2026-05-01 Doyeop Kwak , Jeongsoo Choi , Suyeon Lee , Joon Son Chung

Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab…

Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we…

音频与语音处理 · 电气工程与系统科学 2021-04-06 Zhongxin Bai , Xiao-Lei Zhang

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech…

音频与语音处理 · 电气工程与系统科学 2022-02-07 Naijun Zheng , Na Li , Xixin Wu , Lingwei Meng , Jiawen Kang , Haibin Wu , Chao Weng , Dan Su , Helen Meng

Virtual assistants are becoming increasingly important speech-driven Information Retrieval platforms that assist users with various tasks. We discuss open problems and challenges with respect to modeling spoken information queries for…

信息检索 · 计算机科学 2023-04-27 Christophe Van Gysel

Speaker diarization is a task to label an audio or video recording with the identity of the speaker at each given time stamp. In this work, we propose a novel machine learning framework to conduct real-time multi-speaker diarization and…

声音 · 计算机科学 2023-02-23 Baihan Lin , Xinxin Zhang

In this paper, we describe the top-scoring submissions for team RTZR VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22) in the closed dataset, speaker verification Track 1. The top performed system is a fusion of 7 models, which…

音频与语音处理 · 电气工程与系统科学 2022-09-22 Sangwon Suh , Sunjong Park

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

计算机视觉与模式识别 · 计算机科学 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

This paper proposes an improved approach for open-set speaker identification based on pretrained speaker foundation models. Building upon the previous Speaker Reciprocal Points Learning framework (V1), we first introduce an enhanced…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Zhiyong Chen , Shuhang Wu , Yingjie Duan , Xinkang Xu , Xinhui Hu

The Multitarget Challenge aims to assess how well current speech technology is able to determine whether or not a recorded utterance was spoken by one of a large number of 'blacklisted' speakers. It is a form of multi-target speaker…

音频与语音处理 · 电气工程与系统科学 2018-07-19 Suwon Shon , Najim Dehak , Douglas Reynolds , James Glass

Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Ruijie Tao , Rohan Kumar Das , Haizhou Li

We present the first edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthetic speech. This challenge drew 22 participating teams from academia…

声音 · 计算机科学 2022-07-05 Wen-Chin Huang , Erica Cooper , Yu Tsao , Hsin-Min Wang , Tomoki Toda , Junichi Yamagishi

Despite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Ju-ho Kim , Hye-jin Shim , Jungwoo Heo , Ha-Jin Yu

We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidevice home environments to provide context for the INTERSPEECH…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Gregory Ciccarelli , Jarred Barber , Arun Nair , Israel Cohen , Tao Zhang

Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-world applications, such assumptions often do not hold.…