English
Related papers

Related papers: Helsinki Speech Challenge 2024

200 papers

While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training…

The increasing realism of synthetic speech, driven by advancements in text-to-speech models, raises ethical concerns regarding impersonation and disinformation. Audio watermarking offers a promising solution via embedding…

Machine Learning · Computer Science 2024-11-14 Hongbin Liu , Moyang Guo , Zhengyuan Jiang , Lun Wang , Neil Zhenqiang Gong

This work presents a novel deep-learning-based pipeline for the inverse problem of image deblurring, leveraging augmentation and pre-training with synthetic data. Our results build on our winning submission to the recent Helsinki Deblur…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Theophil Trippe , Martin Genzel , Jan Macdonald , Maximilian März

This paper presents the submission of the S4 team to the Singing Voice Conversion Challenge 2025 (SVCC2025)-a novel singing style conversion system that advances fine-grained style conversion and control within in-domain settings. To…

Sound · Computer Science 2026-04-08 Zhetao Hu , Yiquan Zhou , Wenyu Wang , Zhiyu Wu , Xin Gao , Jihua Zhu

Conversational AI is constrained in many real-world settings where only one side of a dialogue can be recorded, such as telemedicine, call centers, and smart glasses. We formalize this as the one-sided conversation problem (1SC): inferring…

Computation and Language · Computer Science 2026-04-20 Victoria Ebert , Rishabh Singh , Tuochao Chen , Noah A. Smith , Shyamnath Gollakota

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be…

Sound · Computer Science 2026-02-06 Xueping Zhang , Han Yin , Yang Xiao , Lin Zhang , Ting Dang , Rohan Kumar Das , Ming Li

Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content,…

We present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-labels obtained by clustering the outputs of a CPC network…

The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-12 Jiangyan Yi , Chu Yuan Zhang , Jianhua Tao , Chenglong Wang , Xinrui Yan , Yong Ren , Hao Gu , Junzuo Zhou

We present VoiceRestore, a novel approach to restoring the quality of speech recordings using flow-matching Transformers trained in a self-supervised manner on synthetic data. Our method tackles a wide range of degradations frequently found…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Stanislav Kirdey

Noise suppression (NS) algorithms are effective in improving speech quality in many cases. However, aggressive noise suppression can damage the target speech, reducing both speech intelligibility and quality despite removing the noise. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-11 Kyungguen Byun , Jason Filos , Erik Visser , Sunkuk Moon

For a speech-enhancement algorithm, it is highly desirable to simultaneously improve perceptual quality and recognition rate. Thanks to computational costs and model complexities, it is challenging to train a model that effectively…

Machine Learning · Computer Science 2018-02-19 Rasool Fakoor , Xiaodong He , Ivan Tashev , Shuayb Zarar

This paper presents the multi-speaker multi-lingual few-shot voice cloning system developed by THU-HCSI team for LIMMITS'24 Challenge. To achieve high speaker similarity and naturalness in both mono-lingual and cross-lingual scenarios, we…

Sound · Computer Science 2024-04-26 Yixuan Zhou , Shuoyi Zhou , Shun Lei , Zhiyong Wu , Menglin Wu

The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation…

Detecting partial deepfake speech is challenging because manipulations occur only in short regions while the surrounding audio remains authentic. However, existing detection methods are fundamentally limited by the quality of available…

Sound · Computer Science 2025-12-16 Menglu Li , Majd Alber , Ramtin Asgarianamiri , Lian Zhao , Xiao-Ping Zhang

High-quality speech corpora are essential foundations for most speech applications. However, such speech data are expensive and limited since they are collected in professional recording environments. In this work, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Haoyu Li , Yang Ai , Junichi Yamagishi

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

The automatic speaker verification spoofing and countermeasures (ASVspoof) challenge series is a community-led initiative which aims to promote the consideration of spoofing and the development of countermeasures. ASVspoof 2021 is the 4th…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Héctor Delgado , Nicholas Evans , Tomi Kinnunen , Kong Aik Lee , Xuechen Liu , Andreas Nautsch , Jose Patino , Md Sahidullah , Massimiliano Todisco , Xin Wang , Junichi Yamagishi

We present a system for non-intrusive prediction of speech quality in noisy and enhanced speech, developed for Track 3 of the VoiceMOS 2024 Challenge. The task required estimating the ITU-T P.835 metrics SIG, BAK, and OVRL without reference…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Marie Kunešová , Aleš Pražák , Jan Lehečka

We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker recognition technology is able to diarise and recognize…

‹ Prev 1 4 5 6 7 8 10 Next ›