English
Related papers

Related papers: Augmenting Polish Automatic Speech Recognition Sys…

200 papers

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interlocutor which is known as entrainment. To produce human-like…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-22 Jian Cong , Shan Yang , Na Hu , Guangzhi Li , Lei Xie , Dan Su

Speaker verification systems are vulnerable to spoofing attacks which presents a major problem in their real-life deployment. To date, most of the proposed synthetic speech detectors (SSDs) have weighted the importance of different segments…

Sound · Computer Science 2016-10-11 Ali Khodabakhsh , Cenk Demiroglu

Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Many of those systems…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-23 Ünal Ege Gaznepoglu , Nils Peters

This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST)…

Computation and Language · Computer Science 2026-01-29 Zhaolin Li , Yining Liu , Danni Liu , Tuan Nam Nguyen , Enes Yavuz Ugan , Tu Anh Dinh , Carlos Mullov , Alexander Waibel , Jan Niehues

Collecting high-quality training data is essential for fine-tuning Large Language Models (LLMs). However, acquiring such data is often costly and time-consuming, especially for non-English languages such as Italian. Recently, researchers…

Computation and Language · Computer Science 2025-04-01 Fatemeh Mohammadi , Tommaso Romano , Samira Maghool , Paolo Ceravolo

In this study, we present recent developments on ESPnet: End-to-End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a…

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-11 Vikas Tokala , Eric Grinstein , Mike Brookes , Simon Doclo , Jesper Jensen , Patrick A. Naylor

Recent breakthroughs in multi-talker ASR (MT-ASR) and speaker diarization (SD) rely on synthetic data to mitigate the scarcity of large-scale conversational recordings, yet the impact of specific simulation choices remains poorly…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-18 Alexander Polok , Ivan Medennikov , Jan Černocký , Shinji Watanabe , Lukáš Burget , Samuele Cornell

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas,…

Sound · Computer Science 2025-07-18 Zhoulin Ji , Chenhao Lin , Hang Wang , Chao Shen

Self-supervised learning models have revolutionized the field of speech processing. However, the process of fine-tuning these models on downstream tasks requires substantial computational resources, particularly when dealing with multiple…

Computation and Language · Computer Science 2024-06-24 Varsha Suresh , Salah Aït-Mokhtar , Caroline Brun , Ioan Calapodescu

Many neural text-to-speech architectures can synthesize nearly natural speech from text inputs. These architectures must be trained with tens of hours of annotated and high-quality speech data. Compiling such large databases for every new…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Kishor Kayyar Lakshminarayana , Christian Dittmar , Nicola Pia , Emanuël Habets

Supervised training of an automated medical image analysis system often requires a large amount of expert annotations that are hard to collect. Moreover, the proportions of data available across different classes may be highly imbalanced…

Computer Vision and Pattern Recognition · Computer Science 2019-12-10 Yuan Xue , Jiarong Ye , Rodney Long , Sameer Antani , Zhiyun Xue , Xiaolei Huang

This paper describes our submitted systems to the ASVspoof 5 Challenge Track 1: Speech Deepfake Detection - Open Condition, which consists of a stand-alone speech deepfake (bonafide vs spoof) detection task. Recently, large-scale…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-25 Theophile Stourbe , Victor Miara , Theo Lepage , Reda Dehak

Cross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language. This is a challenging task, and resulting speech can suffer from reduced naturalness, accented speech, and/or loss…

Sound · Computer Science 2022-04-04 Marcel de Korte , Jaebok Kim , Aki Kunikoshi , Adaeze Adigwe , Esther Klabbers

Data augmentation methods for Natural Language Processing tasks are explored in recent years, however they are limited and it is hard to capture the diversity on sentence level. Besides, it is not always possible to perform data…

Computation and Language · Computer Science 2022-05-20 M. Şafak Bilici , Mehmet Fatih Amasyali

Collecting and annotating datasets for pixel-level semantic segmentation tasks are highly labor-intensive. Data augmentation provides a viable solution by enhancing model generalization without additional real-world data collection.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Huy Che , Dinh-Duy Phan , Duc-Khai Lam

This paper describes our RoyalFlush system for the track of multi-speaker automatic speech recognition (ASR) in the M2MeT challenge. We adopted the serialized output training (SOT) based multi-speakers ASR system with large-scale simulation…

Sound · Computer Science 2022-02-25 Shuaishuai Ye , Peiyao Wang , Shunfei Chen , Xinhui Hu , Xinkang Xu

The awareness for biased ASR datasets or models has increased notably in recent years. Even for English, despite a vast amount of available training data, systems perform worse for non-native speakers. In this work, we improve an…

Computation and Language · Computer Science 2023-03-03 Philipp Klumpp , Pooja Chitkara , Leda Sarı , Prashant Serai , Jilong Wu , Irina-Elena Veliche , Rongqing Huang , Qing He

In this paper, we present ConvoGen: an innovative framework for generating synthetic conversational data using multi-agent systems. Our method leverages few-shot learning and introduces iterative sampling from a dynamically updated few-shot…

Computation and Language · Computer Science 2025-05-12 Reem Gody , Mahmoud Goudy , Ahmed Y. Tawfik

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various…

Sound · Computer Science 2025-11-03 Jiarong Du , Zhan Jin , Peijun Yang , Juan Liu , Zhuo Li , Xin Liu , Ming Li