English
Related papers

Related papers: VANPY: Voice Analysis Framework

200 papers

The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting…

Sound · Computer Science 2025-03-25 Emma Coletta , Davide Salvi , Viola Negroni , Daniele Ugo Leonzio , Paolo Bestagini

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it…

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

Voice User Interfaces (VUIs) are increasingly popular and built into smartphones, home assistants, and Internet of Things (IoT) devices. Despite offering an always-on convenient user experience, VUIs raise new security and privacy concerns…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-07 Ranya Aloufi , Hamed Haddadi , David Boyle

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatically. In this work…

Sound · Computer Science 2022-11-01 Luigi Attorresi , Davide Salvi , Clara Borrelli , Paolo Bestagini , Stefano Tubaro

Achieving nuanced and accurate emulation of human voice has been a longstanding goal in artificial intelligence. Although significant progress has been made in recent years, the mainstream of speech synthesis models still relies on…

Sound · Computer Science 2024-03-04 Weiwei Lin , Chenhang He , Man-Wai Mak , Jiachen Lian , Kong Aik Lee

Many commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Shareef Babu Kalluri , Deepu Vijayasenan , Sriram Ganapathy , Ragesh Rajan M , Prashant Krishnan

Although speech recognition algorithms have developed quickly in recent years, achieving high transcription accuracy across diverse audio formats and acoustic environments remains a major challenge. This work explores how incorporating…

Sound · Computer Science 2025-03-31 Aniket Abhishek Soni

Voice anonymization systems aim to protect speaker privacy by obscuring vocal traits while preserving the linguistic content relevant for downstream applications. However, because these linguistic cues remain intact, they can be exploited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Ahmad Aloradi , Ünal Ege Gaznepoglu , Emanuël A. P. Habets , Daniel Tenbrinck

This paper investigates the temporal excitation patterns of creaky voice. Creaky voice is a voice quality frequently used as a phrase-boundary marker, but also as a means of portraying attitude, affective states and even social status.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-02 Thomas Drugman , John Kane , Christer Gobl

For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 Tianchi Liu , Kong Aik Lee , Qiongqiong Wang , Haizhou Li

Voice conversion (VC) has made progress in feature disentanglement, but it is still difficult to balance timbre and content information. This paper evaluates the pre-trained model features commonly used in voice conversion, and proposes an…

Sound · Computer Science 2025-04-09 Wenyu Wang , Yiquan Zhou , Jihua Zhu , Hongwu Ding , Jiacheng Xu , Shihao Li

This study explores the potential of using acoustic features of segmental speech sounds to detect deepfake audio. These features are highly interpretable because of their close relationship with human articulatory processes and are expected…

Sound · Computer Science 2025-12-12 Tianle Yang , Chengzhe Sun , Siwei Lyu , Phil Rose

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

Sound · Computer Science 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Siyin Wang , Wenyi Yu , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Lu Lu , Yu Tsao , Junichi Yamagishi , Yuxuan Wang , Chao Zhang

We investigate the feasibility of a singing voice synthesis (SVS) system by using a decomposed framework to improve flexibility in generating singing voices. Due to data-driven approaches, SVS performs a music score-to-waveform mapping;…

Sound · Computer Science 2024-07-15 Lester Phillip Violeta , Taketo Akama

This paper presents a software allowing to describe voices using a continuous Voice Femininity Percentage (VFP). This system is intended for transgender speakers during their voice transition and for voice therapists supporting them in this…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-24 David Doukhan , Simon Devauchelle , Lucile Girard-Monneron , Mía Chávez Ruz , V. Chaddouk , Isabelle Wagner , Albert Rilliard

In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper…

Sound · Computer Science 2024-08-29 Yiyang Zhao , Shuai Wang , Guangzhi Sun , Zehua Chen , Chao Zhang , Mingxing Xu , Thomas Fang Zheng

Story understanding and analysis have long been challenging areas within Natural Language Understanding. Automated narrative analysis requires deep computational semantic representations along with syntactic processing. Moreover, the large…

Computation and Language · Computer Science 2025-11-18 Taimur Khan , Ramoza Ahsan , Mohib Hameed

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-26 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Junbo Zhang , Jian Luan