English
Related papers

Related papers: KunquDB: An Attempt for Speaker Verification in th…

200 papers

Researches indicate that text-dependent speaker verification (TD-SV) often outperforms text-independent verification (TI-SV) in short speech scenarios. However, collecting large-scale fixed text speech data is challenging, and as speech…

Sound · Computer Science 2023-12-05 Litong Zheng , Feng Hong , Weijie Xu

The Chinese numerical string corpus, serves as a valuable resource for speaker verification, particularly in financial transactions. Researches indicate that in short speech scenarios, text-dependent speaker verification (TD-SV)…

Sound · Computer Science 2024-05-22 Litong Zheng , Feng Hong , Weijie Xu , Wan Zheng

The data-driven computational research on automatic jingju (also known as Beijing or Peking opera) singing evaluation lacks a suitable and comprehensive a cappella singing audio dataset. In this work, we present an a cappella singing audio…

Sound · Computer Science 2017-08-15 Rong Gong , Rafael Caro Repetto , Xavier Serra

The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English…

Peking Opera has been the most dominant form of Chinese performing art since around 200 years ago. A Peking Opera singer usually exhibits a very strong personal style via introducing improvisation and expressiveness on stage which leads the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Yusong Wu , Shengchen Li , Chengzhu Yu , Heng Lu , Chao Weng , Liqiang Zhang , Dong Yu

Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both generate and understand audio. However, preserving key…

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations.…

Computation and Language · Computer Science 2025-11-13 Yang Chen , Hui Wang , Shiyao Wang , Junyang Chen , Jiabei He , Jiaming Zhou , Xi Yang , Yequan Wang , Yonghua Lin , Yong Qin

This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as the training data, which relies on techniques and is not able…

Computation and Language · Computer Science 2019-12-30 Yusong Wu , Shengchen Li , Chengzhu Yu , Heng Lu , Chao Weng , Liqiang Zhang , Dong Yu

Recently, instruction-following audio-language models have received broad attention for audio interaction with humans. However, the absence of pre-trained audio models capable of handling diverse audio types and tasks has hindered progress…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Yunfei Chu , Jin Xu , Xiaohuan Zhou , Qian Yang , Shiliang Zhang , Zhijie Yan , Chang Zhou , Jingren Zhou

The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consist of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese…

Computation and Language · Computer Science 2020-04-09 Hao Zhou , Chujie Zheng , Kaili Huang , Minlie Huang , Xiaoyan Zhu

Dialectal Arabic (DA) speech data vary widely in domain coverage, dialect labeling practices, and recording conditions, complicating cross-dataset comparison and model evaluation. To characterize this landscape, we conduct a computational…

Computation and Language · Computer Science 2026-01-30 Peter Sullivan , AbdelRahim Elmadany , Alcides Alcoba Inciarte , Muhammad Abdul-Mageed

Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets that pair spoken sentences with richly ambiguous text. To…

Computation and Language · Computer Science 2025-06-10 Haotian Guo , Jing Han , Yongfeng Tu , Shihao Gao , Shengfan Shen , Wulong Xiang , Weihao Gan , Zixing Zhang

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth…

Sound · Computer Science 2023-06-06 Jianrong Wang , Yuchen Huo , Li Liu , Tianyi Xu , Qi Li , Sen Li

This article describes a system for analyzing acoustic data to assist in the diagnosis and classification of children's speech sound disorders (SSDs) using a computer. The analysis concentrated on identifying and categorizing four distinct…

Sound · Computer Science 2022-07-07 Yao-Ming Kuo , Shanq-Jang Ruan , Yu-Chin Chen , Ya-Wen Tu

This paper introduces a novel multi-Agent framework that automates the end to end production of Qinqiang opera by integrating Large Language Models , visual generation, and Text to Speech synthesis. Three specialized agents collaborate in…

Artificial Intelligence · Computer Science 2025-04-23 Gengxian Cao , Fengyuan Li , Hong Duan , Ye Yang , Bofeng Wang , Donghe Li

Humans naturally attribute utterances of direct speech to their speaker in literary works. When attributing quotes, we process contextual information but also access mental representations of characters that we build and revise throughout…

Computation and Language · Computer Science 2025-01-24 Gaspard Michel , Elena V. Epure , Romain Hennequin , Christophe Cerisara

Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental…

We introduce LibriConvo, a simulated multi-speaker conversational dataset based on speaker-aware conversation simulation (SASC), designed to support training and evaluation of speaker diarization and automatic speech recognition (ASR)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Máté Gedeon , Péter Mihajlik

With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studies (CCS). While existing research primarily focuses on text and visual modalities, the audio…

Computation and Language · Computer Science 2026-04-14 Yexing Du , Kaiyuan Liu , Bihe Zhang , Youcheng Pan , Bo Yang , Liangyu Huo , Xiyuan Zhang , Jian Xie , Daojing He , Yang Xiang , Ming Liu , Bing Qin

In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have…

‹ Prev 1 2 3 10 Next ›