中文
相关论文

相关论文: KunquDB: An Attempt for Speaker Verification in th…

200 篇论文

Researches indicate that text-dependent speaker verification (TD-SV) often outperforms text-independent verification (TI-SV) in short speech scenarios. However, collecting large-scale fixed text speech data is challenging, and as speech…

声音 · 计算机科学 2023-12-05 Litong Zheng , Feng Hong , Weijie Xu

The Chinese numerical string corpus, serves as a valuable resource for speaker verification, particularly in financial transactions. Researches indicate that in short speech scenarios, text-dependent speaker verification (TD-SV)…

声音 · 计算机科学 2024-05-22 Litong Zheng , Feng Hong , Weijie Xu , Wan Zheng

The data-driven computational research on automatic jingju (also known as Beijing or Peking opera) singing evaluation lacks a suitable and comprehensive a cappella singing audio dataset. In this work, we present an a cappella singing audio…

声音 · 计算机科学 2017-08-15 Rong Gong , Rafael Caro Repetto , Xavier Serra

The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English…

Peking Opera has been the most dominant form of Chinese performing art since around 200 years ago. A Peking Opera singer usually exhibits a very strong personal style via introducing improvisation and expressiveness on stage which leads the…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Yusong Wu , Shengchen Li , Chengzhu Yu , Heng Lu , Chao Weng , Liqiang Zhang , Dong Yu

Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both generate and understand audio. However, preserving key…

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations.…

计算与语言 · 计算机科学 2025-11-13 Yang Chen , Hui Wang , Shiyao Wang , Junyang Chen , Jiabei He , Jiaming Zhou , Xi Yang , Yequan Wang , Yonghua Lin , Yong Qin

This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as the training data, which relies on techniques and is not able…

计算与语言 · 计算机科学 2019-12-30 Yusong Wu , Shengchen Li , Chengzhu Yu , Heng Lu , Chao Weng , Liqiang Zhang , Dong Yu

Recently, instruction-following audio-language models have received broad attention for audio interaction with humans. However, the absence of pre-trained audio models capable of handling diverse audio types and tasks has hindered progress…

音频与语音处理 · 电气工程与系统科学 2023-12-22 Yunfei Chu , Jin Xu , Xiaohuan Zhou , Qian Yang , Shiliang Zhang , Zhijie Yan , Chang Zhou , Jingren Zhou

The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consist of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese…

计算与语言 · 计算机科学 2020-04-09 Hao Zhou , Chujie Zheng , Kaili Huang , Minlie Huang , Xiaoyan Zhu

Dialectal Arabic (DA) speech data vary widely in domain coverage, dialect labeling practices, and recording conditions, complicating cross-dataset comparison and model evaluation. To characterize this landscape, we conduct a computational…

计算与语言 · 计算机科学 2026-01-30 Peter Sullivan , AbdelRahim Elmadany , Alcides Alcoba Inciarte , Muhammad Abdul-Mageed

Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets that pair spoken sentences with richly ambiguous text. To…

计算与语言 · 计算机科学 2025-06-10 Haotian Guo , Jing Han , Yongfeng Tu , Shihao Gao , Shengfan Shen , Wulong Xiang , Weihao Gan , Zixing Zhang

Audio-visual speech recognition (AVSR) gains increasing attention from researchers as an important part of human-computer interaction. However, the existing available Mandarin audio-visual datasets are limited and lack the depth…

声音 · 计算机科学 2023-06-06 Jianrong Wang , Yuchen Huo , Li Liu , Tianyi Xu , Qi Li , Sen Li

This article describes a system for analyzing acoustic data to assist in the diagnosis and classification of children's speech sound disorders (SSDs) using a computer. The analysis concentrated on identifying and categorizing four distinct…

声音 · 计算机科学 2022-07-07 Yao-Ming Kuo , Shanq-Jang Ruan , Yu-Chin Chen , Ya-Wen Tu

This paper introduces a novel multi-Agent framework that automates the end to end production of Qinqiang opera by integrating Large Language Models , visual generation, and Text to Speech synthesis. Three specialized agents collaborate in…

人工智能 · 计算机科学 2025-04-23 Gengxian Cao , Fengyuan Li , Hong Duan , Ye Yang , Bofeng Wang , Donghe Li

Humans naturally attribute utterances of direct speech to their speaker in literary works. When attributing quotes, we process contextual information but also access mental representations of characters that we build and revise throughout…

计算与语言 · 计算机科学 2025-01-24 Gaspard Michel , Elena V. Epure , Romain Hennequin , Christophe Cerisara

Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental…

We introduce LibriConvo, a simulated multi-speaker conversational dataset based on speaker-aware conversation simulation (SASC), designed to support training and evaluation of speaker diarization and automatic speech recognition (ASR)…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Máté Gedeon , Péter Mihajlik

With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studies (CCS). While existing research primarily focuses on text and visual modalities, the audio…

计算与语言 · 计算机科学 2026-04-14 Yexing Du , Kaiyuan Liu , Bihe Zhang , Youcheng Pan , Bo Yang , Liangyu Huo , Xiyuan Zhang , Jian Xie , Daojing He , Yang Xiang , Ming Liu , Bing Qin

In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have…

‹ 上一页 1 2 3 10 下一页 ›