中文
相关论文

相关论文: Analyzing phonetic structure of Mandarin using Aud…

200 篇论文

Speech communication systems based on Voice-over-IP technology are frequently used by native as well as non-native speakers of a target language, e.g. in international phone calls or telemeetings. Frequently, such calls also occur in a…

多媒体 · 计算机科学 2020-10-27 Babak Naderi , Gabriel Mittag , Rafael Zequeira Jim\a'enez , Sebastian Möller

Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.In this paper, we aim to fill in the gap and design a…

声音 · 计算机科学 2023-07-19 Haoxin Ma , Jiangyan Yi , Chenglong Wang , Xinrui Yan , Jianhua Tao , Tao Wang , Shiming Wang , Ruibo Fu

In Mandarin, the tonal contours of monosyllabic words produced in isolation or in careful speech are characterized by four lexical tones: a high-level tone (T1), a rising tone (T2), a dipping tone (T3) and a falling tone (T4). However, in…

计算与语言 · 计算机科学 2024-10-22 Xiaoyun Jin , Mirjam Ernestus , R. Harald Baayen

The choice of modeling units is crucial for automatic speech recognition (ASR) tasks. In mandarin scenarios, the Chinese characters represent meaning but are not directly related to the pronunciation. Thus only considering the writing of…

计算与语言 · 计算机科学 2022-10-19 Yuting Yang , Binbin Du , Yuke Li

The Mandarin Chinese language is known to be strongly influenced by a rich set of regional accents, while Mandarin speech with each accent is quite low resource. Hence, an important task in Mandarin speech recognition is to appropriately…

音频与语音处理 · 电气工程与系统科学 2024-06-17 Xurong Xie , Xiang Sui , Xunying Liu , Lan Wang

The development of multi-modal large language models (LLMs) leads to intelligent approaches capable of speech interactions. As one of the most widely spoken languages globally, Mandarin is supported by most models to enhance their…

计算与语言 · 计算机科学 2025-11-18 Heyang Liu , Ziyang Cheng , Yuhao Wang , Hongcheng Liu , Yiqi Li , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

As the structure of Chinese characters are very different, it is very difficult to input Chinese characters into computer quickly and conveniently. The conventional keyboard does not support the pictorial characters in Chinese language.…

人机交互 · 计算机科学 2013-10-14 Umakant Mishra

Services of personalized TTS systems for the Mandarin-speaking speech impaired are rarely mentioned. Taiwan started the VoiceBanking project in 2020, aiming to build a complete set of services to deliver personalized Mandarin TTS systems to…

音频与语音处理 · 电气工程与系统科学 2023-08-30 Jia-Jyu Su , Pang-Chen Liao , Yen-Ting Lin , Wu-Hao Li , Guan-Ting Liou , Cheng-Che Kao , Wei-Cheng Chen , Jen-Chieh Chiang , Wen-Yang Chang , Pin-Han Lin , Chen-Yu Chiang

Speaker anonymization aims to protect speaker identity while preserving content information and the intelligibility of speech. However, most speaker anonymization systems (SASs) are developed and evaluated using only English, resulting in…

音频与语音处理 · 电气工程与系统科学 2025-07-02 Zhe Zhang , Wen-Chin Huang , Xin Wang , Xiaoxiao Miao , Junichi Yamagishi

This study probes the phonetic and phonological knowledge of lexical tones in TTS models through two experiments. Controlled stimuli for testing tonal coarticulation and tone sandhi in Mandarin were fed into Tacotron 2 and WaveGlow to…

计算与语言 · 计算机科学 2019-12-24 Jian Zhu

This paper proposes a framework for evaluating large language models (LLMs) on Chinese topic constructions, focusing on their sensitivity to island constraints. Drawing inspiration from Tian et al. (2024), we outline an experimental design…

计算与语言 · 计算机科学 2025-04-22 Xiaodong Yang

The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely employed as speech tokenizers, the impact of frame rates on…

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech…

He and Rao (2013) reported a raising phenomenon of /a/ in /Xan/ (X being a consonant or a vowel) in Chengdu dialect of Mandarin, i.e. /a/ is realized as [epsilon] for young speakers but [ae] for older speakers, but they offered no acoustic…

计算与语言 · 计算机科学 2018-03-13 Hai Hu , Yiwen Zhang

Singing accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental…

Transliteration is a task of translating named entities from a language to another, based on phonetic similarity. The task has embraced deep learning approaches in recent years, yet, most ignore the phonetic features of the involved…

计算与语言 · 计算机科学 2022-01-24 Shi Cheng , Zhuofei Ding , Songpeng Yan

Speech recognition in mixed language has difficulties to adapt end-to-end framework due to the lack of data and overlapping phone sets, for example in words such as "one" in English and "w\`an" in Chinese. We propose a CTC-based end-to-end…

计算与语言 · 计算机科学 2018-10-31 Genta Indra Winata , Andrea Madotto , Chien-Sheng Wu , Pascale Fung

This paper analyzes dysarthric speech datasets from three languages with different prosodic systems: English, Korean, and Tamil. We inspect 39 acoustic measurements which reflect three speech dimensions including voice quality,…

计算与语言 · 计算机科学 2022-11-04 Eun Jung Yeo , Sunhee Kim , Minhwa Chung

Code-switching (CS) is a common phenomenon and recognizing CS speech is challenging. But CS speech data is scarce and there' s no common testbed in relevant research. This paper describes the design and main outcomes of the ASRU 2019…

音频与语音处理 · 电气工程与系统科学 2020-07-14 Xian Shi , Qiangze Feng , Lei Xie

This paper presents a multi-label stuttering detection system trained on multi-corpus, multilingual data in English, German, and Mandarin.By leveraging annotated stuttering data from three languages and four corpora, the model captures…

声音 · 计算机科学 2026-03-31 Felix Haas , Sebastian P. Bayerl