中文
相关论文

相关论文: Balalaika: Data-Centric, Prosody-Aware Annotation …

200 篇论文

Recent advances in Text-To-Speech (TTS) synthesis have achieved near-human speech quality in neutral speaking styles. However, most existing approaches either depend on costly emotion annotations or optimize surrogate objectives that fail…

计算与语言 · 计算机科学 2026-04-08 Qing Yang , Zhenghao Liu , Yangfan Du , Pengcheng Huang , Tong Xiao

Most End-to-End SLU methods depend on the pretrained ASR or language model features for intent prediction. However, other essential information in speech, such as prosody, is often ignored. Recent research has shown improved results in…

计算与语言 · 计算机科学 2023-05-16 Shangeth Rajaa

Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional…

声音 · 计算机科学 2026-01-14 Runchuan Ye , Yixuan Zhou , Renjie Yu , Zijian Lin , Kehan Li , Xiang Li , Xin Liu , Guoyang Zeng , Zhiyong Wu

We propose Text-Aligned Speech Tokens with Multiple Layer-Aggregation (TASLA), which is a text-aligned speech tokenization framework that aims to address the problem that under a low-frame-rate and text-aligned regime, single-source speech…

声音 · 计算机科学 2025-10-17 Ming-Hao Hsu , Liang-Hsuan Tseng , Hung-yi Lee , Zhizheng Wu

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

声音 · 计算机科学 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

While modern Automatic Speech Recognition (ASR) systems achieve high accuracy on benchmark corpora, their performance often degrades when there is real-world variability. This work focuses on variability arising due to accented,…

计算与语言 · 计算机科学 2026-05-19 Sicheng Jin , Dipankar Srirag , Aditya Joshi

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The second goal is to explain…

音频与语音处理 · 电气工程与系统科学 2019-06-26 Gabriel Meseguer-Brocal , Alice Cohen-Hadria , Geoffroy Peeters

Goal-oriented dialog systems enable users to complete specific goals like requesting information about a movie or booking a ticket. Typically the dialog system pipeline contains multiple ML models, including natural language understanding,…

Prosody plays a crucial role in speech perception, influencing both human understanding and automatic speech recognition (ASR) systems. Despite its importance, prosodic stress remains under-studied due to the challenge of efficiently…

声音 · 计算机科学 2025-03-06 Samuel S. Sohn , Sten Knutsen , Karin Stromswold

The availability of prosodic information from speech signals is useful in a wide range of applications. However, deriving this information from speech signals can be a laborious task involving manual intervention. Therefore, the current…

Individuals with cerebral palsy (CP) and amyotrophic lateral sclerosis (ALS) frequently face challenges with articulation, leading to dysarthria and resulting in atypical speech patterns. In healthcare settings, communication breakdowns…

计算与语言 · 计算机科学 2024-11-11 Macarious Hui , Jinda Zhang , Aanchan Mohan

In this paper, a novel approach is proposed to automatically construct parallel discourse corpus for dialogue machine translation. Firstly, the parallel subtitle data and its corresponding monolingual movie script data are crawled and…

计算与语言 · 计算机科学 2016-05-24 Longyue Wang , Xiaojun Zhang , Zhaopeng Tu , Andy Way , Qun Liu

The design of asynchronous circuits typically requires a judicious definition of signals and modules, combined with a proper specification of their timing constraints, which can be a complex and error-prone process, using standard Hardware…

硬件体系结构 · 计算机科学 2023-08-09 Carsten Nielsen , Zhe Su , Giacomo Indiveri

In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, discarding prosodic…

人机交互 · 计算机科学 2026-03-11 SangYeop Jeong , Yeongseo Na , Seung Gyu Jeong , Jin-Woo Jeong , Seong-Eun Kim

This paper introduces a novel neural network-based speech coding system that can process noisy speech effectively. The proposed source-aware neural audio coding (SANAC) system harmonizes a deep autoencoder-based source separation model and…

音频与语音处理 · 电气工程与系统科学 2020-11-11 Haici Yang , Kai Zhen , Seungkwon Beack , Minje Kim

Automatic Speech Recognition and Text-to-Speech systems are primarily trained in a supervised fashion and require high-quality, accurately labeled speech datasets. In this work, we examine common problems with speech data and introduce a…

音频与语音处理 · 电气工程与系统科学 2022-01-10 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg

While deep learning-based text-to-speech (TTS) models such as VITS have shown excellent results, they typically require a sizable set of high-quality <text, audio> pairs to train, which is expensive to collect. So far, most languages in the…

声音 · 计算机科学 2023-01-05 Xin Yuan , Robin Feng , Mingming Ye

Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of spoken data, the…

音频与语音处理 · 电气工程与系统科学 2025-03-17 Nikola Ljubešić , Peter Rupnik , Danijel Koržinek

Identifying whether an utterance is a statement, question, greeting, and so forth is integral to effective automatic understanding of natural dialog. Little is known, however, about how such dialog acts (DAs) can be automatically classified…

计算与语言 · 计算机科学 2007-05-23 E. Shriberg , R. Bates , A. Stolcke , P. Taylor , D. Jurafsky , K. Ries , N. Coccaro , R. Martin , M. Meteer , C. Van Ess-Dykema

We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the sound recordings of the Czech parliamentary speeches with…

计算与语言 · 计算机科学 2025-09-09 Vladislav Stankov , Matyáš Kopp , Ondřej Bojar