中文
相关论文

相关论文: PRODIS -- a speech database and a phoneme-based la…

200 篇论文

A judicious combination of dictionary learning methods, block sparsity and source recovery algorithm are used in a hierarchical manner to identify the noises and the speakers from a noisy conversation between two people. Conversations are…

声音 · 计算机科学 2016-10-31 K V Vijay Girish , A G Ramakrishnan , T V Ananthapadmanabha

We present a probabilistic model that uses both prosodic and lexical cues for the automatic segmentation of speech into topically coherent units. We propose two methods for combining lexical and prosodic information using hidden Markov…

计算与语言 · 计算机科学 2022-02-28 G. Tur , D. Hakkani-Tur , A. Stolcke , E. Shriberg

This paper investigates the use of word surprisal, a measure of the predictability of a word in a given context, as a feature to aid speech synthesis prosody. We explore how word surprisal extracted from large language models (LLMs)…

音频与语音处理 · 电气工程与系统科学 2023-06-19 Sofoklis Kakouros , Juraj Šimko , Martti Vainio , Antti Suni

The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the…

声音 · 计算机科学 2022-05-31 Syeda Rabia Arshad , Syed Mujtaba Haider , Abdul Basit Mughal

The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. We present forePLay, a novel Polish language dataset for…

计算与语言 · 计算机科学 2025-06-04 Anna Kołos , Katarzyna Lorenc , Emilia Wiśnios , Agnieszka Karlińska

This paper explores the manipulation of prosodic parameters in Text-to-Speech (TTS) systems to achieve controlled speech generation. By leveraging advanced speech processing techniques, we compare TTS-generated audio with human-recorded…

声音 · 计算机科学 2024-10-03 Podakanti Satyajith Chary

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

声音 · 计算机科学 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

This research explores the effects of various training settings on a Polish to English Statistical Machine Translation system for spoken language. Various elements of the TED, Europarl, and OPUS parallel text corpora were used as the basis…

计算与语言 · 计算机科学 2015-10-02 Krzysztof Wołk

Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming. In…

声音 · 计算机科学 2022-06-17 Ziqian Dai , Jianwei Yu , Yan Wang , Nuo Chen , Yanyao Bian , Guangzhi Li , Deng Cai , Dong Yu

Recent neural speech synthesis systems have gradually focused on the control of prosody to improve the quality of synthesized speech, but they rarely consider the variability of prosody and the correlation between prosody and semantics…

音频与语音处理 · 电气工程与系统科学 2020-08-14 Zhen Zeng , Jianzong Wang , Ning Cheng , Jing Xiao

This paper presents a study on mutual speech variation influences in a human-computer setting. The study highlights behavioral patterns in data collected as part of a shadowing experiment, and is performed using a novel end-to-end platform…

人机交互 · 计算机科学 2018-09-14 Eran Raveh , Ingmar Steiner , Iona Gessinger , Bernd Möbius

We introduce LAMBADA, a dataset to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative passages sharing the characteristic that human subjects are…

Although text-to-speech (TTS) systems have significantly improved, most TTS systems still have limitations in synthesizing speech with appropriate phrasing. For natural speech synthesis, it is important to synthesize the speech with a…

音频与语音处理 · 电气工程与系统科学 2023-06-14 Ji-Sang Hwang , Sang-Hoon Lee , Seong-Whan Lee

The main aim of translation is an accurate transfer of meaning so that the result is not only grammatically and lexically correct but also communicatively adequate. This paper stresses the need for discourse analysis the aim of which is to…

cmp-lg · 计算机科学 2008-02-03 Malgorzata E. Stys , Stefan S. Zemke

This study explores prosodic production in latent aphasia, a mild form of aphasia associated with left-hemisphere brain damage (e.g. stroke). Unlike prior research on moderate to severe aphasia, we investigated latent aphasia, which can…

神经元与认知 · 定量生物学 2024-09-25 Cong Zhang , Tong Li , Gayle DeDe , Christos Salis

A lexical database tool tailored for phonological research is described. Database fields include transcriptions, glosses and hyperlinks to speech files. Database queries are expressed using HTML forms, and these permit regular expression…

cmp-lg · 计算机科学 2008-02-03 Steven Bird

Recent work shows promising results in expanding the capabilities of large language models (LLM) to directly understand and synthesize speech. However, an LLM-based strategy for modeling spoken dialogs remains elusive, calling for further…

Spoken communication occurs in a "noisy channel" characterized by high levels of environmental noise, variability within and between speakers, and lexical and syntactic ambiguity. Given these properties of the received linguistic input,…

计算与语言 · 计算机科学 2021-01-26 Stephan C. Meylan , Sathvik Nair , Thomas L. Griffiths

In this paper we introduce a new natural language processing dataset and benchmark for predicting prosodic prominence from written text. To our knowledge this will be the largest publicly available dataset with prosodic labels. We describe…

计算与语言 · 计算机科学 2019-08-07 Aarne Talman , Antti Suni , Hande Celikkanat , Sofoklis Kakouros , Jörg Tiedemann , Martti Vainio

Speech and language technologies offer valuable opportunities for supporting mental health assessment through objective and interpretable cues. We present a systematic feature-based analysis framework leveraging perceptually grounded…

人工智能 · 计算机科学 2026-05-28 Vassilis Lyberatos , Edmund G. Dervakos , Eleni Adamidi , Athanasios Voulodimos , Giorgos Stamou