中文
相关论文

相关论文: Quantifying and Correlating Rhythm Formants in Spe…

200 篇论文

Large pre-trained language models (LMs) have been widely adopted in biomedical and clinical domains, introducing many powerful LMs such as bio-lm and BioELECTRA. However, the applicability of these methods to real clinical use cases is…

计算与语言 · 计算机科学 2022-11-16 Samuel Cahyawijaya , Bryan Wilie , Holy Lovenia , Huan Zhong , MingQian Zhong , Yuk-Yu Nancy Ip , Pascale Fung

Conventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial…

声音 · 计算机科学 2020-02-10 Taejin Park , Kenichi Kumatani , Minhua Wu , Shiva Sundaram

The way infants use auditory cues to learn to speak despite the acoustic mismatch of their vocal apparatus is a hot topic of scientific debate. The simulation of early vocal learning using articulatory speech synthesis offers a way towards…

音频与语音处理 · 电气工程与系统科学 2021-04-05 Branislav Gerazov , Daniel van Niekerk , Anqi Xu , Paul Konstantin Krug , Peter Birkholz , Yi Xu

Random functional-linked types of neural networks (RFLNNs), e.g., the extreme learning machine (ELM) and broad learning system (BLS), which avoid suffering from a time-consuming training process, offer an alternative way of learning in deep…

机器学习 · 计算机科学 2023-04-04 Guang-Yong Chen , Yong-Hang Yu , Min Gan , C. L. Philip Chen , Wenzhong Guo

Large language models (LLMs) and classical machine learning methods offer complementary strengths for predictive modeling, yet their fundamentally different representations and training paradigms hinder effective integration: LLMs rely on…

计算与语言 · 计算机科学 2026-04-21 Yunshuo Tian , Akayou Kitessa , Tanuja Chitnis , Yijun Zhao

The human vocal folds are known to interact with the vocal tract acoustics during voiced speech production; namely a nonlinear source-filter coupling has been observed both by using models and in \emph{in vivo} phonation. These phenomena…

生物物理 · 物理学 2015-11-17 Daniel Aalto , Jarmo Malinen , Martti Vainio

Disorders of voice production have severe effects on the quality of life of the affected individuals. A simulation approach is used to investigate the cause-effect chain in voice production showing typical characteristics of voice such as…

声音 · 计算机科学 2022-07-20 Florian Kraxberger , Andreas Wurzinger , Stefan Schoder

In this work, a recently proposed Head-Related Transfer Function (HRTF)-based Robust Least-Squares Frequency-Invariant (RLSFI) beamformer design is analyzed with respect to its robustness against localization errors, which lead to a…

声音 · 计算机科学 2016-03-30 Hendrik Barfuss , Walter Kellermann

A state-of-the-art 1D acoustic synthesizer has been previously developed, and coupled to speaker-specific biomechanical models of oropharynx in ArtiSynth. As expected, the formant frequencies of the synthesized vowel sounds were shown to be…

声音 · 计算机科学 2015-12-21 Negar M. Harandi , Daniel Aalto , Antti Hannukainen , Jarmo Malinen , Sidney Fels

Recent studies have introduced end-to-end TTS, which integrates the production of context and acoustic features in statistical parametric speech synthesis. As a result, a single neural network replaced laborious feature engineering with…

机器学习 · 计算机科学 2019-02-26 Kohki Mametani , Tsuneo Kato , Seiichi Yamamoto

Can large language models (LLMs) generate continuous numerical features that improve reinforcement learning (RL) trading agents? We build a modular pipeline where a frozen LLM serves as a stateless feature extractor, transforming…

计算与语言 · 计算机科学 2026-04-14 Zhengzhe Yang

Large Language Models (LLMs) are central to reasoning, writing, and decision-support workflows, yet users lack consistent control over how they reason and express outputs. Conventional prompt engineering relies on verbose natural-language…

编程语言 · 计算机科学 2025-10-24 Mostapha Kalami Heris

End-to-end Large Speech Language Models (LSLMs) have demonstrated impressive conversational generation abilities, yet consistently fall short of traditional pipeline systems on semantic understanding benchmarks. In this work, we reveal…

计算与语言 · 计算机科学 2025-10-15 Bajian Xiang , Shuaijiang Zhao , Tingwei Guo , Wei Zou

Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning…

音频与语音处理 · 电气工程与系统科学 2026-01-27 You Zhang , Andrew Francl , Ruohan Gao , Paul Calamia , Zhiyao Duan , Ishwarya Ananthabhotla

Time-frequency (TF) representations of time series are intrinsically subject to the boundary effects. As a result, the structures of signals that are highlighted by the representations are garbled when approaching the boundaries of the TF…

信号处理 · 电气工程与系统科学 2021-02-24 Adrien Meynard , Hau-Tieng Wu

We use the AdS/CFT correspondence to study propagation of sound waves in strongly coupled (2+1) dimensional conformal magnetic fluids. Our computation provides a nontrivial consistency check of the viscous magneto-hydrodynamics of…

高能物理 - 理论 · 物理学 2009-11-13 Evgeny I. Buchbinder , Alex Buchel , Samuel E. Vazquez

Optical tracking was used to characterize acoustic radiation force (ARF) induced displacements in a tissue-mimicking phantom. Amplitude modulated (AM) 3.3 MHz ultrasound was used to induce ARF in the phantom which was embedded with 10…

医学物理 · 物理学 2016-09-03 Visa Suomi , David Edwards , Robin Cleveland

Mental health disorders impose a substantial global socioeconomic burden. While large language models (LLMs) offer 24/7, non-judgmental interactions to address this gap, pretrained models lack contextual coherence and emotional alignment…

计算与语言 · 计算机科学 2026-02-17 Eric Hua Qing Zhang , Julia Ive

This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram frames from linguistic and speaker conditions, our architecture…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Keyu An , Zhiyu Zhang , Changfeng Gao , Yabin Li , Zhendong Peng , Haoxu Wang , Zhihao Du , Han Zhao , Zhifu Gao , Xiangang Li

In recent years, large language models (LLMs) have played an important role in automatic speech recognition (ASR) and text-to-speech (TTS) systems. While reinforcement learning (RL) has significantly enhanced LLM performance in text-based…

声音 · 计算机科学 2025-09-24 Changfeng Gao , Yabin Li , Keyu An , Zhifu Gao , Zhihao Du , Han Zhao , Xiangang Li