中文
相关论文

相关论文: AURORA Model of Formant-to-Tongue Inversion for Di…

200 篇论文

Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the…

声音 · 计算机科学 2024-03-13 Yudong Yang , Rongfeng Su , Xiaokang Liu , Nan Yan , Lan Wang

Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent representations frequently entangle physiologic severity, intervention intensity, observational…

机器学习 · 计算机科学 2026-05-19 Yuanyun Zhang , Shi Li

Despite advances in language and speech technologies, no open-source system enables full speech-to-speech, multi-turn dialogue with integrated tool use and agentic reasoning. We introduce AURA (Agent for Understanding, Reasoning, and…

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems designed for normal speech. Their practical application to atypical task…

音频与语音处理 · 电气工程与系统科学 2023-06-23 Shujie Hu , Xurong Xie , Mengzhe Geng , Mingyu Cui , Jiajun Deng , Guinan Li , Tianzi Wang , Xunying Liu , Helen Meng

Thousands of individuals need surgical removal of their larynx due to critical diseases every year and therefore, require an alternative form of communication to articulate speech sounds after the loss of their voice box. This work…

图像与视频处理 · 电气工程与系统科学 2020-07-01 Pramit Saha , Yadong Liu , Bryan Gick , Sidney Fels

Conventional online surveys provide limited personalization, often resulting in low engagement and superficial responses. Although AI survey chatbots improve convenience, most are still reactive: they rely on fixed dialogue trees or static…

人机交互 · 计算机科学 2025-11-10 Jinwen Tang , Yi Shang

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

计算与语言 · 计算机科学 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Pre-trained audio models excel at detecting acoustic patterns in auscultation sounds but often fail to grasp their clinical significance, limiting their use and performance in diagnostic tasks. To bridge this gap, we introduce AcuLa…

声音 · 计算机科学 2026-04-20 Tsai-Ning Wang , Lin-Lin Chen , Neil Zeghidour , Aaqib Saeed

Acoustic vowel dynamics have some speaker-identifying characteristics, which have been ascribed to individual properties of articulatory strategies: formant transitions have a particular shape because speakers move their articulators, using…

计算与语言 · 计算机科学 2026-05-25 Patrycja Strycharczuk , Justin J. H. Lo , Sam Kirkham

Acoustic articulatory inversion is a major processing challenge, with a wide range of applications from speech synthesis to feedback systems for language learning and rehabilitation. In recent years, deep learning methods have been applied…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

This review is focused on the data-driven approaches applied in different applications of Acoustic-to-Articulatory Inversion (AAI) of speech. This review paper considered the relevant works published in the last ten years (2011-2021). The…

声音 · 计算机科学 2025-04-21 Leena G Pillai , D. Muhammad Noorul Mubarak

Academic advising in higher education is under severe strain, with advisor-to-student ratios commonly exceeding 300:1. These structural bottlenecks limit timely access to guidance, increase the risk of delayed graduation, and contribute to…

We investigate the automatic processing of child speech therapy sessions using ultrasound visual biofeedback, with a specific focus on complementing acoustic features with ultrasound images of the tongue for the tasks of speaker diarization…

音频与语音处理 · 电气工程与系统科学 2019-08-16 Manuel Sam Ribeiro , Aciel Eshky , Korin Richmond , Steve Renals

Virtual cell modeling predicts molecular state changes under genetic perturbations in silico, which is essential for biological mechanism studies. However, existing approaches suffer from unconstrained reasoning, uninterpretable…

定量方法 · 定量生物学 2026-04-23 Zhenyu Wang , Geyan Ye , Wei Liu , Man Tat Alexander Ng

Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modeling. This paper presents a systematic empirical study of…

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques as input (e.g. ultrasound tongue imaging, MRI, lip video). The advantage of lip video is that it is easily…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Frigyes Viktor Arthur , Tamás Gábor Csapó

We present a multilinear statistical model of the human tongue that captures anatomical and tongue pose related shape variations separately. The model is derived from 3D magnetic resonance imaging data of 11 speakers sustaining speech…

计算机视觉与模式识别 · 计算机科学 2018-04-18 Alexander Hewer , Stefanie Wuhrer , Ingmar Steiner , Korin Richmond

We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Ilpo Viertola , Vladimir Iashin , Esa Rahtu

A multi-turn dialogue is composed of multiple utterances from two or more different speaker roles. Thus utterance- and speaker-aware clues are supposed to be well captured in models. However, in the existing retrieval-based multi-turn…

计算与语言 · 计算机科学 2020-12-15 Longxiang Liu , Zhuosheng Zhang , Hai Zhao , Xi Zhou , Xiang Zhou

For articulatory-to-acoustic mapping, typically only limited parallel training data is available, making it impossible to apply fully end-to-end solutions like Tacotron2. In this paper, we experimented with transfer learning and adaptation…

音频与语音处理 · 电气工程与系统科学 2021-07-27 Csaba Zainkó , László Tóth , Amin Honarmandi Shandiz , Gábor Gosztolya , Alexandra Markó , Géza Németh , Tamás Gábor Csapó
‹ 上一页 1 2 3 10 下一页 ›