中文
相关论文

相关论文: GCI detection from raw speech using a fully-convol…

200 篇论文

Many voice disorders induce subharmonic phonation, but voice signal analysis is currently lacking a technique to detect the presence of subharmonics reliably. Distinguishing subharmonic phonation from normal phonation is a challenging task…

音频与语音处理 · 电气工程与系统科学 2025-01-17 Takeshi Ikuma , Melda Kunduk , Brad Story , Andrew J. McWhorter

Space-based gravitational wave (GW) detectors will be able to observe signals from sources that are otherwise nearly impossible from current ground-based detection. Consequently, the well established signal detection method, matched…

广义相对论与量子宇宙学 · 物理学 2023-08-17 Tianyu Zhao , Ruoxi Lyu , He Wang , Zhoujian Cao , Zhixiang Ren

Silent speech interfaces (SSI) are being actively developed to assist individuals with communication impairments who have long suffered from daily hardships and a reduced quality of life. However, silent sentences are difficult to segment…

人机交互 · 计算机科学 2025-09-19 Yudong Xie , Zhifeng Han , Qinfan Xiao , Liwei Liang , Lu-Qi Tao , Tian-Ling Ren

Brain-computer interface (BCI) is the technology that enables the communication between humans and devices by reflecting status and intentions of humans. When conducting imagined speech, the users imagine the pronunciation as if actually…

人机交互 · 计算机科学 2021-12-15 Dae-Hyeok Lee , Sung-Jin Kim , Keon-Woo Lee

Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded…

计算与语言 · 计算机科学 2024-07-22 Stefano Bannò , Rao Ma , Mengjie Qian , Kate M. Knill , Mark J. F. Gales

Silent speech interfaces (SSI) aim to reconstruct the speech signal from a recording of the articulatory movement, such as an ultrasound video of the tongue. Currently, deep neural networks are the most successful technology for this task.…

声音 · 计算机科学 2021-04-26 László Tóth , Amin Honarmandi Shandiz

In conventional target tracking systems, human operators use the estimated target tracks to make higher level inference of the target behaviour/intent. This paper develops syntactic filtering algorithms that assist human operators by…

统计方法学 · 统计学 2011-04-25 Alex Wang , Vikram Krishnamurthy , Bhashyam Balaji

Inspired by recent work in meta-learning and generative teaching networks, we propose a framework called Generative Conversational Networks, in which conversational agents learn to generate their own labelled training data (given some seed…

Speech synthesis is used in a wide variety of industries. Nonetheless, it always sounds flat or robotic. The state of the art methods that allow for prosody control are very cumbersome to use and do not allow easy tuning. To tackle some of…

声音 · 计算机科学 2021-10-08 Enrique Hortal , Rodrigo Brechard Alarcia

The Glottal Source is an important component of voice as it can be considered as the excitation signal to the voice apparatus. Nowadays, new techniques of speech processing such as speech recognition and speech synthesis use the glottal…

其他计算机科学 · 计算机科学 2010-04-20 Nikhil Raj , R. K. Sharma

Non-intrusive speech intelligibility (SI) prediction from binaural signals is useful in many applications. However, most existing signal-based measures are designed to be applied to single-channel signals. Measures specifically designed to…

声音 · 计算机科学 2022-03-23 Alex F. McKinney , Benjamin Cauchi

Current state-of-the-art speech recognition systems build on recurrent neural networks for acoustic and/or language modeling, and rely on feature extraction pipelines to extract mel-filterbanks or cepstral coefficients. In this paper we…

计算与语言 · 计算机科学 2019-04-10 Neil Zeghidour , Qiantong Xu , Vitaliy Liptchinsky , Nicolas Usunier , Gabriel Synnaeve , Ronan Collobert

Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy…

计算与语言 · 计算机科学 2024-08-02 Rui Liu , Yifan Hu , Yi Ren , Xiang Yin , Haizhou Li

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed…

音频与语音处理 · 电气工程与系统科学 2018-09-18 Francois G. Germain , Qifeng Chen , Vladlen Koltun

Open-vocabulary 3D scene understanding is crucial for robotics applications, such as natural language-driven manipulation, human-robot interaction, and autonomous navigation. Existing methods for querying 3D Gaussian Splatting often…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Hairong Yin , Huangying Zhan , Yi Xu , Raymond A. Yeh

Gait is one of the most promising biometrics that aims to identify pedestrians from their walking patterns. However, prevailing methods are susceptible to confounders, resulting in the networks hardly focusing on the regions that reflect…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Huanzhang Dou , Pengyi Zhang , Wei Su , Yunlong Yu , Yining Lin , Xi Li

Signal extraction out of background noise is a common challenge in high precision physics experiments, where the measurement output is often a continuous data stream. To improve the signal to noise ratio of the detection, witness sensors…

广义相对论与量子宇宙学 · 物理学 2020-02-26 Gabriele Vajente , Yiwen Huang , Maximiliano Isi , Jenne C. Driggers , Jeffrey S. Kissel , Marek J. Szczepanczyk , Salvatore Vitale

P300 speller BCIs allow users to compose sentences by selecting target keys on a GUI through the detection of P300 component in their EEG signals following visual stimuli. Most P300 speller BCIs require users to spell words letter by…

人机交互 · 计算机科学 2026-02-19 Jiazhen Hong , Weinan Wang , Laleh Najafizadeh

Generating continuous electroencephalography (EEG) signals through advanced artificial neural networks presents a novel opportunity to enhance brain-computer interface (BCI) technology. This capability has the potential to significantly…

神经元与认知 · 定量生物学 2024-06-11 Omair Ali , Muhammad Saif-ur-Rehman , Marita Metzler , Tobias Glasmachers , Ioannis Iossifidis , Christian Klaes

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis system, to uncover…

计算与语言 · 计算机科学 2018-08-07 Daisy Stanton , Yuxuan Wang , RJ Skerry-Ryan