中文
相关论文

相关论文: Chirp Complex Cepstrum-based Decomposition for Asy…

200 篇论文

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

We present a fully automated, two-stage modular glottal area segmentation framework for high-speed videoendoscopy (HSV) designed for accuracy, generalizability, and real-time playback. Our detection-gated pipeline combines a YOLOv8n glottis…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Harikrishnan Unnikrishnan , Rita Patel

The performance of an Acoustic Scene Classification (ASC) system is highly depending on the latent temporal dynamics of the audio signal. In this paper, we proposed a multiple layers temporal pooling method using CNN feature sequence as…

声音 · 计算机科学 2019-04-04 Liwen Zhang , Jiqing Han

During speech perception, a listener's electroencephalogram (EEG) reflects acoustic-level processing as well as higher-level cognitive factors such as speech comprehension and attention. However, decoding speech from EEG recordings is…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Mike Thornton , Danilo Mandic , Tobias Reichenbach

Decoding speech information from scalp EEG remains difficult due to low SNR and spatial blurring. We present CIPHER (Conformer-based Inference of Phonemes from High-density EEG Representations), a dual-pathway model using (i) ERP features…

计算与语言 · 计算机科学 2026-04-06 Varshith Madishetty

Dysarthric speech recognition often suffers from performance degradation due to the intrinsic diversity of dysarthric severity and extrinsic disparity from normal speech. To bridge these gaps, we propose a Dynamic Phoneme-level Contrastive…

计算与语言 · 计算机科学 2025-02-04 Wonjun Lee , Solee Im , Heejin Do , Yunsu Kim , Jungseul Ok , Gary Geunbae Lee

Interval jitter and spike resampling methods are used to analyze the time scale on which temporal correlations occur. They allow the computation of jitter corrected cross correlograms and the performance of an associated statistically…

神经元与认知 · 定量生物学 2015-03-02 Daniel Jeck , Ernst Niebur

Clinical imaging is routinely used for cochlear implant surgical planning yet lacks the resolution and contrast necessary to visualize the fine intracochlear structures critical for individualized intervention. To address this limitation,…

Clustering a lexicon of words is a well-studied problem in natural language processing (NLP). Word clusters are used to deal with sparse data in statistical language processing, as well as features for solving various NLP tasks (text…

计算与语言 · 计算机科学 2018-08-17 Effi Levi , Saggy Herman , Ari Rappoport

Analysis of signals with oscillatory modes with crossover instantaneous frequencies is a challenging problem in time series analysis. One way to handle this problem is lifting the 2-dimensional time-frequency representation to a…

数值分析 · 数学 2022-06-22 Ziyu Chen , Hau-Tieng Wu

Speech separation has been shown effective for multi-talker speech recognition. Under the ad hoc microphone array setup where the array consists of spatially distributed asynchronous microphones, additional challenges must be overcome as…

声音 · 计算机科学 2021-03-04 Dongmei Wang , Takuya Yoshioka , Zhuo Chen , Xiaofei Wang , Tianyan Zhou , Zhong Meng

A language is constructed of a finite/infinite set of sentences composing of words. Similar to natural languages, Electrocardiogram (ECG) signal, the most common noninvasive tool to study the functionality of the heart and diagnose several…

信号处理 · 电气工程与系统科学 2020-06-17 Sajad Mousavi , Fatemeh Afghah , Fatemeh Khadem , U. Rajendra Acharya

Reconstructing the speech audio envelope from scalp neural recordings (EEG) is a central task for decoding a listener's attentional focus in applications like neuro-steered hearing aids. Current methods for this reconstruction, however,…

声音 · 计算机科学 2026-02-24 Karan Thakkar , Mounya Elhilali

We conduct cluster analysis on a class of locally asymptotically self-similar stochastic processes, which includes multifractional Brownian motion as a representative. When the true number of clusters is supposed to be known, a new…

机器学习 · 统计学 2020-01-15 Qidi Peng , Nan Rao , Ran Zhao

Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper, we propose the usage of an asymmetric analysis-synthesis…

音频与语音处理 · 电气工程与系统科学 2021-06-23 Shanshan Wang , Gaurav Naithani , Archontis Politis , Tuomas Virtanen

Frequency to time mapping is a powerful technique for observing ultrafast phenomena and non-repetitive events in optics. However, many optical sources operate in wavelength regions, or at power levels, that are not compatible with standard…

光学 · 物理学 2019-09-04 Yiqing Xu , Stuart G. Murdoch

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

音频与语音处理 · 电气工程与系统科学 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

We present a new system for simultaneous estimation of keys, chords, and bass notes from music audio. It makes use of a novel chromagram representation of audio that takes perception of loudness into account. Furthermore, it is fully based…

声音 · 计算机科学 2011-07-26 Yizhao Ni , Matt Mcvicar , Raul Santos-Rodriguez , Tijl De Bie

Including local automatic gain control (AGC) circuitry into a silicon cochlea design has been challenging because of transistor mismatch and model complexity. To address this, we present an alternative system-level algorithm that implements…

信号处理 · 电气工程与系统科学 2022-02-15 Ilya Kiselev , Chang Gao , Shih-Chii Liu

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT)…

声音 · 计算机科学 2019-01-03 Xingjian Du , Mengyao Zhu , Xuan Shi , Xinpeng Zhang , Wen Zhang , Jingdong Chen