中文
相关论文

相关论文: CIF: Continuous Integrate-and-Fire for End-to-End …

200 篇论文

The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC)…

音频与语音处理 · 电气工程与系统科学 2026-03-04 Yilei Wu , Changyan Zheng , Xingyu Zhang , Yakun Zhang , Chengshi Zheng , Shuang Yang , Ye Yan , Erwei Yin

Spiking neural networks (SNNs) offer biologically inspired computation but remain underexplored for continuous regression tasks in scientific machine learning. In this work, we introduce and systematically evaluate Quadratic…

神经与进化计算 · 计算机科学 2025-11-11 Ruyin Wan , George Em Karniadakis , Panos Stinis

Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learning-based methods show remarkable performance, but are…

计算机视觉与模式识别 · 计算机科学 2023-08-09 Zhu Liu , Jinyuan Liu , Benzhuang Zhang , Long Ma , Xin Fan , Risheng Liu

In this work, we introduce a simple yet efficient post-processing model for automatic speech recognition (ASR). Our model has Transformer-based encoder-decoder architecture which "translates" ASR model output into grammatically and…

计算与语言 · 计算机科学 2019-10-24 Oleksii Hrinchuk , Mariya Popova , Boris Ginsburg

End-to-end ASR models trained on large amount of data tend to be implicitly biased towards language semantics of the training data. Internal language model estimation (ILME) has been proposed to mitigate this bias for autoregressive models…

音频与语音处理 · 电气工程与系统科学 2023-05-09 Nilaksh Das , Monica Sunkara , Sravan Bodapati , Jinglun Cai , Devang Kulshreshtha , Jeff Farris , Katrin Kirchhoff

Copy-move forgery on speech (CMF), coupled with post-processing techniques, presents a great challenge to the forensic detection and localization of tampered areas. Most of the existing CMF detection approaches necessitate pre-segmentation…

音频与语音处理 · 电气工程与系统科学 2023-09-11 Dong Yang , Mingle Liu , Muyong Cao

Neural contextual biasing effectively improves automatic speech recognition (ASR) for crucial phrases within a speaker's context, particularly those that are infrequent in the training data. This work proposes contextual text injection…

计算与语言 · 计算机科学 2024-06-12 Zhong Meng , Zelin Wu , Rohit Prabhavalkar , Cal Peyser , Weiran Wang , Nanxin Chen , Tara N. Sainath , Bhuvana Ramabhadran

Integrated sensing and communication (ISAC) has emerged as a transformative paradigm, enabling situationally aware and perceptive next-generation wireless networks through the co-design of shared network resources. With the adoption of…

信号处理 · 电气工程与系统科学 2025-05-20 Ahmed Hussain , Asmaa Abdallah , Abdulkadir Celik , Ahmed M. Eltawil

Speech technology is a field that encompasses various techniques and tools used to enable machines to interact with speech, such as automatic speech recognition (ASR), spoken dialog systems, and others, allowing a device to capture spoken…

音频与语音处理 · 电气工程与系统科学 2024-03-14 Amanda Phaladi , Thipe Modipa

We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-time speaker adaptation, and (iii) test-time contextual…

音频与语音处理 · 电气工程与系统科学 2024-10-01 Hao Yen , Shaoshi Ling , Guoli Ye

In this paper, we investigate the benefit that off-the-shelf word embedding can bring to the sequence-to-sequence (seq-to-seq) automatic speech recognition (ASR). We first introduced the word embedding regularization by maximizing the…

计算与语言 · 计算机科学 2020-02-06 Alexander H. Liu , Tzu-Wei Sung , Shun-Po Chuang , Hung-yi Lee , Lin-shan Lee

We propose Citrinet - a new end-to-end convolutional Connectionist Temporal Classification (CTC) based automatic speech recognition (ASR) model. Citrinet is deep residual neural model which uses 1D time-channel separable convolutions…

音频与语音处理 · 电气工程与系统科学 2021-04-06 Somshubra Majumdar , Jagadeesh Balam , Oleksii Hrinchuk , Vitaly Lavrukhin , Vahid Noroozi , Boris Ginsburg

Time series classification (TSC) is home to a number of algorithm groups that utilise different kinds of discriminatory patterns. One of these groups describes classifiers that predict using phase dependant intervals. The time series forest…

机器学习 · 计算机科学 2021-05-11 Matthew Middlehurst , James Large , Anthony Bagnall

Speech foundation models (SFMs), such as Open Whisper-Style Speech Models (OWSM), are trained on massive datasets to achieve accurate automatic speech recognition. However, even SFMs struggle to accurately recognize rare and unseen words.…

声音 · 计算机科学 2025-06-12 Yui Sudo , Yusuke Fujita , Atsushi Kojima , Tomoya Mizumoto , Lianbo Liu

The state-of-art models for speech synthesis and voice conversion are capable of generating synthetic speech that is perceptually indistinguishable from bonafide human speech. These methods represent a threat to the automatic speaker…

机器学习 · 计算机科学 2019-07-11 Moustafa Alzantot , Ziqi Wang , Mani B. Srivastava

Speech emotion recognition~(SER) refers to the technique of inferring the emotional state of an individual from speech signals. SERs continue to garner interest due to their wide applicability. Although the domain is mainly founded on…

音频与语音处理 · 电气工程与系统科学 2022-03-29 Sneha Das , Nicklas Leander Lund , Nicole Nadine Lønfeldt , Anne Katrine Pagsberg , Line H. Clemmensen

End-to-end (E2E) automatic speech recognition (ASR) systems have revolutionized the field by integrating all components into a single neural network, with attention-based encoder-decoder models achieving state-of-the-art performance.…

计算与语言 · 计算机科学 2025-07-01 Duygu Altinok

Up to now, modern Machine Learning is mainly based on fitting high dimensional functions to enormous data sets, taking advantage of huge hardware resources. We show that biologically inspired neuron models such as the…

神经与进化计算 · 计算机科学 2022-10-03 Richard C. Gerum , Achim Schilling

Silent speech interfaces (SSI) are being actively developed to assist individuals with communication impairments who have long suffered from daily hardships and a reduced quality of life. However, silent sentences are difficult to segment…

人机交互 · 计算机科学 2025-09-19 Yudong Xie , Zhifeng Han , Qinfan Xiao , Liwei Liang , Lu-Qi Tao , Tian-Ling Ren

At the pinnacle of computational imaging is the co-optimization of camera and algorithm. This, however, is not the only form of computational imaging. In problems such as imaging through adverse weather, the bigger challenge is how to…

图像与视频处理 · 电气工程与系统科学 2023-10-30 Stanley H. Chan