中文
相关论文

相关论文: Achieving Timestamp Prediction While Recognizing w…

200 篇论文

Mapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) performance in new domains. However, the length of speech…

声音 · 计算机科学 2023-10-10 Jiaxu Zhu , Weinan Tong , Yaoxun Xu , Changhe Song , Zhiyong Wu , Zhao You , Dan Su , Dong Yu , Helen Meng

Weighted finite-state automata (WSFAs) are commonly used in NLP. Failure transitions are a useful extension for compactly representing backoffs or interpolation in $n$-gram models and CRFs, which are special cases of WFSAs. The pathsum in…

数据结构与算法 · 计算机科学 2023-07-12 Anej Svete , Benjamin Dayan , Tim Vieira , Ryan Cotterell , Jason Eisner

Automatic pronunciation error detection (APED) plays an important role in the domain of language learning. As for the previous ASR-based APED methods, the decoded results need to be aligned with the target text so that the errors can be…

音频与语音处理 · 电气工程与系统科学 2021-05-06 Zhan Zhang , Yuehai Wang , Jianyi Yang

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy. Misalignment…

计算与语言 · 计算机科学 2025-07-28 Hsuan-Yu Wang , Pei-Ying Lee , Berlin Chen

For end-to-end Automatic Speech Recognition (ASR) models, recognizing personal or rare phrases can be hard. A promising way to improve accuracy is through spelling correction (or rewriting) of the ASR lattice, where potentially…

计算与语言 · 计算机科学 2024-09-26 Leonid Velikovich , Christopher Li , Diamantino Caseiro , Shankar Kumar , Pat Rondon , Kandarp Joshi , Xavier Velez

In this paper, we introduce a novel self-calibrating integrate-and-fire time encoding machine (S-IF-TEM) that enables simultaneous parameter estimation and signal reconstruction during sampling, thereby effectively mitigating mismatch…

信号处理 · 电气工程与系统科学 2025-09-16 Maya Mekel , Vered Karp , Satish Mulleti , Alejandro Cohen

This paper proposes a tensor-based parametric channel estimation technique for IRS-assisted communication systems with time-varying channel parameters. We exploit the multidimensional structure of the received signal by developing a…

信号处理 · 电气工程与系统科学 2026-05-29 Kenneth B. A. Benício , André L. F. de Almeida , Bruno Sokal , Fazal-E-Asim , Behrooz Makki , Gabor Fodor

Reducing prediction delay for streaming end-to-end ASR models with minimal performance regression is a challenging problem. Constrained alignment is a well-known existing approach that penalizes predicted word boundaries using external…

音频与语音处理 · 电气工程与系统科学 2021-05-12 Jaeyoung Kim , Han Lu , Anshuman Tripathi , Qian Zhang , Hasim Sak

An adaptive time-frequency representation (TFR) with higher energy concentration usually requires higher complexity. Recently, a low-complexity adaptive short-time Fourier transform (ASTFT) based on the chirp rate has been proposed. To…

信息论 · 计算机科学 2017-05-26 Soo-Chang Pei , Shih-Gu Huang

Power system dynamic state estimation is essential to monitoring and controlling power system stability. Kalman filtering approaches are predominant in estimation of synchronous machine dynamic states (i.e. rotor angle and rotor speed).…

系统与控制 · 计算机科学 2017-02-03 Shahrokh Akhlaghi , Ning Zhou

With regard to a three-step estimation procedure, proposed without theoretical discussion by Li and You in Journal of Applied Statistics and Management, for a nonparametric regression model with time-varying regression function, local…

统计理论 · 数学 2020-10-27 Jiyanglin Li , Tao Li

Adaptive Local Iterative Filtering (ALIF) is a currently proposed novel time-frequency analysis tool. It has been empirically shown that ALIF is able to separate components and overcome the mode-mixing problem. However, so far its…

数值分析 · 数学 2020-05-12 Antonio Cicone , Hau-Tieng Wu

The conventional recipe for Automatic Speech Recognition (ASR) models is to 1) train multiple checkpoints on a training set while relying on a validation set to prevent overfitting using early stopping and 2) average several last…

计算与语言 · 计算机科学 2023-08-08 Fangyuan Wang , Ming Hao , Yuhai Shi , Bo Xu

Time Series Foundation Models (TSFMs) advance generalization and data efficiency in time series forecasting by unified large-scale pretraining. But TSFMs remain lacking when adapting to specific downstream forecasting tasks for two reasons.…

信号处理 · 电气工程与系统科学 2026-05-04 Siyang Li , Yize Chen , Zijie Zhu , Yuxin Pan , Yan Guo , Ming Huang , Hui Xiong

Non-autoregressive mechanisms can significantly decrease inference time for speech transformers, especially when the single step variant is applied. Previous work on CTC alignment-based single step non-autoregressive transformer (CASS-NAT)…

音频与语音处理 · 电气工程与系统科学 2021-07-23 Ruchao Fan , Wei Chu , Peng Chang , Jing Xiao , Abeer Alwan

This paper presents the use of non-autoregressive (NAR) approaches for joint automatic speech recognition (ASR) and spoken language understanding (SLU) tasks. The proposed NAR systems employ a Conformer encoder that applies connectionist…

音频与语音处理 · 电气工程与系统科学 2023-04-24 Mohan Li , Rama Doddipatla

Incorporating item-side information, such as category and brand, into sequential recommendation is a well-established and effective approach for improving performance. However, despite significant advancements, current models are generally…

信息检索 · 计算机科学 2026-01-01 Jie Luo , Wenyu Zhang , Xinming Zhang , Yuan Fang

Compute-and-forward (CPF) strategy is one category of network coding in which a relay will compute and forward a linear combination of source messages according to the observed channel coefficients, based on the algebraic structure of…

信号处理 · 电气工程与系统科学 2020-03-09 Lili Wei , Wen Chen

Accurate time-series forecasting for complex physical systems is the backbone of modern industrial monitoring and control, yet deep learning models often lack the physical consistency required in regulated environments.To bridge this gap,…

机器学习 · 计算机科学 2026-05-18 Ramona Rubini , Siavash Khodakarami , Aniruddha Bora , George Em Karniadakis , Michele Dassisti

Finding synthetic artifacts of spoofing data will help the anti-spoofing countermeasures (CMs) system discriminate between spoofed and real speech. The Conformer combines the best of convolutional neural network and the Transformer,…

声音 · 计算机科学 2023-10-31 Yikang Wang , Hiromitsu Nishizaki , Ming Li