中文
相关论文

相关论文: Whispy: Adapting STT Whisper Models to Real-Time E…

200 篇论文

This article presents a whisper speech detector in the far-field domain. The proposed system consists of a long-short term memory (LSTM) neural network trained on log-filterbank energy (LFBE) acoustic features. This model is trained and…

Machine recognition of an atypical speech like whispered speech, is a challenging task. We introduce whisper-to-natural-speech conversion using sequence-to-sequence approach by proposing enhanced transformer architecture, which uses both…

音频与语音处理 · 电气工程与系统科学 2021-04-06 Abhishek Niranjan , Mukesh Sharma , Sai Bharath Chandra Gutha , M Ali Basha Shaik

There is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require additional training or are simply not competitive. The…

音频与语音处理 · 电气工程与系统科学 2025-09-15 Sung-Lin Yeh , Yen Meng , Hao Tang

In this paper, we propose a new class of high-efficiency semantic coded transmission methods for end-to-end speech transmission over wireless channels. We name the whole system as deep speech semantic transmission (DSST). Specifically, we…

声音 · 计算机科学 2022-11-07 Zixuan Xiao , Shengshi Yao , Jincheng Dai , Sixian Wang , Kai Niu , Ping Zhang

Voice interfaces are quickly becoming a common way for people to interact with AI systems. This also brings new security risks, such as prompt injection, social engineering, and harmful voice commands. Traditional security methods rely on…

声音 · 计算机科学 2026-03-10 Sumit Ranjan , Sugandha Sharma , Ubaid Abbas , Puneeth N Ail

Speaker-adaptive Text-to-Speech (TTS) synthesis has attracted considerable attention due to its broad range of applications, such as personalized voice assistant services. While several approaches have been proposed, they often exhibit high…

声音 · 计算机科学 2024-12-31 Wooseok Han , Minki Kang , Changhun Kim , Eunho Yang

Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these models is their…

声音 · 计算机科学 2024-10-14 Vyas Raina , Mark Gales

The conventional paradigm in speech translation starts with a speech recognition step to generate transcripts, followed by a translation step with the automatic transcripts as input. To address various shortcomings of this paradigm, recent…

计算与语言 · 计算机科学 2020-08-31 Matthias Sperber , Hendra Setiawan , Christian Gollan , Udhyakumar Nallasamy , Matthias Paulik

Large-scale in-the-wild speech datasets have become more prevalent in recent years due to increased interest in models that can learn useful features from unlabelled data for tasks such as speech recognition or synthesis. These datasets…

Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in adaptivity,…

计算与语言 · 计算机科学 2024-12-24 Amir Nassereldine , Dancheng Liu , Chenhui Xu , Ruiyang Qin , Yiyu Shi , Jinjun Xiong

As the size of pre-trained speech recognition models increases, running these large models in low-latency or resource-constrained environments becomes challenging. In this work, we leverage pseudo-labelling to assemble a large-scale…

计算与语言 · 计算机科学 2023-11-02 Sanchit Gandhi , Patrick von Platen , Alexander M. Rush

Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs. In this paper, we examine simple extensions to a…

音频与语音处理 · 电气工程与系统科学 2022-06-30 Peter Makarov , Ammar Abbas , Mateusz Łajszczak , Arnaud Joly , Sri Karlapati , Alexis Moinet , Thomas Drugman , Penny Karanasou

Large encoder-decoder models like Whisper achieve strong offline transcription but remain impractical for streaming applications due to high latency. However, due to the accessibility of pre-trained checkpoints, the open Thai ASR landscape…

As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios. However, its encoder-decoder structure hinders its application to…

声音 · 计算机科学 2025-05-06 Haoyu Wang , Guoqiang Hu , Guodong Lin , Wei-Qiang Zhang , Jian Li

Automatic speech recognition has recently seen a significant advancement with large foundational models such as Whisper. However, these models often struggle to perform well in low-resource languages, such as Indian languages. This paper…

计算与语言 · 计算机科学 2024-12-30 Kumud Tripathi , Raj Gothi , Pankaj Wasnik

This paper investigates the in-context learning abilities of the Whisper automatic speech recognition (ASR) models released by OpenAI. A novel speech-based in-context learning (SICL) approach is proposed for test-time adaptation, which can…

音频与语音处理 · 电气工程与系统科学 2024-03-21 Siyin Wang , Chao-Han Huck Yang , Ji Wu , Chao Zhang

After decades of use in dictation and, more recently, ambient documentation, speech is emerging as a primary modality for interacting with technology and AI in healthcare. Yet medical speech recognition remains difficult: systems must…

State of the art (SOTA) neural text to speech (TTS) models can generate natural-sounding synthetic voices. These models are characterized by large memory footprints and substantial number of operations due to the long-standing focus on…

音频与语音处理 · 电气工程与系统科学 2023-05-24 Rowel Atienza

In the realm of automatic speech recognition (ASR), robustness in noisy environments remains a significant challenge. Recent ASR models, such as Whisper, have shown promise, but their efficacy in noisy conditions can be further enhanced.…

声音 · 计算机科学 2024-06-28 Yehoshua Dissen , Shiry Yonash , Israel Cohen , Joseph Keshet

Whisper, despite being trained on 680K hours of web-scaled audio data, faces difficulty in recognising rare words like domain-specific terms, with a solution being contextual biasing through prompting. To improve upon this method, in this…

音频与语音处理 · 电气工程与系统科学 2025-02-19 Yash Jogi , Vaibhav Aggarwal , Shabari S Nair , Yash Verma , Aayush Kubba