中文
相关论文

相关论文: A practical two-stage training strategy for multi-…

200 篇论文

The advent of foundation models have revolutionized various fields, enabling unprecedented task accuracy and flexibility in computational linguistics, computer vision and other domains. Attention mechanism has become an essential component…

分布式、并行与集群计算 · 计算机科学 2025-05-19 Mohammadali Shakerdargah , Shan Lu , Chao Gao , Di Niu

For time-frequency (TF) domain speech enhancement (SE) methods, the overlap-and-add operation in the inverse TF transformation inevitably leads to an algorithmic delay equal to the window size. However, typical causal SE systems fail to…

音频与语音处理 · 电气工程与系统科学 2025-01-22 Yuewei Zhang , Huanbin Zou , Jie Zhu

Casual conversations involving multiple speakers and noises from surrounding devices are common in everyday environments, which degrades the performances of automatic speech recognition systems. These challenging characteristics of…

音频与语音处理 · 电气工程与系统科学 2019-06-24 Nelson Yalta , Shinji Watanabe , Takaaki Hori , Kazuhiro Nakadai , Tetsuya Ogata

We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the independent vector analysis (IVA) paradigm. It uses the fast…

音频与语音处理 · 电气工程与系统科学 2022-04-04 Robin Scheibler , Wangyou Zhang , Xuankai Chang , Shinji Watanabe , Yanmin Qian

To accomplish punctuation restoration, most existing methods focus on introducing extra information (e.g., part-of-speech) or addressing the class imbalance problem. Recently, large-scale transformer-based pre-trained language models (PLMS)…

计算与语言 · 计算机科学 2022-11-10 Yangjun Wu , Kebin Fang , Yao Zhao , Hao Zhang , Lifeng Shi , Mengqi Zhang

Target speaker extraction focuses on extracting a target speech signal from an environment with multiple speakers by leveraging an enrollment. Existing methods predominantly rely on speaker embeddings obtained from the enrollment,…

声音 · 计算机科学 2025-02-13 Ke Xue , Rongfei Fan , Shanping Yu , Chang Sun , Jianping An

Deep learning has dramatically improved the performance of speech recognition systems through learning hierarchies of features optimized for the task at hand. However, true end-to-end learning, where features are learned directly from…

计算与语言 · 计算机科学 2016-04-06 Zhenyao Zhu , Jesse H. Engel , Awni Hannun

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired…

音频与语音处理 · 电气工程与系统科学 2022-04-12 Zehai Tu , Jack Deadman , Ning Ma , Jon Barker

Multilingual end-to-end(E2E) models have shown a great potential in the expansion of the language coverage in the realm of automatic speech recognition(ASR). In this paper, we aim to enhance the multilingual ASR performance in two ways,…

计算与语言 · 计算机科学 2021-10-18 Rimita Lahiri , Kenichi Kumatani , Eric Sun , Yao Qian

The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a…

Retinal blood vessel segmentation is crucial for diagnosing ocular and cardiovascular diseases. Although the introduction of U-Net in 2015 by Olaf Ronneberger significantly advanced this field, yet issues like limited training data,…

图像与视频处理 · 电气工程与系统科学 2025-06-04 Md Tauhidul Islam , Wu Da-Wen , Tang Qing-Qing , Zhao Kai-Yang , Yin Teng , Li Yan-Fei , Shang Wen-Yi , Liu Jing-Yu , Zhang Hai-Xian

Disfluency detection is usually an intermediate step between an automatic speech recognition (ASR) system and a downstream task. By contrast, this paper aims to investigate the task of end-to-end speech recognition and disfluency removal.…

音频与语音处理 · 电气工程与系统科学 2020-09-30 Paria Jamshid Lou , Mark Johnson

We present a method for converting the voices between a set of speakers. Our method is based on training multiple autoencoder paths, where there is a single speaker-independent encoder and multiple speaker-dependent decoders. The…

音频与语音处理 · 电气工程与系统科学 2019-05-13 Orhan Ocal , Oguz H. Elibol , Gokce Keskin , Cory Stephenson , Anil Thomas , Kannan Ramchandran

Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models' performance. In this paper, we propose a decoupled…

声音 · 计算机科学 2020-10-29 Shuai Zhang , Jiangyan Yi , Zhengkun Tian , Ye Bai , Jianhua Tao , Zhengqi wen

End-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a unified model. Although it simplifies training and decoding pipelines, the unified model is hard to…

计算与语言 · 计算机科学 2018-12-06 Zhehuai Chen , Mahaveer Jain , Yongqiang Wang , Michael L. Seltzer , Christian Fuegen

This paper proposes a textless training method for many-to-many multilingual speech-to-speech translation that can also benefit the transfer of pre-trained knowledge to text-based systems, text-to-speech synthesis and text-to-speech…

计算与语言 · 计算机科学 2024-08-20 Minsu Kim , Jeongsoo Choi , Dahun Kim , Yong Man Ro

Streaming processing of speech audio is required for many contemporary practical speech recognition tasks. Even with the large corpora of manually transcribed speech data available today, it is impossible for such corpora to cover…

计算与语言 · 计算机科学 2021-04-12 Rodrigo Cabrera , Xiaofeng Liu , Mohammadreza Ghodsi , Zebulun Matteson , Eugene Weinstein , Anjuli Kannan

Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain…

音频与语音处理 · 电气工程与系统科学 2025-09-15 Peter Vieting , Benedikt Hilmes , Ralf Schlüter , Hermann Ney

Running AI models on smart edge devices can unlock versatile user experiences, but presents challenges due to limited compute and the need to handle multiple tasks simultaneously. This requires a vision encoder with small size but powerful…

Existing multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training and real-data testing, or the real transcribed overlapping…

音频与语音处理 · 电气工程与系统科学 2022-04-08 Xiaofei Wang , Dongmei Wang , Naoyuki Kanda , Sefik Emre Eskimez , Takuya Yoshioka
‹ 上一页 1 8 9 10 下一页 ›