中文
相关论文

相关论文: OWSM v3.1: Better and Faster Open Whisper-Style Sp…

200 篇论文

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and…

Speech foundation models (SFMs), such as Open Whisper-Style Speech Models (OWSM), are trained on massive datasets to achieve accurate automatic speech recognition. However, even SFMs struggle to accurately recognize rare and unseen words.…

声音 · 计算机科学 2025-06-12 Yui Sudo , Yusuke Fujita , Atsushi Kojima , Tomoya Mizumoto , Lianbo Liu

The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are…

计算与语言 · 计算机科学 2024-06-14 Jinchuan Tian , Yifan Peng , William Chen , Kwanghee Choi , Karen Livescu , Shinji Watanabe

The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a…

计算与语言 · 计算机科学 2025-06-03 Yifan Peng , Shakeel Muhammad , Yui Sudo , William Chen , Jinchuan Tian , Chyi-Jiunn Lin , Shinji Watanabe

There has been an increasing interest in large speech models that can perform multiple tasks in a single model. Such models usually adopt an encoder-decoder or decoder-only architecture due to their popularity and good performance in many…

计算与语言 · 计算机科学 2024-08-28 Yifan Peng , Yui Sudo , Muhammad Shakeel , Shinji Watanabe

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. However, most advanced…

Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion…

计算与语言 · 计算机科学 2026-01-19 Chin-Jou Li , Kalvin Chang , Shikhar Bharadwaj , Eunjung Yeo , Kwanghee Choi , Jian Zhu , David Mortensen , Shinji Watanabe

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality…

The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inaccessible training data and code--poses major reproducibility…

OpenAI's Whisper Automated Speech Recognition model excels in generalizing across diverse datasets and domains. However, this broad adaptability can lead to diminished performance in tasks requiring recognition of specific vocabularies.…

人工智能 · 计算机科学 2025-08-12 Vishakha Lall , Yisi Liu

Improvements in training data scale and quality have led to significant advances, yet its influence in speech recognition remains underexplored. In this paper, we present a large-scale dataset, OLMoASR-Pool, and series of models, OLMoASR,…

声音 · 计算机科学 2025-08-29 Huong Ngo , Matt Deitke , Martijn Bartelds , Sarah Pratt , Josh Gardner , Matt Jordan , Ludwig Schmidt

Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of…

计算与语言 · 计算机科学 2025-04-14 Xabier de Zuazo , Eva Navas , Ibon Saratxaga , Inma Hernáez Rioja

This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates…

计算与语言 · 计算机科学 2025-06-03 Yosuke Kashiwagi , Hayato Futami , Emiru Tsunoo , Satoshi Asakawa

Speech foundation models achieve strong generalization across languages and acoustic conditions, but require significant computational resources for inference. In the context of speech foundation models, pruning techniques have been studied…

音频与语音处理 · 电气工程与系统科学 2025-05-27 Masao Someki , Shikhar Bharadwaj , Atharva Anand Joshi , Chyi-Jiunn Lin , Jinchuan Tian , Jee-weon Jung , Markus Müller , Nathan Susanj , Jing Liu , Shinji Watanabe

We investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering. We selected three tasks: audio-visual speech recognition (AVSR), code-switched speech…

音频与语音处理 · 电气工程与系统科学 2023-08-17 Puyuan Peng , Brian Yan , Shinji Watanabe , David Harwath

Text and vision foundation models can perform many tasks in a zero-shot setting, a desirable property that enables these systems to be applied in general and low-resource settings. There has been far less work, however, on the zero-shot…

计算与语言 · 计算机科学 2024-03-29 Rao Ma , Adian Liusie , Mark J. F. Gales , Kate M. Knill

Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these models is their…

声音 · 计算机科学 2024-10-14 Vyas Raina , Mark Gales

Automated Speech Recognition shows superhuman performance for adult English speech on a range of benchmarks, but disappoints when fed children's speech. This has long sat in the way of child-robot interaction. Recent evolutions in…

计算与语言 · 计算机科学 2024-11-20 Ruben Janssens , Eva Verhelst , Giulio Antonio Abbo , Qiaoqiao Ren , Maria Jose Pinto Bernal , Tony Belpaeme

Recognizing whispered speech and converting it to normal speech creates many possibilities for speech interaction. Because the sound pressure of whispered speech is significantly lower than that of normal speech, it can be used as a…

声音 · 计算机科学 2023-03-06 Jun Rekimoto

Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken…

音频与语音处理 · 电气工程与系统科学 2024-08-30 Mohan Li , Cong-Thanh Do , Simon Keizer , Youmna Farag , Svetlana Stoyanchev , Rama Doddipatla
‹ 上一页 1 2 3 10 下一页 ›