中文
相关论文

相关论文: FAMA: The First Large-Scale Open-Science Speech Fo…

200 篇论文

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing speech FMs (SFMs) fall short of full compliance with the…

The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a…

计算与语言 · 计算机科学 2025-06-03 Yifan Peng , Shakeel Muhammad , Yui Sudo , William Chen , Jinchuan Tian , Chyi-Jiunn Lin , Shinji Watanabe

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and…

This paper presents Open Unified Speech Language Models (OpusLMs), a family of open foundational speech language models (SpeechLMs) up to 7B. Initialized from decoder-only text language models, the OpusLMs are continuously pre-trained on…

Recent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits. However, previous…

The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are…

计算与语言 · 计算机科学 2024-06-14 Jinchuan Tian , Yifan Peng , William Chen , Kwanghee Choi , Karen Livescu , Shinji Watanabe

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Haorui He , Zengqiang Shang , Chaoren Wang , Xuyuan Li , Yicheng Gu , Hua Hua , Liwei Liu , Chen Yang , Jiaqi Li , Peiyang Shi , Yuancheng Wang , Kai Chen , Pengyuan Zhang , Zhizheng Wu

Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with…

音频与语音处理 · 电气工程与系统科学 2024-10-17 Orchid Chetia Phukan , Swarup Ranjan Behera , Girish , Mohd Mujtaba Akhtar , Arun Balaji Buduru , Rajesh Sharma

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on…

In this paper, we introduce FAMMA, an open-source benchmark for \underline{f}in\underline{a}ncial \underline{m}ultilingual \underline{m}ultimodal question \underline{a}nswering (QA). Our benchmark aims to evaluate the abilities of large…

计算与语言 · 计算机科学 2025-05-16 Siqiao Xue , Xiaojing Li , Fan Zhou , Qingyang Dai , Zhixuan Chu , Hongyuan Mei

Speech foundation models (SFMs), such as Open Whisper-Style Speech Models (OWSM), are trained on massive datasets to achieve accurate automatic speech recognition. However, even SFMs struggle to accurately recognize rare and unseen words.…

声音 · 计算机科学 2025-06-12 Yui Sudo , Yusuke Fujita , Atsushi Kojima , Tomoya Mizumoto , Lianbo Liu

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the…

The development of Large Speech-Language Models (LSLMs) has been slowed by fragmented architectures and a lack of transparency, hindering the systematic comparison and reproducibility of research. Unlike in the vision-language domain, the…

计算与语言 · 计算机科学 2025-08-22 Yirong Sun , Yizhong Geng , Peidong Wei , Yanjun Chen , Jinghan Yang , Rongfei Chen , Wei Zhang , Xiaoyu Shen

Large Language Models represent state-of-the-art linguistic models designed to equip computers with the ability to comprehend natural language. With its exceptional capacity to capture complex contextual relationships, the LLaMA (Large…

计算与语言 · 计算机科学 2023-12-18 Pierpaolo Basile , Elio Musacchio , Marco Polignano , Lucia Siciliani , Giuseppe Fiameni , Giovanni Semeraro

We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets…

Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion…

计算与语言 · 计算机科学 2026-01-19 Chin-Jou Li , Kalvin Chang , Shikhar Bharadwaj , Eunjung Yeo , Kwanghee Choi , Jian Zhu , David Mortensen , Shinji Watanabe

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality…

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. However, most advanced…

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which…

音频与语音处理 · 电气工程与系统科学 2025-04-18 Prabhat Pandey , Rupak Vignesh Swaminathan , K V Vijay Girish , Arunasish Sen , Jian Xie , Grant P. Strimel , Andreas Schwarz

Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how…

计算与语言 · 计算机科学 2025-03-04 Qingkai Fang , Shoutao Guo , Yan Zhou , Zhengrui Ma , Shaolei Zhang , Yang Feng
‹ 上一页 1 2 3 10 下一页 ›