English
Related papers

Related papers: FAMA: The First Large-Scale Open-Science Speech Fo…

200 papers

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing speech FMs (SFMs) fall short of full compliance with the…

Computation and Language · Computer Science 2024-10-03 Marco Gaido , Sara Papi , Luisa Bentivogli , Alessio Brutti , Mauro Cettolo , Roberto Gretter , Marco Matassoni , Mohamed Nabih , Matteo Negri

The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a…

Computation and Language · Computer Science 2025-06-03 Yifan Peng , Shakeel Muhammad , Yui Sudo , William Chen , Jinchuan Tian , Chyi-Jiunn Lin , Shinji Watanabe

Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data. It generalizes well to various speech recognition and…

This paper presents Open Unified Speech Language Models (OpusLMs), a family of open foundational speech language models (SpeechLMs) up to 7B. Initialized from decoder-only text language models, the OpusLMs are continuously pre-trained on…

Recent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits. However, previous…

The Open Whisper-style Speech Model (OWSM) series was introduced to achieve full transparency in building advanced speech-to-text (S2T) foundation models. To this end, OWSM models are trained on 25 public speech datasets, which are…

Computation and Language · Computer Science 2024-06-14 Jinchuan Tian , Yifan Peng , William Chen , Kwanghee Choi , Karen Livescu , Shinji Watanabe

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Haorui He , Zengqiang Shang , Chaoren Wang , Xuyuan Li , Yicheng Gu , Hua Hua , Liwei Liu , Chen Yang , Jiaqi Li , Peiyang Shi , Yuancheng Wang , Kai Chen , Pengyuan Zhang , Zhizheng Wu

Despite being trained exclusively on speech data, speech foundation models (SFMs) like Whisper have shown impressive performance in non-speech tasks such as audio classification. This is partly because speech shares some common traits with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-17 Orchid Chetia Phukan , Swarup Ranjan Behera , Girish , Mohd Mujtaba Akhtar , Arun Balaji Buduru , Rajesh Sharma

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on…

In this paper, we introduce FAMMA, an open-source benchmark for \underline{f}in\underline{a}ncial \underline{m}ultilingual \underline{m}ultimodal question \underline{a}nswering (QA). Our benchmark aims to evaluate the abilities of large…

Computation and Language · Computer Science 2025-05-16 Siqiao Xue , Xiaojing Li , Fan Zhou , Qingyang Dai , Zhixuan Chu , Hongyuan Mei

Speech foundation models (SFMs), such as Open Whisper-Style Speech Models (OWSM), are trained on massive datasets to achieve accurate automatic speech recognition. However, even SFMs struggle to accurately recognize rare and unseen words.…

Sound · Computer Science 2025-06-12 Yui Sudo , Yusuke Fujita , Atsushi Kojima , Tomoya Mizumoto , Lianbo Liu

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to about one hundred languages which is a small fraction of the…

The development of Large Speech-Language Models (LSLMs) has been slowed by fragmented architectures and a lack of transparency, hindering the systematic comparison and reproducibility of research. Unlike in the vision-language domain, the…

Computation and Language · Computer Science 2025-08-22 Yirong Sun , Yizhong Geng , Peidong Wei , Yanjun Chen , Jinghan Yang , Rongfei Chen , Wei Zhang , Xiaoyu Shen

Large Language Models represent state-of-the-art linguistic models designed to equip computers with the ability to comprehend natural language. With its exceptional capacity to capture complex contextual relationships, the LLaMA (Large…

Computation and Language · Computer Science 2023-12-18 Pierpaolo Basile , Elio Musacchio , Marco Polignano , Lucia Siciliani , Giuseppe Fiameni , Giovanni Semeraro

We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets…

Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion…

Computation and Language · Computer Science 2026-01-19 Chin-Jou Li , Kalvin Chang , Shikhar Bharadwaj , Eunjung Yeo , Kwanghee Choi , Jian Zhu , David Mortensen , Shinji Watanabe

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality…

Computation and Language · Computer Science 2025-09-24 Run Luo , Ting-En Lin , Haonan Zhang , Yuchuan Wu , Xiong Liu , Min Yang , Yongbin Li , Longze Chen , Jiaming Li , Lei Zhang , Xiaobo Xia , Hamid Alinejad-Rokny , Fei Huang

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. However, most advanced…

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-18 Prabhat Pandey , Rupak Vignesh Swaminathan , K V Vijay Girish , Arunasish Sen , Jian Xie , Grant P. Strimel , Andreas Schwarz

Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of exploration on how…

Computation and Language · Computer Science 2025-03-04 Qingkai Fang , Shoutao Guo , Yan Zhou , Zhengrui Ma , Shaolei Zhang , Yang Feng
‹ Prev 1 2 3 10 Next ›