中文
相关论文

相关论文: Prompting the Hidden Talent of Web-Scale Speech Mo…

200 篇论文

We equip a smaller Language Model to generalise to answering challenging compositional questions that have not been seen in training. To do so we propose a combination of multitask supervised pretraining on up to 93 tasks designed to…

计算与语言 · 计算机科学 2023-08-22 Tim Hartill , Neset Tan , Michael Witbrock , Patricia J. Riddle

Recent advances in automatic speech recognition (ASR) and speech enhancement have led to a widespread assumption that improving perceptual audio quality should directly benefit recognition accuracy. In this work, we rigorously examine…

声音 · 计算机科学 2026-03-06 Akif Islam , Raufun Nahar , Md. Ekramul Hamid

Speaker verification is a task of confirming an individual's identity through the analysis of their voice. Whispered speech differs from phonated speech in acoustic characteristics, which degrades the performance of speaker verification…

声音 · 计算机科学 2026-05-08 Magdalena Gołębiowska , Piotr Syga

End-to-end automatic speech recognition (ASR) and large language models, such as Whisper and GPT-2, have recently been scaled to use vast amounts of training data. Despite the large amount of training data, infrequent content words that…

计算与语言 · 计算机科学 2023-06-06 Guangzhi Sun , Xianrui Zheng , Chao Zhang , Philip C. Woodland

Language models can be prompted to perform a wide variety of zero- and few-shot learning problems. However, performance varies significantly with the choice of prompt, and we do not yet understand why this happens or how to pick the best…

计算与语言 · 计算机科学 2024-09-16 Hila Gonen , Srini Iyer , Terra Blevins , Noah A. Smith , Luke Zettlemoyer

This paper proposes AS-ASR, a lightweight aphasia-specific speech recognition framework based on Whisper-tiny, tailored for low-resource deployment on edge devices. Our approach introduces a hybrid training strategy that systematically…

音频与语音处理 · 电气工程与系统科学 2026-02-03 Chen Bao , Chuanbing Huo , Qinyu Chen , Chang Gao

Current speech encoding pipelines often rely on an additional text-based LM to get robust representations of human communication, even though SotA speech-to-text models often have a LM within. This work proposes an approach to improve the…

Pre-trained multilingual speech foundation models, like Whisper, have shown impressive performance across different languages. However, adapting these models to new or specific languages is computationally extensive and faces catastrophic…

计算与语言 · 计算机科学 2024-08-21 Tianyi Xu , Kaixun Huang , Pengcheng Guo , Yu Zhou , Longtao Huang , Hui Xue , Lei Xie

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms…

音频与语音处理 · 电气工程与系统科学 2024-04-11 Ziyue Jiang , Jinglin Liu , Yi Ren , Jinzheng He , Zhenhui Ye , Shengpeng Ji , Qian Yang , Chen Zhang , Pengfei Wei , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself. As such, it is unclear how well speech-based retrieval can work in practice -- both in an absolute…

计算与语言 · 计算机科学 2021-06-16 Ramon Sanabria , Austin Waters , Jason Baldridge

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great potential to employ…

音频与语音处理 · 电气工程与系统科学 2023-12-15 Yifan Yang , Feiyu Shen , Chenpeng Du , Ziyang Ma , Kai Yu , Daniel Povey , Xie Chen

In the landscape of modern machine learning, frozen pre-trained models provide stability and efficiency but often underperform on specific tasks due to mismatched data distributions. This paper introduces the Whisperer, a novel visual…

机器学习 · 计算机科学 2026-03-06 Samandar Samandarov , Nazirjon Ismoiljonov , Abdullah Sattorov , Temirlan Sabyrbayev

We present an efficient end-to-end approach for holistic Automatic Speaking Assessment (ASA) of multi-part second-language tests, developed for the 2025 Speak & Improve Challenge. Our system's main novelty is the ability to process all four…

计算与语言 · 计算机科学 2025-10-07 Nhan Phan , Anusha Porwal , Yaroslav Getman , Ekaterina Voskoboinik , Tamás Grósz , Mikko Kurimo

Automatic Speech Recognition systems have made significant progress with large-scale pre-trained models. However, most current systems focus solely on transcribing the speech without identifying speaker roles, a function that is critical…

音频与语音处理 · 电气工程与系统科学 2025-06-13 Anfeng Xu , Tiantian Feng , Shrikanth Narayanan

Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to…

音频与语音处理 · 电气工程与系统科学 2025-05-08 Jinpeng Li , Yu Pu , Qi Sun , Wei-Qiang Zhang

Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions remains a challenge.…

计算与语言 · 计算机科学 2024-08-16 Mohamed Osman , Daniel Z. Kaplan , Tamer Nadeem

In this paper, we present empirical studies on conversational recommendation tasks using representative large language models in a zero-shot setting with three primary contributions. (1) Data: To gain insights into model behavior in…

Contextual automatic speech recognition (ASR) systems allow for recognizing out-of-vocabulary (OOV) words, such as named entities or rare words. However, it remains challenging due to limited training data and ambiguous or inconsistent…

计算与语言 · 计算机科学 2025-09-03 Changsong Liu , Yizhou Peng , Eng Siong Chng

Large-scale auto-regressive language models pretrained on massive text have demonstrated their impressive ability to perform new natural language tasks with only a few text examples, without the need for fine-tuning. Recent studies further…

音频与语音处理 · 电气工程与系统科学 2022-04-15 Heting Gao , Junrui Ni , Kaizhi Qian , Yang Zhang , Shiyu Chang , Mark Hasegawa-Johnson

Despite the growing advancements in Automatic Speech Recognition (ASR) models, the development of robust models for underrepresented languages, such as Nepali, remains a challenge. This research focuses on making an exhaustive and…

计算与语言 · 计算机科学 2024-11-20 Sanjay Rijal , Shital Adhikari , Manish Dahal , Manish Awale , Vaghawan Ojha