中文
相关论文

相关论文: Emilia: A Large-Scale, Extensive, Multilingual, an…

200 篇论文

Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of spoken data, the…

音频与语音处理 · 电气工程与系统科学 2025-03-17 Nikola Ljubešić , Peter Rupnik , Danijel Koržinek

We introduce API Pack, a massive multi-programming language dataset containing over one million instruction-API calls for improving the API call generation capabilities of large language models. Our evaluation highlights three key findings:…

计算与语言 · 计算机科学 2025-02-17 Zhen Guo , Adriana Meza Soria , Wei Sun , Yikang Shen , Rameswar Panda

Talk2AI is a large-scale longitudinal dataset of 3,080 conversations (totaling 30,800 turns) between human participants and Large Language Models (LLMs), designed to support research on persuasion, opinion change, and human-AI interaction.…

With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully capturing the rich…

计算与语言 · 计算机科学 2025-09-18 Haoyu Wang , Guangyan Zhang , Jiale Chen , Jingyu Li , Yuehai Wang , Yiwen Guo

This paper introduces Multilingual LibriSpeech (MLS) dataset, a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages, including about 44.5K hours of…

音频与语音处理 · 电气工程与系统科学 2020-12-22 Vineel Pratap , Qiantong Xu , Anuroop Sriram , Gabriel Synnaeve , Ronan Collobert

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of…

The GPT-4o represents a significant milestone in enabling real-time interaction with large language models (LLMs) through speech, its remarkable low latency and high fluency not only capture attention but also stimulate research interest in…

计算与语言 · 计算机科学 2024-12-04 Shuaijiang Zhao , Tingwei Guo , Bajian Xiang , Tongtang Wan , Qiang Niu , Wei Zou , Xiangang Li

Current pre-trained language models (PLM) are typically trained with static data, ignoring that in real-world scenarios, streaming data of various sources may continuously grow. This requires PLMs to integrate the information from all the…

计算与语言 · 计算机科学 2022-07-12 Yujia Qin , Jiajie Zhang , Yankai Lin , Zhiyuan Liu , Peng Li , Maosong Sun , Jie Zhou

The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient. This work enhances OWSM by integrating YODAS, a…

计算与语言 · 计算机科学 2025-06-03 Yifan Peng , Shakeel Muhammad , Yui Sudo , William Chen , Jinchuan Tian , Chyi-Jiunn Lin , Shinji Watanabe

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

High-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations. Human recordings incur high costs and privacy concerns, while synthetic approaches often lack…

计算与语言 · 计算机科学 2025-04-01 Minghan Wang , Ye Bai , Yuxia Wang , Thuy-Trang Vu , Ehsan Shareghi , Gholamreza Haffari

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Bingsong Bai , Qihang Lu , Wenbing Yang , Zihan Sun , Yueran Hou , Peilei Jia , Songbai Pu , Ruibo Fu , Yingming Gao , Ya Li , Jun Gao

In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. Addressing the scarcity of samples containing both…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Vahid Noroozi , Zhehuai Chen , Somshubra Majumdar , Steve Huang , Jagadeesh Balam , Boris Ginsburg

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale…

With the emergence of increasingly powerful large language models, there is a burgeoning interest in leveraging these models for casual conversation and role-play applications. However, existing conversational and role-playing datasets…

计算与语言 · 计算机科学 2023-08-14 Tear Gosling , Alpin Dale , Yinhe Zheng

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

声音 · 计算机科学 2024-02-08 Dan Lyth , Simon King

Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more…

计算与语言 · 计算机科学 2026-04-13 Zhepeng Cen , Haolin Chen , Shiyu Wang , Zuxin Liu , Zhiwei Liu , Jielin Qiu , Ding Zhao , Silvio Savarese , Caiming Xiong , Huan Wang , Weiran Yao

Research on large language models (LLMs) has shown remarkable performance in domains such as mathematics, programming, and literary creation. However, most studies have focused on semantic memory-based question answering, neglecting LLMs'…

计算与语言 · 计算机科学 2025-02-25 WenTao Liu , Ruohua Zhang , Aimin Zhou , Feng Gao , JiaLi Liu

Objectives: Despite the recent adoption of large language models (LLMs) for biomedical information extraction, challenges in prompt engineering and algorithms persist, with no dedicated software available. To address this, we developed…

机器学习 · 计算机科学 2025-04-02 Enshuo Hsu , Kirk Roberts

This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and…