中文
相关论文

相关论文: Emilia: A Large-Scale, Extensive, Multilingual, an…

200 篇论文

Personalized dialogue systems are an essential step toward better human-machine interaction. Existing personalized dialogue agents rely on properly designed conversational datasets, which are mostly monolingual (e.g., English), which…

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

声音 · 计算机科学 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks, yet the majority of high-performing models remain closed-source or partially open, limiting transparency and reproducibility. In this work,…

At the beginning era of large language model, it is quite critical to generate a high-quality financial dataset to fine-tune a large language model for financial related tasks. Thus, this paper presents a carefully designed data creation…

计算与语言 · 计算机科学 2023-08-04 Ziao Wang , Jianning Wang , Junda Wu , Xiaofeng Zhang

Large language models (LLMs) have great potential for synthetic data generation. This work shows that useful data can be synthetically generated even for tasks that cannot be solved directly by LLMs: for problems with structured outputs, it…

计算与语言 · 计算机科学 2023-10-31 Martin Josifoski , Marija Sakota , Maxime Peyrard , Robert West

Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent…

计算与语言 · 计算机科学 2026-02-25 Dilara Torunoğlu-Selamet , Dogukan Arslan , Rodrigo Wilkens , Wei He , Doruk Eryiğit , Thomas Pickard , Adriana S. Pagano , Aline Villavicencio , Gülşen Eryiğit , Ágnes Abuczki , Aida Cardoso , Alesia Lazarenka , Dina Almassova , Amalia Mendes , Anna Kanellopoulou , Antoni Brosa-Rodríguez , Baiba Saulite , Beata Wojtowicz , Bolette Pedersen , Carlos Manuel Hidalgo-Ternero , Chaya Liebeskind , Danka Jokić , Diego Alves , Eleni Triantafyllidi , Erik Velldal , Fred Philippy , Giedre Valunaite Oleskeviciene , Ieva Rizgeliene , Inguna Skadina , Irina Lobzhanidze , Isabell Stinessen Haugen , Jauza Akbar Krito , Jelena M. Marković , Johanna Monti , Josue Alejandro Sauca , Kaja Dobrovoljc , Kingsley O. Ugwuanyi , Laura Rituma , Lilja Øvrelid , Maha Tufail Agro , Manzura Abjalova , Maria Chatzigrigoriou , María del Mar Sánchez Ramos , Marija Pendevska , Masoumeh Seyyedrezaei , Mehrnoush Shamsfard , Momina Ahsan , Muhammad Ahsan Riaz Khan , Nathalie Carmen Hau Norman , Nilay Erdem Ayyıldız , Nina Hosseini-Kivanani , Noémi Ligeti-Nagy , Numaan Naeem , Olha Kanishcheva , Olha Yatsyshyna , Daniil Orel , Petra Giommarelli , Petya Osenova , Radovan Garabik , Regina E. Semou , Rozane Rebechi , Salsabila Zahirah Pranida , Samia Touileb , Sanni Nimb , Sarfraz Ahmad , Sarvinoz Sharipova , Shahar Golan , Shaoxiong Ji , Sopuruchi Christian Aboh , Srdjan Sucur , Stella Markantonatou , Sussi Olsen , Vahide Tajalli , Veronika Lipp , Voula Giouli , Yelda Yeşildal Eraydın , Zahra Saaberi , Zhuohan Xie

Cross-lingual open-ended generation - responding in a language different from that of the query - is an important yet understudied problem. This work proposes XL-Instruct, a novel technique for generating high-quality synthetic data, and…

计算与语言 · 计算机科学 2025-09-30 Vivek Iyer , Pinzhen Chen , Ricardo Rei , Alexandra Birch

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximately 36.7k hours of English speech for 22.05 kHz training,…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Ryan Langman , Xuesong Yang , Paarth Neekhara , Shehzeen Hussain , Edresson Casanova , Evelina Bakhturina , Jason Li

Large language models (LLMs) have proven to be effective tools for a wide range of natural language processing (NLP) applications. Although many LLMs are multilingual, most remain English-centric and perform poorly on low-resource…

计算与语言 · 计算机科学 2026-02-03 Tan Sang Nguyen , Muhammad Reza Qorib , Hwee Tou Ng

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works,…

音频与语音处理 · 电气工程与系统科学 2026-04-22 Gaspard Michel , Elena V. Epure , Christophe Cerisara

Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such as high costs in voice data collection, weakness in dynamic…

计算与语言 · 计算机科学 2025-02-19 Ailin Huang , Boyong Wu , Bruce Wang , Chao Yan , Chen Hu , Chengli Feng , Fei Tian , Feiyu Shen , Jingbei Li , Mingrui Chen , Peng Liu , Ruihang Miao , Wang You , Xi Chen , Xuerui Yang , Yechang Huang , Yuxiang Zhang , Zheng Gong , Zixin Zhang , Hongyu Zhou , Jianjian Sun , Brian Li , Chengting Feng , Changyi Wan , Hanpeng Hu , Jianchang Wu , Jiangjie Zhen , Ranchen Ming , Song Yuan , Xuelin Zhang , Yu Zhou , Bingxin Li , Buyun Ma , Hongyuan Wang , Kang An , Wei Ji , Wen Li , Xuan Wen , Xiangwen Kong , Yuankai Ma , Yuanwei Liang , Yun Mou , Bahtiyar Ahmidi , Bin Wang , Bo Li , Changxin Miao , Chen Xu , Chenrun Wang , Dapeng Shi , Deshan Sun , Dingyuan Hu , Dula Sai , Enle Liu , Guanzhe Huang , Gulin Yan , Heng Wang , Haonan Jia , Haoyang Zhang , Jiahao Gong , Junjing Guo , Jiashuai Liu , Jiahong Liu , Jie Feng , Jie Wu , Jiaoren Wu , Jie Yang , Jinguo Wang , Jingyang Zhang , Junzhe Lin , Kaixiang Li , Lei Xia , Li Zhou , Liang Zhao , Longlong Gu , Mei Chen , Menglin Wu , Ming Li , Mingxiao Li , Mingliang Li , Mingyao Liang , Na Wang , Nie Hao , Qiling Wu , Qinyuan Tan , Ran Sun , Shuai Shuai , Shaoliang Pang , Shiliang Yang , Shuli Gao , Shanshan Yuan , Siqi Liu , Shihong Deng , Shilei Jiang , Sitong Liu , Tiancheng Cao , Tianyu Wang , Wenjin Deng , Wuxun Xie , Weipeng Ming , Wenqing He , Wen Sun , Xin Han , Xin Huang , Xiaomin Deng , Xiaojia Liu , Xin Wu , Xu Zhao , Yanan Wei , Yanbo Yu , Yang Cao , Yangguang Li , Yangzhen Ma , Yanming Xu , Yaoyu Wang , Yaqiang Shi , Yilei Wang , Yizhuang Zhou , Yinmin Zhong , Yang Zhang , Yaoben Wei , Yu Luo , Yuanwei Lu , Yuhe Yin , Yuchu Luo , Yuanhao Ding , Yuting Yan , Yaqi Dai , Yuxiang Yang , Zhe Xie , Zheng Ge , Zheng Sun , Zhewei Huang , Zhichao Chang , Zhisheng Guan , Zidong Yang , Zili Zhang , Binxing Jiao , Daxin Jiang , Heung-Yeung Shum , Jiansheng Chen , Jing Li , Shuchang Zhou , Xiangyu Zhang , Xinhao Zhang , Yibo Zhu

Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods continue to suffer from residual interference in complex…

声音 · 计算机科学 2026-02-02 Kai Li , Jintao Cheng , Chang Zeng , Zijun Yan , Helin Wang , Zixiong Su , Bo Zheng , Xiaolin Hu

Modern language models (LMs) must be trained on many orders of magnitude more words of training data than human children receive before they begin to produce useful behavior. Assessing the nature and origins of this "data gap" requires…

计算与语言 · 计算机科学 2026-04-01 Steven Y. Feng , Alvin W. M. Tan , Michael C. Frank

Large language models (LLMs) serve as powerful tools for design, providing capabilities for both task automation and design assistance. Recent advancements have shown tremendous potential for facilitating LLM integration into the chip…

Whisper generation is constrained by the difficulty of data collection. Because whispered speech has low acoustic amplitude, high-fidelity recording is challenging. In this paper, we introduce WhispSynth, a large-scale multilingual corpus…

声音 · 计算机科学 2026-03-17 Tianyi Tan , Jiaxin Ye , Yuanming Zhang , Xiaohuai Le , Xianjun Xia , Chuanzeng Huang , Jing Lu

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based…

计算与语言 · 计算机科学 2025-08-08 Wenqian Cui , Dianzhi Yu , Xiaoqi Jiao , Ziqiao Meng , Guangyan Zhang , Qichao Wang , Yiwen Guo , Irwin King

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged as a promising alternative to the autoregressive paradigm,…

This paper presents LOLA, a massively multilingual large language model trained on more than 160 languages using a sparse Mixture-of-Experts Transformer architecture. Our architectural and implementation choices address the challenge of…

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations.…

计算与语言 · 计算机科学 2025-09-08 Wei Chu , Yuanzhe Dong , Ke Tan , Dong Han , Xavier Menendez-Pidal , Ruchao Fan , Chenfeng Miao , Chanwoo Kim , Bhiksha Raj , Rita Singh

Text-to-speech (TTS) development is limited by scarcity of high-quality, publicly available speech data for most languages outside a few high-resource languages. We present Nord-Parl-TTS, an open TTS dataset for Finnish and Swedish based on…

音频与语音处理 · 电气工程与系统科学 2026-02-10 Zirui Li , Jens Edlund , Yicheng Gu , Nhan Phan , Lauri Juvela , Mikko Kurimo