中文
相关论文

相关论文: A Survey on Spoken Italian Datasets and Corpora

200 篇论文

The purpose of this project was to derive a reliable estimate of the frequency of occurrence of the 30 phonemes - plus consonant geminated counterparts - of the Italian language, based on four selected written texts. Since no comparable…

音频与语音处理 · 电气工程与系统科学 2021-01-19 Javi Arango , Alec DeCaprio , Sunwoo Baik , Luca De Nardis , Stefanie Shattuck-Hufnagel , Maria Gabriella Di Benedetto

Automatic summarization has consistently attracted attention due to its versatility and wide application in various downstream tasks. Despite its popularity, we find that annotation efforts have largely been disjointed, and have lacked…

计算与语言 · 计算机科学 2025-02-12 Noam Dahan , Gabriel Stanovsky

Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these models often exhibit…

计算与语言 · 计算机科学 2025-05-22 Yi-Cheng Lin , Wei-Chih Chen , Hung-yi Lee

Spoken Dialogue Models (SDMs) have recently attracted significant attention for their ability to generate voice responses directly to users' spoken queries. Despite their increasing popularity, there exists a gap in research focused on…

计算与语言 · 计算机科学 2025-10-07 Chengqian Ma , Wei Tao , Yiwen Guo

Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to generate and manipulate human language, highlighting their potential across various applications. Evaluating LLMs in languages other than…

计算与语言 · 计算机科学 2024-06-26 Fabio Mercorio , Mario Mezzanzanica , Daniele Potertì , Antonio Serino , Andrea Seveso

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and…

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

声音 · 计算机科学 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

Sentiment analysis, the automated process of determining emotions or opinions expressed in text, has seen extensive exploration in the field of natural language processing. However, one aspect that has remained underrepresented is the…

计算与语言 · 计算机科学 2024-09-16 Mouad Jbel , Mourad Jabrane , Imad Hafidi , Abdulmutallib Metrane

The recent emergence and adoption of Machine Learning technology, and specifically of Large Language Models, has drawn attention to the need for systematic and transparent management of language data. This work proposes an approach to…

Multilingual semantic parsing is a cost-effective method that allows a single model to understand different languages. However, researchers face a great imbalance of availability of training data, with English being resource rich, and other…

计算与语言 · 计算机科学 2021-06-15 Menglin Xia , Emilio Monti

In this work we present SignIT, a new dataset to study the task of Italian Sign Language (LIS) recognition. The dataset is composed of 644 videos covering 3.33 hours. We manually annotated videos considering a taxonomy of 94 distinct sign…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Alessia Micieli , Giovanni Maria Farinella , Francesco Ragusa

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED,…

声音 · 计算机科学 2021-02-04 Weiquan Fan , Xiangmin Xu , Xiaofen Xing , Weidong Chen , Dongyan Huang

This paper analyses language modeling in spoken dialogue systems for accessing a database. The use of several language models obtained by exploiting dialogue predictions gives better results than the use of a single model for the whole…

cmp-lg · 计算机科学 2008-02-03 Cosmin Popovici , Paolo Baggia

Open conversations are one of the most engaging forms of teaching. However, creating those conversations in educational software is a complex endeavor, especially if we want to address the needs of different audiences. While language models…

计算与语言 · 计算机科学 2024-04-17 Donya Rooein , Dirk Hovy

The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which are used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings. The paper…

计算与语言 · 计算机科学 2024-03-21 Michal Mochtak , Peter Rupnik , Nikola Ljubešić

Current research in machine learning and artificial intelligence is largely centered on modeling and performance evaluation, less so on data collection. However, recent research demonstrated that limitations and biases in data may…

人工智能 · 计算机科学 2025-02-18 Eleonora Mancini , Ana Tanevska , Andrea Galassi , Alessio Galatolo , Federico Ruggeri , Paolo Torroni

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected. One reason for this is that evaluation datasets do not…

计算与语言 · 计算机科学 2024-06-05 Chunlan Ma , Ayyoob ImaniGooghari , Haotian Ye , Renhao Pei , Ehsaneddin Asgari , Hinrich Schütze

Software development relies heavily on text-based communication, making sentiment analysis a valuable tool for understanding team dynamics and supporting trustworthy AI-driven analytics in requirements engineering. However, existing…

软件工程 · 计算机科学 2025-07-11 Martin Obaidi , Marc Herrmann , Jil Klünder , Kurt Schneider

Scaling semantic parsing models for task-oriented dialog systems to new languages is often expensive and time-consuming due to the lack of available datasets. Available datasets suffer from several shortcomings: a) they contain few…

计算与语言 · 计算机科学 2021-01-28 Haoran Li , Abhinav Arora , Shuohui Chen , Anchit Gupta , Sonal Gupta , Yashar Mehdad

We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information…

计算与语言 · 计算机科学 2025-09-09 Jinrui Yang , Timothy Baldwin , Trevor Cohn