English
Related papers

Related papers: A Survey on Spoken Italian Datasets and Corpora

200 papers

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale…

Computation and Language · Computer Science 2024-06-25 Alkis Koudounas , Moreno La Quatra , Lorenzo Vaiani , Luca Colomba , Giuseppe Attanasio , Eliana Pastor , Luca Cagliero , Elena Baralis

The widespread use of conversational and question answering systems made it necessary to improve the performances of speaker intent detection and understanding of related semantic slots, i.e., Spoken Language Understanding (SLU). Often,…

Computation and Language · Computer Science 2019-07-18 Valentina Bellomaria , Giuseppe Castellucci , Andrea Favalli , Raniero Romagnoli

We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large…

Computation and Language · Computer Science 2026-04-10 Matteo Rinaldi , Rossella Varvara , Viviana Patti

Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets…

Computation and Language · Computer Science 2025-08-19 Michael Flor , Xinyi Liu , Anna Feldman

Italy exhibits rich linguistic diversity across its territory due to the distinct regional languages spoken in different areas. Recent advances in self-supervised learning provide new opportunities to analyze Italy's linguistic varieties…

Computation and Language · Computer Science 2025-11-13 Moreno La Quatra , Alkis Koudounas , Elena Baralis , Sabato Marco Siniscalchi

During the past decade, several areas of speech and language understanding have witnessed substantial breakthroughs from the use of data-driven models. In the area of dialogue systems, the trend is less obvious, and most practical systems…

Computation and Language · Computer Science 2017-03-22 Iulian Vlad Serban , Ryan Lowe , Peter Henderson , Laurent Charlin , Joelle Pineau

In the era of digital healthcare, the huge volumes of textual information generated every day in hospitals constitute an essential but underused asset that could be exploited with task-specific, fine-tuned biomedical language representation…

Computation and Language · Computer Science 2023-07-12 Tommaso Mario Buonocore , Claudio Crema , Alberto Redolfi , Riccardo Bellazzi , Enea Parimbelli

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Haorui He , Zengqiang Shang , Chaoren Wang , Xuyuan Li , Yicheng Gu , Hua Hua , Liwei Liu , Chen Yang , Jiaqi Li , Peiyang Shi , Yuancheng Wang , Kai Chen , Pengyuan Zhang , Zhizheng Wu

This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root system that…

Computation and Language · Computer Science 2024-02-29 Yang Liu , Jiahuan Cao , Chongyu Liu , Kai Ding , Lianwen Jin

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on…

Research in question answering datasets and models has gained a lot of attention in the research community. Many of them release their own question answering datasets as well as the models. There is tremendous progress that we have seen in…

Computation and Language · Computer Science 2021-12-28 Andreas Chandra , Affandy Fahrizain , Ibrahim , Simon Willyanto Laufried

Computational historical linguistics seeks to systematically understand processes of sound change, including during periods at which little to no formal recording of language is attested. At the same time, few computational resources exist…

Computation and Language · Computer Science 2024-04-26 Stephen Bothwell , Brian DuSell , David Chiang , Brian Krostenko

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a…

Computation and Language · Computer Science 2025-03-06 Jiyue Jiang , Alfred Kar Yin Truong , Yanyu Chen , Qinghang Bao , Sheng Wang , Pengan Chen , Jiuming Wang , Lingpeng Kong , Yu Li , Chuan Wu

Collecting high-quality training data is essential for fine-tuning Large Language Models (LLMs). However, acquiring such data is often costly and time-consuming, especially for non-English languages such as Italian. Recently, researchers…

Computation and Language · Computer Science 2025-04-01 Fatemeh Mohammadi , Tommaso Romano , Samira Maghool , Paolo Ceravolo

Italy is characterized by a one-of-a-kind linguistic diversity landscape in Europe, which implicitly encodes local knowledge, cultural traditions, artistic expressions and history of its speakers. However, most local languages and dialects…

Computation and Language · Computer Science 2023-11-21 Alan Ramponi

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted…

This paper presents Fauno, the first and largest open-source Italian conversational Large Language Model (LLM). Our goal with Fauno is to democratize the study of LLMs in Italian, demonstrating that obtaining a fine-tuned conversational bot…

Computation and Language · Computer Science 2023-06-27 Andrea Bacciu , Giovanni Trappolini , Andrea Santilli , Emanuele Rodolà , Fabrizio Silvestri

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional…

Computation and Language · Computer Science 2018-04-11 Pete Warden

Making sure that users understand privacy policies that impact them is a key challenge for a real GDPR deployment. Research studies are mostly carried in English, but in Europe and elsewhere, users speak a language that is not English.…

Cryptography and Security · Computer Science 2023-02-13 Francesco Ciclosi , Silvia Vidor , Fabio Massacci

This paper presents an analysis of the distribution of spoken language in the V3C video retrieval benchmark dataset based on automatically generated transcripts. It finds that a large portion of the dataset is covered by spoken language.…

Multimedia · Computer Science 2022-12-16 Luca Rossetto
‹ Prev 1 2 3 10 Next ›