中文
相关论文

相关论文: A Survey on Spoken Italian Datasets and Corpora

200 篇论文

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale…

The widespread use of conversational and question answering systems made it necessary to improve the performances of speaker intent detection and understanding of related semantic slots, i.e., Spoken Language Understanding (SLU). Often,…

计算与语言 · 计算机科学 2019-07-18 Valentina Bellomaria , Giuseppe Castellucci , Andrea Favalli , Raniero Romagnoli

We present "Testimole-conversational" a massive collection of discussion boards messages in the Italian language. The large size of the corpus, more than 30B word-tokens (1996-2024), renders it an ideal dataset for native Italian Large…

计算与语言 · 计算机科学 2026-04-10 Matteo Rinaldi , Rossella Varvara , Viviana Patti

Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets…

计算与语言 · 计算机科学 2025-08-19 Michael Flor , Xinyi Liu , Anna Feldman

Italy exhibits rich linguistic diversity across its territory due to the distinct regional languages spoken in different areas. Recent advances in self-supervised learning provide new opportunities to analyze Italy's linguistic varieties…

计算与语言 · 计算机科学 2025-11-13 Moreno La Quatra , Alkis Koudounas , Elena Baralis , Sabato Marco Siniscalchi

During the past decade, several areas of speech and language understanding have witnessed substantial breakthroughs from the use of data-driven models. In the area of dialogue systems, the trend is less obvious, and most practical systems…

计算与语言 · 计算机科学 2017-03-22 Iulian Vlad Serban , Ryan Lowe , Peter Henderson , Laurent Charlin , Joelle Pineau

In the era of digital healthcare, the huge volumes of textual information generated every day in hospitals constitute an essential but underused asset that could be exploited with task-specific, fine-tuned biomedical language representation…

计算与语言 · 计算机科学 2023-07-12 Tommaso Mario Buonocore , Claudio Crema , Alberto Redolfi , Riccardo Bellazzi , Enea Parimbelli

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Haorui He , Zengqiang Shang , Chaoren Wang , Xuyuan Li , Yicheng Gu , Hua Hua , Liwei Liu , Chen Yang , Jiaqi Li , Peiyang Shi , Yuancheng Wang , Kai Chen , Pengyuan Zhang , Zhizheng Wu

This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root system that…

计算与语言 · 计算机科学 2024-02-29 Yang Liu , Jiahuan Cao , Chongyu Liu , Kai Ding , Lianwen Jin

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on…

Research in question answering datasets and models has gained a lot of attention in the research community. Many of them release their own question answering datasets as well as the models. There is tremendous progress that we have seen in…

计算与语言 · 计算机科学 2021-12-28 Andreas Chandra , Affandy Fahrizain , Ibrahim , Simon Willyanto Laufried

Computational historical linguistics seeks to systematically understand processes of sound change, including during periods at which little to no formal recording of language is attested. At the same time, few computational resources exist…

计算与语言 · 计算机科学 2024-04-26 Stephen Bothwell , Brian DuSell , David Chiang , Brian Krostenko

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 million native speakers, Cantonese is still considered a…

Collecting high-quality training data is essential for fine-tuning Large Language Models (LLMs). However, acquiring such data is often costly and time-consuming, especially for non-English languages such as Italian. Recently, researchers…

计算与语言 · 计算机科学 2025-04-01 Fatemeh Mohammadi , Tommaso Romano , Samira Maghool , Paolo Ceravolo

Italy is characterized by a one-of-a-kind linguistic diversity landscape in Europe, which implicitly encodes local knowledge, cultural traditions, artistic expressions and history of its speakers. However, most local languages and dialects…

计算与语言 · 计算机科学 2023-11-21 Alan Ramponi

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted…

This paper presents Fauno, the first and largest open-source Italian conversational Large Language Model (LLM). Our goal with Fauno is to democratize the study of LLMs in Italian, demonstrating that obtaining a fine-tuned conversational bot…

计算与语言 · 计算机科学 2023-06-27 Andrea Bacciu , Giovanni Trappolini , Andrea Santilli , Emanuele Rodolà , Fabrizio Silvestri

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional…

计算与语言 · 计算机科学 2018-04-11 Pete Warden

Making sure that users understand privacy policies that impact them is a key challenge for a real GDPR deployment. Research studies are mostly carried in English, but in Europe and elsewhere, users speak a language that is not English.…

密码学与安全 · 计算机科学 2023-02-13 Francesco Ciclosi , Silvia Vidor , Fabio Massacci

This paper presents an analysis of the distribution of spoken language in the V3C video retrieval benchmark dataset based on automatically generated transcripts. It finds that a large portion of the dataset is covered by spoken language.…

多媒体 · 计算机科学 2022-12-16 Luca Rossetto
‹ 上一页 1 2 3 10 下一页 ›