中文
相关论文

相关论文: A Survey on Spoken Italian Datasets and Corpora

200 篇论文

Conversational search is a relatively young area of research that aims at automating an information-seeking dialogue. In this paper we help to position it with respect to other research areas within conversational Artificial Intelligence…

信息检索 · 计算机科学 2021-06-09 Svitlana Vakulenko , Evangelos Kanoulas , Maarten de Rijke

While Italian is a high-resource language, there are few Italian-native benchmarks to evaluate generative Large Language Models (LLMs) in this language. This work presents three new benchmarks: Invalsi MATE to evaluate models performance on…

计算与语言 · 计算机科学 2024-09-18 Giovanni Puccetti , Maria Cassese , Andrea Esuli

Subtitles are essential for video accessibility and audience engagement. Modern Automatic Speech Recognition (ASR) systems, built upon Encoder-Decoder neural network architectures and trained on massive amounts of data, have progressively…

计算与语言 · 计算机科学 2025-12-23 Alessandro Lucca , Francesco Pierri

The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English,…

Progress in Machine Learning is often driven by the availability of large datasets, and consistent evaluation metrics for comparing modeling approaches. To this end, we present a repository of conversational datasets consisting of hundreds…

Almost none of the 2,000+ languages spoken in Africa have widely available automatic speech recognition systems, and the required data is also only available for a few languages. We have experimented with two techniques which may provide…

计算与语言 · 计算机科学 2022-10-05 Sandy Ritchie , You-Chi Cheng , Mingqing Chen , Rajiv Mathews , Daan van Esch , Bo Li , Khe Chai Sim

Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization remains loosely defined. The…

计算与语言 · 计算机科学 2025-10-20 Fabian Retkowski , Maike Züfle , Andreas Sudmann , Dinah Pfau , Shinji Watanabe , Jan Niehues , Alexander Waibel

The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because…

Shopping online is more and more frequent in our everyday life. For e-commerce search systems, understanding natural language coming through voice assistants, chatbots or from conversational search is an essential ability to understand what…

信息检索 · 计算机科学 2023-02-14 Andrea Papenmeier , Dagmar Kern , Daniel Hienert , Alfred Sliwa , Ahmet Aker , Norbert Fuhr

Large Language Models represent state-of-the-art linguistic models designed to equip computers with the ability to comprehend natural language. With its exceptional capacity to capture complex contextual relationships, the LLaMA (Large…

计算与语言 · 计算机科学 2023-12-18 Pierpaolo Basile , Elio Musacchio , Marco Polignano , Lucia Siciliani , Giuseppe Fiameni , Giovanni Semeraro

A vast majority of the world's 7,000 spoken languages are predicted to become extinct within this century, including the endangered language of Ladin from the Italian Alps. Linguists who work to preserve a language's phonetic and…

音频与语音处理 · 电气工程与系统科学 2021-08-31 Zane Durante , Leena Mathur , Eric Ye , Sichong Zhao , Tejas Ramdas , Khalil Iskarous

Large language models (LLMs) show increasing potential in education, yet benchmarks for non-English languages in specialized domains remain scarce. We introduce MedBench-IT, the first comprehensive benchmark for evaluating LLMs on Italian…

计算与语言 · 计算机科学 2025-09-10 Ruggero Marino Lazzaroni , Alessandro Angioi , Michelangelo Puliga , Davide Sanna , Roberto Marras

The advent of generative AI tools has had a profound impact on societies globally, transcending geographical boundaries. Understanding these tools' global reception and utilization is crucial for service providers and policymakers in…

计算机与社会 · 计算机科学 2025-03-24 Taichi Murayama , Kunihiro Miyazaki , Yasuko Matsubara , Yasushi Sakurai

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

The rapid development of Large Language Models (LLMs) demonstrates remarkable multilingual capabilities in natural language processing, attracting global attention in both academia and industry. To mitigate potential discrimination and…

Recent advances in data science, machine learning, and artificial intelligence, such as the emergence of large language models, are leading to an increasing demand for data that can be processed by such models. While data sources are…

机器学习 · 计算机科学 2023-09-13 Paul Bilokon , Oleksandr Bilokon , Saeed Amen

Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs…

计算机与社会 · 计算机科学 2024-07-12 Henry Tari , Danial Khan , Justus Rutten , Darian Othman , Rishabh Kaushal , Thales Bertaglia , Adriana Iamnitchi

Traditionally, large language models have been either trained on general web crawls or domain-specific data. However, recent successes of generative large language models, have shed light on the benefits of cross-domain datasets. To examine…

Word representation is fundamental in NLP tasks, because it is precisely from the coding of semantic closeness between words that it is possible to think of teaching a machine to understand text. Despite the spread of word embedding…

Translations often carry traces of the source language, a phenomenon known as translationese. We introduce the first freely available English-to-Swedish dataset contrasting translationese sentences with idiomatic alternatives, designed to…

计算与语言 · 计算机科学 2026-03-10 Jenny Kunz , Anja Jarochenko , Marcel Bollmann