English
Related papers

Related papers: Larth: Dataset and Machine Translation for Etrusca…

200 papers

Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored…

Computation and Language · Computer Science 2025-03-13 Marco Giordano , Claudia Rinaldi

With 17,000 pairs of Sicilian-English translated sentences, Arba Sicula developed the first neural machine translator for the Sicilian language. Using small subword vocabularies, we trained small Transformer models with high dropout…

Computation and Language · Computer Science 2021-10-06 Eryk Wdowiak

Palaeohispanic languages are those spoken in the Iberian Peninsula before the arrival of the Romans in the 3rd Century B.C. Their study was really put on motion after G\'omez Moreno deciphered the Iberian Levantine script, one of the…

Computation and Language · Computer Science 2026-04-16 Gonzalo Martínez-Fernández , Jose F Quesada , Agustín Riscos-Núñez , Francisco José Salguero-Lamillar

The widespread use of conversational and question answering systems made it necessary to improve the performances of speaker intent detection and understanding of related semantic slots, i.e., Spoken Language Understanding (SLU). Often,…

Computation and Language · Computer Science 2019-07-18 Valentina Bellomaria , Giuseppe Castellucci , Andrea Favalli , Raniero Romagnoli

Computational historical linguistics seeks to systematically understand processes of sound change, including during periods at which little to no formal recording of language is attested. At the same time, few computational resources exist…

Computation and Language · Computer Science 2024-04-26 Stephen Bothwell , Brian DuSell , David Chiang , Brian Krostenko

We present the first parallel dataset for English-Tulu translation. Tulu, classified within the South Dravidian linguistic family branch, is predominantly spoken by approximately 2.5 million individuals in southwestern India. Our dataset is…

Computation and Language · Computer Science 2024-03-29 Manu Narayanan , Noëmi Aepli

We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century. In a series of case studies, we illustrate the affordances…

Computation and Language · Computer Science 2020-09-22 David Bamman , Patrick J. Burns

This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…

Computation and Language · Computer Science 2018-07-30 Marcely Zanon Boito , Antonios Anastasopoulos , Marika Lekakou , Aline Villavicencio , Laurent Besacier

The growing interest in argument mining and computational argumentation brings with it a plethora of Natural Language Understanding (NLU) tasks and corresponding datasets. However, as with many other NLU tasks, the dominant language is…

Computation and Language · Computer Science 2020-10-14 Orith Toledo-Ronen , Matan Orbach , Yonatan Bilu , Artem Spector , Noam Slonim

We present the first neural machine translation system for translation between the endangered Erzya language and Russian and the dataset collected by us to train and evaluate it. The BLEU scores are 17 and 19 for translation to Erzya and…

Computation and Language · Computer Science 2022-09-21 David Dale

We present a survey covering the state of the art in low-resource machine translation research. There are currently around 7000 languages spoken in the world and almost all language pairs lack significant resources for training machine…

Computation and Language · Computer Science 2022-02-08 Barry Haddow , Rachel Bawden , Antonio Valerio Miceli Barone , Jindřich Helcl , Alexandra Birch

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale…

Computation and Language · Computer Science 2024-06-25 Alkis Koudounas , Moreno La Quatra , Lorenzo Vaiani , Luca Colomba , Giuseppe Attanasio , Eliana Pastor , Luca Cagliero , Elena Baralis

This paper describes "TLT-school" a corpus of speech utterances collected in schools of northern Italy for assessing the performance of students learning both English and German. The corpus was recorded in the years 2017 and 2018 from…

Computation and Language · Computer Science 2020-01-23 Roberto Gretter , Marco Matassoni , Stefano Bannò , Daniele Falavigna

Current research into spoken language translation (SLT),or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limited set of language…

Large-scale pretrained language models have become ubiquitous in Natural Language Processing. However, most of these models are available either in high-resource languages, in particular English, or as multilingual models that compromise…

Computation and Language · Computer Science 2020-09-21 Stefan Daniel Dumitrescu , Andrei-Marius Avram , Sampo Pyysalo

In this paper, we present DIETA, a small, decoder-only Transformer model with 0.5 billion parameters, specifically designed and trained for Italian-English machine translation. We collect and curate a large parallel corpus consisting of…

Computation and Language · Computer Science 2026-01-27 Pranav Kasela , Marco Braga , Alessandro Ghiotto , Andrea Pilzer , Marco Viviani , Alessandro Raganato

While there are more than 7000 languages in the world, most translation research efforts have targeted a few high-resource languages. Commercial translation systems support only one hundred languages or fewer, and do not make these models…

Computation and Language · Computer Science 2024-10-24 Thamme Gowda , Zhao Zhang , Chris A Mattmann , Jonathan May

This study focuses on the generation of Persian named entity datasets through the application of machine translation on English datasets. The generated datasets were evaluated by experimenting with one monolingual and one multilingual…

Computation and Language · Computer Science 2025-02-21 Amir Sartipi , Afsaneh Fatemi

A vast majority of the world's 7,000 spoken languages are predicted to become extinct within this century, including the endangered language of Ladin from the Italian Alps. Linguists who work to preserve a language's phonetic and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-31 Zane Durante , Leena Mathur , Eric Ye , Sichong Zhao , Tejas Ramdas , Khalil Iskarous

Given the prevalence of crowd sourced labor in creating Natural Language processing datasets, these aforementioned sets have become increasingly large. For instance, the SQUAD dataset currently sits at over 80,000 records. However, because…

Computation and Language · Computer Science 2023-04-28 Will Rieger
‹ Prev 1 2 3 10 Next ›