English
Related papers

Related papers: FreCDo: A Large Corpus for French Cross-Domain Dia…

200 papers

In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their…

Computation and Language · Computer Science 2026-04-08 Cherifa Ben Khelil , Jean-Yves Antoine , Anaïs Halftermeyer , Frédéric Rayar , Mathieu Thebaud

Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet their adaptation to specialized fields remains challenging, particularly for non-English languages. This study investigates domain-adaptive…

Computation and Language · Computer Science 2026-04-09 Aidan Mannion , Cécile Macaire , Armand Violle , Stéphane Ohayon , Xavier Tannier , Didier Schwab , Lorraine Goeuriot , François Portet

This paper presents a publicly available corpus of French encyclopedic history texts annotated according to the Berkeley FrameNet formalism. The main difference in our approach compared to previous works on semantic parsing with FrameNet is…

Computation and Language · Computer Science 2018-12-20 Gabriel Marzinotto , Jeremy Auguste , Frederic Bechet , Géraldine Damnati , Alexis Nasr

Fake news has altered society in negative ways in politics and culture. It has adversely affected both online social network systems as well as offline communities and conversations. Using automatic machine learning classification models is…

Computation and Language · Computer Science 2020-03-13 Kai Nakamura , Sharon Levy , William Yang Wang

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article…

Computation and Language · Computer Science 2025-05-01 Michelle Wastl , Jannis Vamvas , Selena Calleri , Rico Sennrich

Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation…

Computation and Language · Computer Science 2023-10-17 Nuno M. Guerreiro , Ricardo Rei , Daan van Stigt , Luisa Coheur , Pierre Colombo , André F. T. Martins

Objective. Epidemiological studies require data that are in alignment with the classifications established for occupations or economic activities. The classifications usually include hundreds of codes and titles. Manual coding of raw data…

Computation and Language · Computer Science 2020-12-15 Nenad Savic , Nicolas Bovio , Fabian Gilbert , Irina Guseva Canu

Teachers and students are increasingly relying on online learning resources to supplement the ones provided in school. This increase in the breadth and depth of available resources is a great thing for students, but only provided they are…

Computation and Language · Computer Science 2023-04-17 Antoine Lefebvre-Brossard , Stephane Gazaille , Michel C. Desmarais

It is challenging to control the quality of online information due to the lack of supervision over all the information posted online. Manual checking is almost impossible given the vast number of posts made on online media and how quickly…

Computation and Language · Computer Science 2022-03-16 Rini Anggrainingsih , Ghulam Mubashar Hassan , Amitava Datta

This article presents multilingual deep learning models for identifying web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introduce the Multilingual CORE corpora, which contain over 72,000…

Computation and Language · Computer Science 2026-02-10 Erik Henriksson , Amanda Myntti , Saara Hellström , Anni Eskelinen , Selcen Erten-Johansson , Veronika Laippala

In this work, we introduce the methods proposed by the UnibucKernel team in solving the Social Media Variety Geolocation task featured in the 2020 VarDial Evaluation Campaign. We address only the second subtask, which targets a data set…

Computation and Language · Computer Science 2020-10-09 Mihaela Gaman , Radu Tudor Ionescu

Disfluency correction (DC) is the process of removing disfluent elements like fillers, repetitions and corrections from spoken utterances to create readable and interpretable text. DC is a vital post-processing step applied to Automatic…

Computation and Language · Computer Science 2023-10-26 Vineet Bhat , Preethi Jyothi , Pushpak Bhattacharyya

The exponential growth of online textual content across diverse domains has necessitated advanced methods for automated text classification. Large Language Models (LLMs) based on transformer architectures have shown significant success in…

Computation and Language · Computer Science 2025-09-09 Zhyar Rzgar K Rostam , Gábor Kertész

Old French is a typical example of an under-resourced historic languages, that furtherly displays animportant amount of linguistic variation. In this paper, we present the current results of a long going project (2015-...) and describe how…

Computation and Language · Computer Science 2021-09-24 Jean-Baptiste Camps , Thibault Clérice , Frédéric Duval , Lucence Ing , Naomi Kanaoka , Ariane Pinche

We present a new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5,700+ sentences) between native English and French speakers,…

Computation and Language · Computer Science 2019-06-03 Rachel Bawden , Sophie Rosset , Thomas Lavergne , Eric Bilinski

We leverage generative large language models for language learning applications, focusing on estimating the difficulty of foreign language texts and simplifying them to lower difficulty levels. We frame both tasks as prediction problems and…

Computation and Language · Computer Science 2024-07-26 Henri Jamet , Yash Raj Shrestha , Michalis Vlachos

The traditional approach to morphological inflection (the task of modifying a base word (lemma) to express grammatical categories) has been, for decades, to consider lexical entries of lemma-tag-form triples uniformly, lacking any…

Computation and Language · Computer Science 2025-10-28 Tomáš Sourada , Jana Straková

We present SDS-200, a corpus of Swiss German dialectal speech with Standard German text translations, annotated with dialect, age, and gender information of the speakers. The dataset allows for training speech translation, dialect…

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

Computation and Language · Computer Science 2019-12-02 Masato Hagiwara , Masato Mita