中文
相关论文

相关论文: FreCDo: A Large Corpus for French Cross-Domain Dia…

200 篇论文

In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their…

计算与语言 · 计算机科学 2026-04-08 Cherifa Ben Khelil , Jean-Yves Antoine , Anaïs Halftermeyer , Frédéric Rayar , Mathieu Thebaud

Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet their adaptation to specialized fields remains challenging, particularly for non-English languages. This study investigates domain-adaptive…

This paper presents a publicly available corpus of French encyclopedic history texts annotated according to the Berkeley FrameNet formalism. The main difference in our approach compared to previous works on semantic parsing with FrameNet is…

计算与语言 · 计算机科学 2018-12-20 Gabriel Marzinotto , Jeremy Auguste , Frederic Bechet , Géraldine Damnati , Alexis Nasr

Fake news has altered society in negative ways in politics and culture. It has adversely affected both online social network systems as well as offline communities and conversations. Using automatic machine learning classification models is…

计算与语言 · 计算机科学 2020-03-13 Kai Nakamura , Sharon Levy , William Yang Wang

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article…

计算与语言 · 计算机科学 2025-05-01 Michelle Wastl , Jannis Vamvas , Selena Calleri , Rico Sennrich

Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation…

计算与语言 · 计算机科学 2023-10-17 Nuno M. Guerreiro , Ricardo Rei , Daan van Stigt , Luisa Coheur , Pierre Colombo , André F. T. Martins

Objective. Epidemiological studies require data that are in alignment with the classifications established for occupations or economic activities. The classifications usually include hundreds of codes and titles. Manual coding of raw data…

计算与语言 · 计算机科学 2020-12-15 Nenad Savic , Nicolas Bovio , Fabian Gilbert , Irina Guseva Canu

Teachers and students are increasingly relying on online learning resources to supplement the ones provided in school. This increase in the breadth and depth of available resources is a great thing for students, but only provided they are…

计算与语言 · 计算机科学 2023-04-17 Antoine Lefebvre-Brossard , Stephane Gazaille , Michel C. Desmarais

It is challenging to control the quality of online information due to the lack of supervision over all the information posted online. Manual checking is almost impossible given the vast number of posts made on online media and how quickly…

计算与语言 · 计算机科学 2022-03-16 Rini Anggrainingsih , Ghulam Mubashar Hassan , Amitava Datta

This article presents multilingual deep learning models for identifying web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introduce the Multilingual CORE corpora, which contain over 72,000…

In this work, we introduce the methods proposed by the UnibucKernel team in solving the Social Media Variety Geolocation task featured in the 2020 VarDial Evaluation Campaign. We address only the second subtask, which targets a data set…

计算与语言 · 计算机科学 2020-10-09 Mihaela Gaman , Radu Tudor Ionescu

Disfluency correction (DC) is the process of removing disfluent elements like fillers, repetitions and corrections from spoken utterances to create readable and interpretable text. DC is a vital post-processing step applied to Automatic…

计算与语言 · 计算机科学 2023-10-26 Vineet Bhat , Preethi Jyothi , Pushpak Bhattacharyya

The exponential growth of online textual content across diverse domains has necessitated advanced methods for automated text classification. Large Language Models (LLMs) based on transformer architectures have shown significant success in…

计算与语言 · 计算机科学 2025-09-09 Zhyar Rzgar K Rostam , Gábor Kertész

Old French is a typical example of an under-resourced historic languages, that furtherly displays animportant amount of linguistic variation. In this paper, we present the current results of a long going project (2015-...) and describe how…

计算与语言 · 计算机科学 2021-09-24 Jean-Baptiste Camps , Thibault Clérice , Frédéric Duval , Lucence Ing , Naomi Kanaoka , Ariane Pinche

We present a new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5,700+ sentences) between native English and French speakers,…

计算与语言 · 计算机科学 2019-06-03 Rachel Bawden , Sophie Rosset , Thomas Lavergne , Eric Bilinski

We leverage generative large language models for language learning applications, focusing on estimating the difficulty of foreign language texts and simplifying them to lower difficulty levels. We frame both tasks as prediction problems and…

计算与语言 · 计算机科学 2024-07-26 Henri Jamet , Yash Raj Shrestha , Michalis Vlachos

The traditional approach to morphological inflection (the task of modifying a base word (lemma) to express grammatical categories) has been, for decades, to consider lexical entries of lemma-tag-form triples uniformly, lacking any…

计算与语言 · 计算机科学 2025-10-28 Tomáš Sourada , Jana Straková

We present SDS-200, a corpus of Swiss German dialectal speech with Standard German text translations, annotated with dialect, age, and gender information of the speakers. The dataset allows for training speech translation, dialect…

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

计算与语言 · 计算机科学 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

计算与语言 · 计算机科学 2019-12-02 Masato Hagiwara , Masato Mita