English
Related papers

Related papers: The WiLI benchmark dataset for written language id…

200 papers

Text simplification is a valuable technique. However, current research is limited to sentence simplification. In this paper, we define and investigate a new task of document-level text simplification, which aims to simplify a document…

Computation and Language · Computer Science 2021-10-12 Renliang Sun , Hanqi Jin , Xiaojun Wan

Identifying words which may cause difficulty for a reader is an essential step in most lexical text simplification systems prior to lexical substitution and can also be used for assessing the readability of a text. This task is commonly…

Computation and Language · Computer Science 2022-11-04 Matthew Shardlow , Richard Evans , Marcos Zampieri

With the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word…

Computer Vision and Pattern Recognition · Computer Science 2015-05-13 Baoguang Shi , Cong Yao , Chengquan Zhang , Xiaowei Guo , Feiyue Huang , Xiang Bai

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75…

Computation and Language · Computer Science 2025-03-03 Michael Dinzinger , Laura Caspari , Kanishka Ghosh Dastidar , Jelena Mitrović , Michael Granitzer

This paper introduces IGGA, a dataset of 160 industry guidelines and policy statements for the use of Generative AIs (GAIs) and Large Language Models (LLMs) in industry and workplace settings, collected from official company websites, and…

Computers and Society · Computer Science 2025-03-19 Junfeng Jiao , Saleh Afroogh , Kevin Chen , David Atkinson , Amit Dhurandhar

This paper presents an approach to classify documents in any language into an English topical label space, without any text categorization training data. The approach, Cross-Lingual Dataless Document Classification (CLDDC) relies on mapping…

Computation and Language · Computer Science 2016-11-15 Yangqiu Song , Stephen Mayhew , Dan Roth

We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a…

Computation and Language · Computer Science 2026-05-01 Jaap Jumelet , Leonie Weissweiler , Joakim Nivre , Arianna Bisazza

English is the predominant language on the web, powering nearly half of the world's top ten million websites. Support for multilingual content is nevertheless growing, with many websites increasingly combining English with regional or…

Computation and Language · Computer Science 2025-08-27 Masudul Hasan Masud Bhuiyan , Matteo Varvello , Yasir Zaki , Cristian-Alexandru Staicu

Recent years have seen a growing number of publications that analyse Natural Language Inference (NLI) datasets for superficial cues, whether they undermine the complexity of the tasks underlying those datasets and how they impact those…

Computation and Language · Computer Science 2020-06-01 Viktor Schlegel , Goran Nenadic , Riza Batista-Navarro

Motivated by the sparsity of NLP resources for Eastern European languages, we present a broad index of existing Eastern European language resources (90+ datasets and 45+ models) published as a github repository open for updates from the…

Computation and Language · Computer Science 2022-05-12 Alexey Tikhonov , Alex Malkhasov , Andrey Manoshin , George Dima , Réka Cserháti , Md. Sadek Hossain Asif , Matt Sárdi

This article introduces a corpus of cuneiform texts from which the dataset for the use of the Cuneiform Language Identification (CLI) 2019 shared task was derived as well as some preliminary language identification experiments conducted…

Computation and Language · Computer Science 2019-03-14 Tommi Jauhiainen , Heidi Jauhiainen , Tero Alstola , Krister Lindén

Relation classification is one of the key topics in information extraction, which can be used to construct knowledge bases or to provide useful information for question answering. Current approaches for relation classification are mainly…

Computation and Language · Computer Science 2020-10-20 Abdullatif Köksal , Arzucan Özgür

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This…

Computation and Language · Computer Science 2024-10-17 Tarek Naous , Michael J. Ryan , Anton Lavrouk , Mohit Chandra , Wei Xu

Over the past years, deep learning methods allowed for new state-of-the-art results in ad-hoc information retrieval. However such methods usually require large amounts of annotated data to be effective. Since most standard ad-hoc…

Information Retrieval · Computer Science 2020-03-18 Jibril Frej , Didier Schwab , Jean-Pierre Chevallet

As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, including on tasks like machine translation…

Wikipedia can be edited by anyone and thus contains various quality sentences. Therefore, Wikipedia includes some poor-quality edits, which are often marked up by other editors. While editors' reviews enhance the credibility of Wikipedia,…

Computation and Language · Computer Science 2024-01-02 Kenichiro Ando , Satoshi Sekine , Mamoru Komachi

Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems extend this coverage under certain circumstances, but for most…

Computation and Language · Computer Science 2026-02-10 Rasul Dent , Pedro Ortiz Suarez , Thibault Clérice , Benoît Sagot

Increasingly, web content is automatically generated by large language models (LLMs) with little human input. We call this "LLM-dominant" content. Since LLMs plagiarize and hallucinate, LLM-dominant content can be unreliable and unethical.…

Networking and Internet Architecture · Computer Science 2025-10-13 Sichang Steven He , Ramesh Govindan , Harsha V. Madhyastha

This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand…

Human label variation (Plank 2022), or annotation disagreement, exists in many natural language processing (NLP) tasks. To be robust and trusted, NLP models need to identify such variation and be able to explain it. To this end, we created…

Computation and Language · Computer Science 2023-04-26 Nan-Jiang Jiang , Chenhao Tan , Marie-Catherine de Marneffe