English
Related papers

Related papers: WikiMulti: a Corpus for Cross-Lingual Summarizatio…

200 papers

Acknowledged as one of the most successful online cooperative projects in human society, Wikipedia has obtained rapid growth in recent years and desires continuously to expand content and disseminate knowledge values for everyone globally.…

Computation and Language · Computer Science 2022-10-25 Hoang Thang Ta , Alexander Gelbukha , Grigori Sidorov

In recent years, the field of document understanding has progressed a lot. A significant part of this progress has been possible thanks to the use of language models pretrained on large amounts of documents. However, pretraining corpora…

Computation and Language · Computer Science 2023-06-07 Michał Turski , Tomasz Stanisławek , Karol Kaczmarek , Paweł Dyda , Filip Graliński

Relation classification is one of the key topics in information extraction, which can be used to construct knowledge bases or to provide useful information for question answering. Current approaches for relation classification are mainly…

Computation and Language · Computer Science 2020-10-20 Abdullatif Köksal , Arzucan Özgür

This overview describes the official results of the CL-SciSumm Shared Task 2018 -- the first medium-scale shared task on scientific document summarization in the computational linguistics (CL) domain. This year, the dataset comprised 60…

Computation and Language · Computer Science 2019-09-05 Kokil Jaidka , Michihiro Yasunaga , Muthu Kumar Chandrasekaran , Dragomir Radev , Min-Yen Kan

High-quality scientific extreme summary (TLDR) facilitates effective science communication. How do large language models (LLMs) perform in generating them? How are LLM-generated summaries different from those written by human experts?…

Computation and Language · Computer Science 2025-12-30 Zhuoqi Lyu , Qing Ke

Document summarization is a task to shorten texts into concise and informative summaries. This paper introduces a novel dataset designed for summarizing multiple scientific articles into a section of a survey. Our contributions are: (1)…

Query-specific article generation is the task of, given a search query, generate a single article that gives an overview of the topic. We envision such articles as an alternative to presenting a ranking of search results. While generative…

Information Retrieval · Computer Science 2023-10-20 Connor Lennox , Sumanta Kashyapi , Laura Dietz

Identifying cross-language plagiarism is challenging, especially for distant language pairs and sense-for-sense translations. We introduce the new multilingual retrieval model Cross-Language Ontology-Based Similarity Analysis (CL-OSA) for…

Computation and Language · Computer Science 2021-12-17 Johannes Stegmüller , Fabian Bauer-Marquart , Norman Meuschke , Terry Ruas , Moritz Schubotz , Bela Gipp

We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web crawls from the…

Retrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification. Previous works have mostly tackled the task monolingually,…

Computation and Language · Computer Science 2025-09-23 Alan Ramponi , Marco Rovera , Robert Moro , Sara Tonelli

While impressive performance has been achieved on the task of Answer Sentence Selection (AS2) for English, the same does not hold for languages that lack large labeled datasets. In this work, we propose Cross-Lingual Knowledge Distillation…

Computation and Language · Computer Science 2025-01-06 Shivanshu Gupta , Yoshitomo Matsubara , Ankit Chadha , Alessandro Moschitti

Previous research in multi-document news summarization has typically concentrated on collating information that all sources agree upon. However, the summarization of diverse information dispersed across multiple articles about an event…

Computation and Language · Computer Science 2024-03-26 Kung-Hsiang Huang , Philippe Laban , Alexander R. Fabbri , Prafulla Kumar Choubey , Shafiq Joty , Caiming Xiong , Chien-Sheng Wu

Nowadays, digital news articles are widely available, published by various editors and often written in different languages. This large volume of diverse and unorganized information makes human reading very difficult or almost impossible.…

Computation and Language · Computer Science 2020-04-20 Mathis Linger , Mhamed Hajaiej

Summarization is the task of compressing source document(s) into coherent and succinct passages. This is a valuable tool to present users with concise and accurate sketch of the top ranked documents related to their queries. Query-based…

Computation and Language · Computer Science 2020-10-27 Sayali Kulkarni , Sheide Chammas , Wan Zhu , Fei Sha , Eugene Ie

Cross-lingual text summarization aims at generating a document summary in one language given input in another language. It is a practically important but under-explored task, primarily due to the dearth of available data. Existing methods…

Computation and Language · Computer Science 2020-06-30 Zi-Yi Dou , Sachin Kumar , Yulia Tsvetkov

Today's research progress in the field of multi-document summarization is obstructed by the small number of available datasets. Since the acquisition of reference summaries is costly, existing datasets contain only hundreds of samples at…

Computation and Language · Computer Science 2020-02-18 Diego Antognini , Boi Faltings

The abstractive methods lack of creative ability is particularly a problem in automatic text summarization. The summaries generated by models are mostly extracted from the source articles. One of the main causes for this problem is the lack…

Computation and Language · Computer Science 2022-06-10 Xiaojun Liu , Shunan Zang , Chuang Zhang , Xiaojun Chen , Yangyang Ding

The milestone improvements brought about by deep representation learning and pre-training techniques have led to large performance gains across downstream NLP, IR and Vision tasks. Multimodal modeling techniques aim to leverage large…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Krishna Srinivasan , Karthik Raman , Jiecao Chen , Michael Bendersky , Marc Najork

Defining psycholinguistic characteristics in written texts is a task gaining increasing attention from researchers. One of the most widely used tools in the current field is Linguistic Inquiry and Word Count (LIWC) that originally was…

Computation and Language · Computer Science 2026-01-29 Elina Sigdel , Anastasia Panfilova

An edit summary is a succinct comment written by a Wikipedia editor explaining the nature of, and reasons for, an edit to a Wikipedia page. Edit summaries are crucial for maintaining the encyclopedia: they are the first thing seen by…

Computation and Language · Computer Science 2024-08-20 Marija Šakota , Isaac Johnson , Guosheng Feng , Robert West