English
Related papers

Related papers: Infini-News: Efficiently Queryable Access to 1.3 B…

200 papers

This paper presents a novel methodological framework for detecting and classifying latent constructs, including frames, narratives, and topics, from textual data using Open-Source Large Language Models (LLMs). The proposed hybrid approach…

Computation and Language · Computer Science 2025-04-01 Maël Kubli

We present NewsQs (news-cues), a dataset that provides question-answer pairs for multiple news documents. To create NewsQs, we augment a traditional multi-document summarization dataset with questions automatically generated by a T5-Large…

Computation and Language · Computer Science 2024-06-18 Alyssa Hwang , Kalpit Dixit , Miguel Ballesteros , Yassine Benajiba , Vittorio Castelli , Markus Dreyer , Mohit Bansal , Kathleen McKeown

Automatic text summarization aims to produce a brief but crucial summary for the input documents. Both extractive and abstractive methods have witnessed great success in English datasets in recent years. However, there has been a minimal…

Computation and Language · Computer Science 2021-10-22 Danqing Wang , Jiaze Chen , Xianze Wu , Hao Zhou , Lei Li

Cross-lingual information retrieval (CLIR) helps users find documents in languages different from their queries. This is especially important in academic search, where key research is often published in non-English languages. We present…

Information Retrieval · Computer Science 2025-11-20 Francisco Valentini , Diego Kozlowski , Vincent Larivière

Fake news, rumor, incorrect information, and misinformation detection are nowadays crucial issues as these might have serious consequences for our social fabrics. The rate of such information is increasing rapidly due to the availability of…

Computation and Language · Computer Science 2018-11-13 Arjun Roy , Kingshuk Basak , Asif Ekbal , Pushpak Bhattacharyya

We present a novel approach to data preparation for developing multilingual Indic large language model. Our meticulous data acquisition spans open-source and proprietary sources, including Common Crawl, Indic books, news articles, and…

Text summarization models are approaching human levels of fidelity. Existing benchmarking corpora provide concordant pairs of full and abridged versions of Web, news or, professional content. To date, all summarization datasets operate…

Computation and Language · Computer Science 2022-06-01 Seyed Ali Bahrainian , Sheridan Feucht , Carsten Eickhoff

Causal structure discovery methods are commonly applied to structured data where the causal variables are known and where statistical testing can be used to assess the causal relationships. By contrast, recovering a causal structure from…

Computation and Language · Computer Science 2024-10-10 Gaël Gendron , Jože M. Rožanec , Michael Witbrock , Gillian Dobbie

The accelerating pace of research on autoregressive generative models has produced thousands of papers, making manual literature surveys and reproduction studies increasingly impractical. We present a fully open-source, reproducible…

Information Retrieval · Computer Science 2025-08-07 Faruk Alpay , Bugra Kilictas , Hamdi Alakkad

With the ongoing growth in number of digital articles in a wider set of languages and the expanding use of different languages, we need annotation methods that enable browsing multi-lingual corpora. Multilingual probabilistic topic models…

Computation and Language · Computer Science 2021-01-11 Carlos Badenes-Olmedo , Jose-Luis Redondo García , Oscar Corcho

News sources undergo the process of selecting newsworthy information when covering a certain topic. The process inevitably exhibits selection biases, i.e. news sources' typical patterns of choosing what information to include in news…

Computation and Language · Computer Science 2023-04-10 Sihao Chen , William Bruno , Dan Roth

We describe a strategy for identifying the universe of research publications relevant to the application and development of artificial intelligence. The approach leverages the arXiv corpus of scientific preprints, in which authors choose…

Digital Libraries · Computer Science 2020-05-29 James Dunham , Jennifer Melot , Dewey Murdick

Linguistic diversity across the world creates a disparity with the availability of good quality digital language resources thereby restricting the technological benefits to majority of human population. The lack or absence of data resources…

Computation and Language · Computer Science 2025-10-16 Prawaal Sharma , Navneet Goyal , Poonam Goyal , Vishnupriyan R

News archives are an invaluable primary source for placing current events in historical context. But current search engine tools do a poor job at uncovering broad themes and narratives across documents. We present Rookie: a practical…

Human-Computer Interaction · Computer Science 2017-08-08 Abram Handler , Brendan O'Connor

Fake news is dramatically increased in social media in recent years. This has prompted the need for effective fake news detection algorithms. Capsule neural networks have been successful in computer vision and are receiving attention for…

Computation and Language · Computer Science 2020-02-05 Mohammad Hadi Goldani , Saeedeh Momtazi , Reza Safabakhsh

In order to explore the suitability of a fine-grained classification of journal articles by exploiting multiple sources of information, articles are organized in a two-layer multiplex. The first layer conveys similarities based on the…

Digital Libraries · Computer Science 2024-01-02 Alberto Baccini , Federica Baccini , Lucio Barabesi , Martina Cioni , Eugenio Petrovich , Daria Pignalosa

Large-scale data sets on scholarly publications are the basis for a variety of bibliometric analyses and natural language processing (NLP) applications. Especially data sets derived from publication's full-text have recently gained…

Digital Libraries · Computer Science 2023-11-06 Tarek Saier , Johan Krause , Michael Färber

Obtaining sufficient information in one's mother tongue is crucial for satisfying the information needs of the users. While high-resource languages have abundant online resources, the situation is less than ideal for very low-resource…

Computation and Language · Computer Science 2024-05-03 Abhinaba Bala , Ashok Urlana , Rahul Mishra , Parameswari Krishnamurthy

Classifying the same event reported by different countries is of significant importance for public opinion control and intelligence gathering. Due to the diverse types of news, relying solely on transla-tors would be costly and inefficient,…

Computation and Language · Computer Science 2023-05-31 Lin Wu , Rui Li , Wong-Hing Lam

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. Relating visual and textual items on document pages has…