English
Related papers

Related papers: Reliable Access to Massive Restricted Texts: Exper…

200 papers

This data paper introduces MajinBook, an open catalogue designed to facilitate the use of shadow libraries-such as Library Genesis and Z-Library-for computational social science and cultural analytics. By linking metadata from these vast,…

Computation and Language · Computer Science 2026-05-13 Antoine Mazières , Thierry Poibeau

We propose a bootstrap-based robust high-confidence level upper bound (Robust H-CLUB) for assessing the risks of large portfolios. The proposed approach exploits rank-based and quantile-based estimators, and can be viewed as a robust…

Statistics Theory · Mathematics 2015-01-13 Jianqing Fan , Fang Han , Han Liu , Byron Vickers

Training modern neural networks or models typically requires averaging over a sample of high-dimensional vectors. Poisoning attacks can skew or bias the average vectors used to train the model, forcing the model to learn specific patterns…

Cryptography and Security · Computer Science 2024-12-17 Sarthak Choudhary , Aashish Kolluri , Prateek Saxena

This paper is devoted to the adaptation of generative large language models for the Tajik language, a low-resource language with Cyrillic script. To overcome the shortage of digital text resources, the author created and publicly released…

Computation and Language · Computer Science 2026-05-06 Mullosharaf K. Arabov

Computational narrative analysis aims to capture rhythm, tension, and emotional dynamics in literary texts. Existing large language models can generate long stories but overly focus on causal coherence, neglecting the complex story arcs and…

Computation and Language · Computer Science 2026-05-06 Mingzhe Lu , Yiwen Wang , Yanbing Liu , Qi You , Chong Liu , Ruize Qin , Haoyu Dong , Wenyu Zhang , Jiarui Zhang , Yue Hu , Yunpeng Li

Main memory column-stores have proven to be efficient for processing analytical queries. Still, there has been much less work in the context of clusters. Using only a single machine poses several restrictions: Processing power and data…

Databases · Computer Science 2017-09-18 Demian Hespe , Martin Weidner , Jonathan Dees , Peter Sanders

We study the ability of state-of-the art models to answer constraint satisfaction queries for information retrieval (e.g., 'a list of ice cream shops in San Diego'). In the past, such queries were considered to be tasks that could only be…

We consider the problem of producing compact architectures for text classification, such that the full model fits in a limited amount of memory. After considering different solutions inspired by the hashing literature, we propose a method…

Computation and Language · Computer Science 2016-12-19 Armand Joulin , Edouard Grave , Piotr Bojanowski , Matthijs Douze , Hérve Jégou , Tomas Mikolov

The exponential growth of scientific production makes secondary literature abridgements increasingly demanding. We introduce a new open-source framework for systematic reviews that significantly reduces time and workload for collecting and…

Digital Libraries · Computer Science 2022-02-24 Angelo D'Ambrosio , Hajo Grundmann , Tjibbe Donker

Existing methods for evaluating large language models face challenges such as data contamination, sensitivity to prompts, and the high cost of benchmark creation. To address this, we propose a lossless data compression based evaluation…

Computation and Language · Computer Science 2024-02-06 Yucheng Li , Yunhao Guo , Frank Guerin , Chenghua Lin

With an ever growing number of heterogeneous applicational services running on equally heterogeneous computational systems, the problem of resource management becomes more essential. Although current solutions consider some network and time…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-08-05 Rui Eduardo Lopes , Duarte Raposo , Pedro V. Teixeira , Susana Sargento

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a web-scale crawl of…

Computation and Language · Computer Science 2018-03-01 Alexander Panchenko , Eugen Ruppert , Stefano Faralli , Simone Paolo Ponzetto , Chris Biemann

Citation analysis is one of the most frequently used methods in research evaluation. We are seeing significant growth in citation analysis through bibliometric metadata, primarily due to the availability of citation databases such as the…

Digital Libraries · Computer Science 2020-09-01 Sehrish Iqbal , Saeed-Ul Hassan , Naif Radi Aljohani , Salem Alelyani , Raheel Nawaz , Lutz Bornmann

Cassandra is a popular structured storage system with high-performance, scalability and high availability, and is usually used to store data that has some sortable attributes. When deploying and configuring Cassandra, it is important to…

Databases · Computer Science 2018-10-03 Jialin Qiao , Xiangdong Huang , Lei Rui , Jianmin Wang

Cross-referencing, which links passages of text to other related passages, can be a valuable study aid for facilitating comprehension of a text. However, cross-referencing requires first, a comprehensive thematic knowledge of the entire…

Computation and Language · Computer Science 2019-05-21 Jeffrey Lund , Piper Armstrong , Wilson Fearn , Stephen Cowley , Emily Hales , Kevin Seppi

Scholarship on underresourced languages bring with them a variety of challenges which make access to the full spectrum of source materials and their evaluation difficult. For Coptic in particular, large scale analyses and any kind of…

Computation and Language · Computer Science 2023-06-22 Caroline T. Schroeder , Amir Zeldes

Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of the datasets used to train these models, or to support their…

The role of trust within Human-Computer Interaction is being redefined. With the increasing omnipresence, autonomy, and opacity of technology, users often struggle to understand the capabilities and limitations of systems. In this article,…

Human-Computer Interaction · Computer Science 2026-04-08 Gabriela Beltrão , Debora F. de Souza , Sonia Sousa , David Lamas

This paper presents results of our experiments for the next utterance ranking on the Ubuntu Dialog Corpus -- the largest publicly available multi-turn dialog corpus. First, we use an in-house implementation of previously reported models to…

Computation and Language · Computer Science 2015-11-04 Rudolf Kadlec , Martin Schmid , Jan Kleindienst

The Web Based File Clustering and Indexing for Mindoro State University aim to organize data circulated over the Web into groups or collections to facilitate data availability and access and at the same time meet user preferences. The main…

Information Retrieval · Computer Science 2022-02-15 Christie A. Luzon , Luisito Lolong Lacatan , Harold Y. Bangalisan , Jayvee M. Osapdin