English
Related papers

Related papers: A Collaborative Ecosystem for Digital Coptic Studi…

200 papers

The metaphor studies community has developed numerous valuable labelled corpora in various languages over the years. Many of these resources are not only unknown to the NLP community, but are also often not easily shared among the…

Computation and Language · Computer Science 2025-03-11 Joanne Boisson , Arif Mehmood , Jose Camacho-Collados

We investigate an abstract conceptualisation of DigitalEcosystems from a computer science perspective. We then provide a conceptual framework for the cross pollination of ideas, concepts and understanding between different classes of…

Multiagent Systems · Computer Science 2011-05-10 Gerard Briscoe

In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of document creation time.…

Computation and Language · Computer Science 2026-02-13 Mariia Fedorova , Andrey Kutuzov , Khonzoda Umarova

The task of discovering topics in text corpora has been dominated by Latent Dirichlet Allocation and other Topic Models for over a decade. In order to apply these approaches to massive text corpora, the vocabulary needs to be reduced…

Computation and Language · Computer Science 2019-08-08 Gibran Fuentes-Pineda , Ivan Vladimir Meza-Ruiz

Literature search is arguably one of the most important phases of the academic and non-academic research. The increase in the number of published papers each year makes manual search inefficient and furthermore insufficient. Hence,…

Information Retrieval · Computer Science 2012-09-27 Onur Küçüktunç , Erik Saule , Kamer Kaya , Ümit V. Çatalyürek

The COVID-19 pandemic triggered a wave of novel scientific literature that is impossible to inspect and study in a reasonable time frame manually. Current machine learning methods offer to project such body of literature into the vector…

Information Retrieval · Computer Science 2021-10-19 Blaž Škrlj , Marko Jukič , Nika Eržen , Senja Pollak , Nada Lavrač

Most previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a…

Information Retrieval · Computer Science 2007-05-23 Oren Kurland , Lillian Lee

Over the past decades, the competition for academic resources has gradually intensified, and worsened with the current financial crisis. To optimize the resource allocation, individualized assessment of research results is being actively…

Applications · Statistics 2010-03-18 Ge Wang , Jiansheng Yang

Generative large-scale language models create the fifth paradigm of scientific research, organically combine data science and computational intelligence, transform the research paradigm of natural language processing and multimodal…

Digital Libraries · Computer Science 2024-04-30 Jiangfeng Liu , Ziyi Wang , Jing Xie , Lei Pei

Collaborative work on unstructured or semi-structured documents, such as in literature corpora or source code, often involves agreed upon templates containing metadata. These templates are not consistent across users and over time.…

Digital Libraries · Computer Science 2016-08-24 Owen Kaser , Daniel Lemire

Keeping track of scientific challenges, advances and emerging directions is a fundamental part of research. However, researchers face a flood of papers that hinders discovery of important knowledge. In biomedicine, this directly impacts…

The rapid growth of AI conferences is straining an already fragile peer-review system, leading to heavy reviewer workloads, expertise mismatches, inconsistent evaluation standards, superficial or templated reviews, and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jing Yang , Qiyao Wei , Jiaxin Pei

Witnesses of medieval literary texts, preserved in manuscript, are layered objects , being almost exclusively copies of copies. This results in multiple and hard to distinguish linguistic strata -- the author's scripta interacting with the…

Computation and Language · Computer Science 2018-02-06 Jean-Baptiste Camps

Subtask distillation is an emerging paradigm in which compact, specialized models are extracted from large, general-purpose 'foundation models' for deployment in environments with limited resources or in standalone computer systems.…

Machine Learning · Computer Science 2026-02-03 Pattarawat Chormai , Klaus-Robert Müller , Grégoire Montavon

This paper presents the first systematic study of strategies for translating Coptic into French. Our comprehensive pipeline systematically evaluates: pivot versus direct translation, the impact of pre-training, the benefits of multi-version…

Computation and Language · Computer Science 2026-05-14 Nasma Chaoui , Richard Khoury

We present a system that allows life-science researchers to search a linguistically annotated corpus of scientific texts using patterns over dependency graphs, as well as using patterns over token sequences and a powerful variant of boolean…

Computation and Language · Computer Science 2020-06-09 Hillel Taub-Tabib , Micah Shlain , Shoval Sadde , Dan Lahav , Matan Eyal , Yaara Cohen , Yoav Goldberg

New ways of documenting and describing language via electronic media coupled with new ways of distributing the results via the World-Wide Web offer a degree of access to language resources that is unparalleled in history. At the same time,…

Computation and Language · Computer Science 2007-05-23 Gary Simons , Steven Bird

Scientific workflow systems are increasingly popular for expressing and executing complex data analysis pipelines over large datasets, as they offer reproducibility, dependability, and scalability of analyses by automatic parallelization on…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-09 Mario Sänger , Ninon De Mecquenem , Katarzyna Ewa Lewińska , Vasilis Bountris , Fabian Lehmann , Ulf Leser , Thomas Kosch

This paper presents the "Leipzig Corpus Miner", a technical infrastructure for supporting qualitative and quantitative content analysis. The infrastructure aims at the integration of 'close reading' procedures on individual documents with…

Computation and Language · Computer Science 2017-07-12 Andreas Niekler , Gregor Wiedemann , Gerhard Heyer

Currently, the world is in the midst of a severe global pandemic, which has affected all aspects of people's lives. As a result, there is a deluge of COVID-related digital media articles published in the United States, due to the disparate…

Computation and Language · Computer Science 2021-06-18 Xiangpeng Wan , Michael C. Lucic , Hakim Ghazzai , Yehia Massoud