English
Related papers

Related papers: Mapping Mutable Genres in Structurally Complex Vol…

200 papers

While coreference resolution is attracting more interest than ever from computational literature researchers, representative datasets of fully annotated long documents remain surprisingly scarce. In this paper, we introduce a new annotated…

Computation and Language · Computer Science 2025-10-20 Antoine Bourgois , Thierry Poibeau

We release 70,509 high-quality social networks extracted from multilingual fiction and nonfiction narratives. We additionally provide metadata for $\sim$30,000 of these texts (73\% nonfiction and 27\% fiction) written between 1800 and 1999…

Computation and Language · Computer Science 2025-04-01 Sil Hamilton , Rebecca M. M. Hicke , David Mimno , Matthew Wilkens

Academic literature reviews have traditionally relied on techniques such as keyword searches and accumulation of relevant back-references, using databases like Google Scholar or IEEEXplore. However, both the precision and accuracy of these…

Human-Computer Interaction · Computer Science 2024-08-27 Hongye An , Arpit Narechania , Emily Wall , Kai Xu

Contemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theory building remain underexplored. In this work, we formulate…

Computation and Language · Computer Science 2026-04-10 Peter Jansen , Peter Clark , Doug Downey , Daniel S. Weld

Collaborative work on unstructured or semi-structured documents, such as in literature corpora or source code, often involves agreed upon templates containing metadata. These templates are not consistent across users and over time.…

Digital Libraries · Computer Science 2016-08-24 Owen Kaser , Daniel Lemire

Understanding how people represent categories is a core problem in cognitive science. Decades of research have yielded a variety of formal theories of categories, but validating them with naturalistic stimuli is difficult. The challenge is…

Computer Vision and Pattern Recognition · Computer Science 2018-05-22 Joshua C. Peterson , Jordan W. Suchow , Krisha Aghi , Alexander Y. Ku , Thomas L. Griffiths

A text stream is an ordered sequence of text documents generated over time. A massive amount of such text data is generated by online social platforms every day. Designing an algorithm for such text streams to extract useful information is…

Information Retrieval · Computer Science 2024-09-04 Jay Kumar

Machine learning (ML) algorithms have become integral to decision making in various domains, including healthcare, finance, education, and law enforcement. However, concerns about fairness and bias in these systems pose significant ethical…

Machine Learning · Computer Science 2024-12-18 Ahmed Rashed , Abdelkrim Kallich , Mohamed Eltayeb

In recent times, organizations purport to undergo unprecedented transformations owing to the adoption of digital technologies. Consequently, there has been a substantial effort in academia attempting to better understand the phenomenon of…

Digital Libraries · Computer Science 2022-12-09 Mekhala Egodawele , Darshana Sedera , Vinh Bui

We illustrate the use of machine learning techniques to analyze, structure, maintain, and evolve a large online corpus of academic literature. An emerging field of research can be identified as part of an existing corpus, permitting the…

Information Retrieval · Computer Science 2009-11-10 Paul Ginsparg , Paul Houle , Thorsten Joachims , Jae-Hoon Sul

The vision of the Semantic Web (SW) is gradually unfolding and taking shape through a web of linked data, a part of which is built by capturing semantics stored in existing knowledge organization systems (KOS), subject metadata and resource…

Information Retrieval · Computer Science 2017-05-22 Aida Slavic

Most of the fastest-growing string collections today are repetitive, that is, most of the constituent documents are similar to many others. As these collections keep growing, a key approach to handling them is to exploit their…

Information Retrieval · Computer Science 2017-05-22 Travis Gagie , Aleksi Hartikainen , Kalle Karhu , Juha Kärkkäinen , Gonzalo Navarro , Simon J. Puglisi , Jouni Sirén

Structural concept complexity, class overlap, and data scarcity are some of the most important factors influencing the performance of classifiers under class imbalance conditions. When these effects were uncovered in the early 2000s,…

Machine Learning · Computer Science 2021-07-30 Kushankur Ghosh , Colin Bellinger , Roberto Corizzo , Bartosz Krawczyk , Nathalie Japkowicz

An increasing number of media streaming services have expanded their offerings to include entities of multiple content types. For instance, audio streaming services that started by offering music only, now also offer podcasts, merchandise…

Information Retrieval · Computer Science 2024-08-20 Jan Malte Lichtenberg , Giuseppe Di Benedetto , Matteo Ruffini

Due to the black-box nature of large language models (LLMs) and the realism of their generated content, issues such as hallucinations, bias, unfairness, and copyright infringement have become significant. In this context, sourcing…

Computation and Language · Computer Science 2026-01-01 Liang Pang , Jia Gu , Sunhao Dai , Zihao Wei , Zenghao Duan , Kangxi Wu , Zhiyi Yin , Jun Xu , Huawei Shen , Xueqi Cheng

We are living in an era of "big literature", where scientific literature is expanding exponentially. While this growth presents new opportunities, it complicates mapping global scientific research landscapes, as manual review methods become…

Neural and Evolutionary Computing · Computer Science 2025-11-17 Mingyu Huang , Shasha Zhou , Ke Li

Literary artefacts are generally indexed and searched based on titles, meta data and keywords over the years. This searching and indexing works well when user/reader already knows about that particular creative textual artefact or document.…

Information Retrieval · Computer Science 2021-06-28 Hrishikesh Kulkarni , Bradly Alicea

Despite the advancements in search engine features, ranking methods, technologies, and the availability of programmable APIs, current-day open-access digital libraries still rely on crawl-based approaches for acquiring their underlying…

Information Retrieval · Computer Science 2016-04-19 Sujatha Das Gollapalli , Krutarth Patel , Cornelia Caragea

The digital transformation of the scientific publishing industry has led to dramatic improvements in content discoverability and information analytics. Unfortunately, these improvements have not been uniform across research areas. The…

Digital Libraries · Computer Science 2023-06-22 Natallia Kokash , Matteo Romanello , Ernest Suyver , Giovanni Colavizza

Accurate attribution of authorship is crucial for maintaining the integrity of digital content, improving forensic investigations, and mitigating the risks of misinformation and plagiarism. Addressing the imperative need for proper…

Computers and Society · Computer Science 2026-05-27 Baixiang Huang , Canyu Chen , Kai Shu
‹ Prev 1 8 9 10 Next ›