中文
相关论文

相关论文: MITAO: a tool for enabling scholars in the Humanit…

200 篇论文

Interpretive scholars generate knowledge from text corpora by manually sampling documents, applying codes, and refining and collating codes into categories until meaningful themes emerge. Given a large corpus, machine learning could help…

Topic models are a family of statistical-based algorithms to summarize, explore and index large collections of text documents. After a decade of research led by computer scientists, topic models have spread to social science as a new…

计算与语言 · 计算机科学 2018-04-04 Ryan Wesslen

The connection between texts is referred to as intertextuality in literary theory, which served as an important theoretical basis in many digital humanities studies. Over the past decade, advancements in natural language processing have…

计算与语言 · 计算机科学 2025-11-03 Siyu Duan

Topic models are widely used analysis techniques for clustering documents and surfacing thematic elements of text corpora. These models remain challenging to optimize and often require a "human-in-the-loop" approach where domain experts use…

人机交互 · 计算机科学 2021-01-08 Anamaria Crisan , Michael Correll

Human-in-the-loop topic modelling incorporates users' knowledge into the modelling process, enabling them to refine the model iteratively. Recent research has demonstrated the value of user feedback, but there are still issues to consider,…

计算与语言 · 计算机科学 2023-04-05 Zheng Fang , Lama Alqazlan , Du Liu , Yulan He , Rob Procter

Topic modelling has become increasingly popular for summarizing text data, such as social media posts and articles. However, topic modelling is usually completed in one shot. Assessing the quality of resulting topics is challenging. No…

This paper aims to answer one central question: to what extent can open-source generative text models be used in a workflow to approximate thematic analysis in social science research? To answer this question, we present the Generative…

计算与语言 · 计算机科学 2024-10-08 Andrew Katz , Gabriella Coloyan Fleming , Joyce Main

The essential task of Topic Detection and Tracking (TDT) is to organize a collection of news media into clusters of stories that pertain to the same real-world event. To apply TDT models to practical applications such as search engines and…

信息检索 · 计算机科学 2021-10-15 Doug Beeferman , Hang Jiang

Textual content (including titles, annotations, and captions) plays a central role in helping readers understand a visualization by emphasizing, contextualizing, or summarizing the depicted data. Yet, existing visualization tools provide…

人机交互 · 计算机科学 2025-02-12 Arjun Srinivasan , Vidya Setlur , Arvind Satyanarayan

We are presenting a set of multilingual text analysis tools that can help analysts in any field to explore large document collections quickly in order to determine whether the documents contain information of interest, and to find the…

计算与语言 · 计算机科学 2007-05-23 Camelia Ignat , Bruno Pouliquen , Ralf Steinberger , Tomaz Erjavec

Topic modelling is a text mining technique for identifying salient themes from a number of documents. The output is commonly a set of topics consisting of isolated tokens that often co-occur in such documents. Manual effort is often…

计算与语言 · 计算机科学 2024-04-26 Lowri Williams , Eirini Anthi , Laura Arman , Pete Burnap

Ancient Chinese texts present an area of enormous challenge and opportunity for humanities scholars interested in exploiting computational methods to assist in the development of new insights and interpretations of culturally significant…

计算与语言 · 计算机科学 2017-02-06 Colin Allen , Hongliang Luo , Jaimie Murdock , Jianghuai Pu , Xiaohong Wang , Yanjie Zhai , Kun Zhao

Topic models represent groups of documents as a list of words (the topic labels). This work asks whether an alternative approach to topic labeling can be developed that is closer to a natural language description of a topic than a word…

计算与语言 · 计算机科学 2022-11-11 Domenic Rosati

While many researchers use Large Language Models (LLMs) through chat-based access, their real potential lies in leveraging LLMs via application programming interfaces (APIs). This paper conceptualizes LLMs as universal text processing…

计算与语言 · 计算机科学 2026-03-23 Ivan Zupic

In this paper, we present Pilaster (https://visusal.github.io/pilaster/), a collection of citation metadata extracted from publications in visualization for the digital humanities. The collection is generated from a seed set of relevant…

人机交互 · 计算机科学 2020-09-08 Alejandro Benito-Santos , Roberto Therón

Qualitative research is an approach to understanding social phenomenon based around human interpretation of data, particularly text. Probabilistic topic modelling is a machine learning approach that is also based around the analysis of text…

人机交互 · 计算机科学 2022-10-04 Marco Gillies , Dhiraj Murthy , Harry Brenton , Rapheal Olaniyan

With massive texts on social media, users and analysts often rely on topic modeling techniques to quickly extract key themes and gain insights. Traditional topic modeling techniques, such as Latent Dirichlet Allocation (LDA), provide…

数据库 · 计算机科学 2025-08-12 Fei Ye , Jiapan Liu , Yinan Jing , Zhenying He , Weirao Wang , X. Sean Wang

The number of documents available into Internet moves each day up. For this reason, processing this amount of information effectively and expressibly becomes a major concern for companies and scientists. Methods that represent a textual…

A lot of real-world phenomena are complex and cannot be captured by single task annotations. This causes a need for subsequent annotations, with interdependent questions and answers describing the nature of the subject at hand. Even in the…

计算与语言 · 计算机科学 2020-10-05 Moritz Wolf , Dana Ruiter , Ashwin Geet D'Sa , Liane Reiners , Jan Alexandersson , Dietrich Klakow

We present Twitmo, a package that provides a broad range of methods to collect, pre-process, analyze and visualize geo-tagged Twitter data. Twitmo enables the user to collect geo-tagged Tweets from Twitter and and provides a comprehensive…

‹ 上一页 1 2 3 10 下一页 ›