中文
相关论文

相关论文: WikiDoMiner: Wikipedia Domain-specific Miner

200 篇论文

Wikidata, like Wikipedia, is a knowledge base that anyone can edit. This open collaboration model is powerful in that it reduces barriers to participation and allows a large number of people to contribute. However, it exposes the knowledge…

信息检索 · 计算机科学 2017-03-14 Amir Sarabadani , Aaron Halfaker , Dario Taraborelli

HITS adapted algorithm for synonym search, the program architecture, and the program work evaluation with test examples are presented in the paper. Synarcher program for synonym (and related terms) search in the text corpus of special…

信息检索 · 计算机科学 2007-05-23 A. Krizhanovsky

A growing body of work has highlighted the important role that Wikipedia's volunteer-created content plays in helping search engines achieve their core goal of addressing the information needs of millions of people. In this paper, we report…

计算机与社会 · 计算机科学 2020-04-23 Nicholas Vincent , Brent Hecht

The traditional entity extraction problem lies in the ability of extracting named entities from plain text using natural language processing techniques and intensive training from large document collections. Examples of named entities…

信息检索 · 计算机科学 2007-11-21 Anne-Marie Vercoustre , James A. Thom , Jovan Pehcevski

We introduce an open-domain topic classification system that accepts user-defined taxonomy in real time. Users will be able to classify a text snippet with respect to any candidate labels they want, and get instant response from our web…

计算与语言 · 计算机科学 2023-07-03 Hantian Ding , Jinrui Yang , Yuqian Deng , Hongming Zhang , Dan Roth

Very important breakthroughs in data centric deep learning algorithms led to impressive performance in transactional point applications of Artificial Intelligence (AI) such as Face Recognition, or EKG classification. With all due…

人工智能 · 计算机科学 2018-05-23 Moshe BenBassat

Wikipedia's contents are based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close…

数字图书馆 · 计算机科学 2020-11-24 Harshdeep Singh , Robert West , Giovanni Colavizza

With the exponential increase in online scientific literature, identifying reliable domain-specific data has become increasingly important but also very challenging. Manual data collection and filtering for domain-specific scientific…

信息检索 · 计算机科学 2026-03-10 Nikita Gautam , Doina Caragea , Ignacio Ciampitti , Federico Gomez

Hyperlinks constitute the backbone of the Web; they enable user navigation, information discovery, content ranking, and many other crucial services on the Internet. In particular, hyperlinks found within Wikipedia allow the readers to…

计算机与社会 · 计算机科学 2021-06-01 Martin Gerlach , Marshall Miller , Rita Ho , Kosta Harlan , Djellel Difallah

Today's research progress in the field of multi-document summarization is obstructed by the small number of available datasets. Since the acquisition of reference summaries is costly, existing datasets contain only hundreds of samples at…

计算与语言 · 计算机科学 2020-02-18 Diego Antognini , Boi Faltings

We present Wikipedia-based Polyglot Dirichlet Allocation (WikiPDA), a crosslingual topic model that learns to represent Wikipedia articles written in any language as distributions over a common set of language-independent topics. It…

计算与语言 · 计算机科学 2021-02-16 Tiziano Piccardi , Robert West

The limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In…

计算与语言 · 计算机科学 2022-07-25 Haichao Zhu , Li Dong , Furu Wei , Bing Qin , Ting Liu

Aspect-based summarization is the task of generating focused summaries based on specific points of interest. Such summaries aid efficient analysis of text, such as quickly understanding reviews or opinions from different angles. However,…

计算与语言 · 计算机科学 2020-11-17 Hiroaki Hayashi , Prashant Budania , Peng Wang , Chris Ackerson , Raj Neervannan , Graham Neubig

Wikipedia is among the largest examples of collective intelligence on the Web with over 61 million articles covering over 320 languages. Although edited and maintained by an active workforce of human volunteers, Wikipedia is highly reliant…

人机交互 · 计算机科学 2025-09-29 Neal Reeves , Elena Simperl

In this work, we study disagreements in discussions around Wikidata, an online knowledge community that builds the data backend of Wikipedia. Discussions are essential in collaborative work as they can increase contributor performance and…

人机交互 · 计算机科学 2025-05-20 Elisavet Koutsiana , Tushita Yadav , Nitisha Jain , Albert Meroño-Peñuela , Elena Simperl

Despite being vast repositories of factual information, cross-domain knowledge graphs, such as Wikidata and the Google Knowledge Graph, only sparsely provide short synoptic descriptions for entities. Such descriptions that briefly identify…

计算与语言 · 计算机科学 2019-04-17 Rajarshi Bhowmik , Gerard de Melo

Knowledge graphs have recently become the state-of-the-art tool for representing the diverse and complex knowledge of the world. Examples include the proprietary knowledge graphs of companies such as Google, Facebook, IBM, or Microsoft, but…

人工智能 · 计算机科学 2020-02-28 Tom Hanika , Maximilian Marx , Gerd Stumme

Deploying dense retrieval models efficiently is becoming increasingly important across various industries. This is especially true for enterprise search services, where customizing search engines to meet the time demands of different…

信息检索 · 计算机科学 2024-01-24 Chen Huang , Duanyu Feng , Wenqiang Lei , Jiancheng Lv

There is much debate on how public participation and expertise can be brought together in collaborative knowledge environments. One of the experiments addressing the issue directly is Citizendium. In seeking to harvest the strengths (and…

数字图书馆 · 计算机科学 2010-08-25 Tom Morris , Daniel Mietchen

Deep Research Agents (DRAs) have demonstrated remarkable capabilities in autonomous information retrieval and report generation, showing great potential to assist humans in complex research tasks. Current evaluation frameworks primarily…

计算与语言 · 计算机科学 2026-02-04 Shaohan Wang , Benfeng Xu , Licheng Zhang , Mingxuan Du , Chiwei Zhu , Xiaorui Wang , Zhendong Mao , Yongdong Zhang