中文
相关论文

相关论文: Product/Brand extraction from WikiPedia

200 篇论文

The traditional entity extraction problem lies in the ability of extracting named entities from plain text using natural language processing techniques and intensive training from large document collections. Examples of named entities…

信息检索 · 计算机科学 2007-11-21 Anne-Marie Vercoustre , James A. Thom , Jovan Pehcevski

Keyphrase is an efficient representation of the main idea of documents. While background knowledge can provide valuable information about documents, they are rarely incorporated in keyphrase extraction methods. In this paper, we propose…

计算与语言 · 计算机科学 2018-03-28 Yang Yu , Vincent Ng

The World Wide Web caters to the needs of billions of users in heterogeneous groups. Each user accessing the World Wide Web might have his / her own specific interest and would expect the web to respond to the specific requirements. The…

信息检索 · 计算机科学 2017-11-22 K. S. Kuppusamy , G. Aghila

In this paper, we propose a method to extract descriptions of technical terms from Web pages in order to utilize the World Wide Web as an encyclopedia. We use linguistic patterns and HTML text structures to extract text fragments containing…

计算与语言 · 计算机科学 2007-05-23 Atsushi Fujii , Tetsuya Ishikawa

We present a new concept - Wikiometrics - the derivation of metrics and indicators from Wikipedia. Wikipedia provides an accurate representation of the real world due to its size, structure, editing policy and popularity. We demonstrate an…

数字图书馆 · 计算机科学 2016-01-11 Gilad Katz , Lior Rokach

One of the most impressive human endeavors of the past two decades is the collection and categorization of human knowledge in the free and accessible format that is Wikipedia. In this work we ask what makes a term worthy of entering this…

计算与语言 · 计算机科学 2020-09-18 Yonatan Bilu , Shai Gretz , Edo Cohen , Noam Slonim

Wikipedia is a useful knowledge source that benefits many applications in language processing and knowledge representation. An important feature of Wikipedia is that of categories. Wikipedia pages are assigned different categories according…

计算与语言 · 计算机科学 2017-04-26 Yanqing Chen , Steven Skiena

The extraction of individual reference strings from the reference section of scientific publications is an important step in the citation extraction pipeline. Current approaches divide this task into two steps by first detecting the…

信息检索 · 计算机科学 2017-05-24 Martin Körner

When it comes to factual knowledge about a wide range of domains, Wikipedia is often the prime source of information on the web. DBpedia and YAGO, as large cross-domain knowledge graphs, encode a subset of that knowledge by creating an…

信息检索 · 计算机科学 2020-04-02 Nicolas Heist , Heiko Paulheim

Online user profiling is a very active research field, catalyzing great interest by both scientists and practitioners. In this paper, in particular, we look at approaches able to mine social media activities of users to create a rich user…

信息检索 · 计算机科学 2018-09-27 Christian Torrero , Carlo Caprini , Daniele Miorandi

Performing effective preference-based data retrieval requires detailed and preferentially meaningful structurized information about the current user as well as the items under consideration. A common problem is that representations of items…

人工智能 · 计算机科学 2011-01-13 Joachim Selke , Wolf-Tilo Balke

Extraction of financial and economic events from text has previously been done mostly using rule-based methods, with more recent works employing machine learning techniques. This work is in line with this latter approach, leveraging…

This report argues that, even in the simplest cases, IE is an ontology-driven process. It is not a mere text filtering method based on simple pattern matching and keywords, because the extracted pieces of texts are interpreted with respect…

人工智能 · 计算机科学 2016-08-16 Claire Nédellec , Adeline Nazarenko

Cross-document event coreference resolution is a foundational task for NLP applications involving multi-text processing. However, existing corpora for this task are scarce and relatively small, while annotating only modest-size clusters of…

计算与语言 · 计算机科学 2021-05-03 Alon Eirew , Arie Cattan , Ido Dagan

Text is the main method of communicating information in the digital age. Messages, blogs, news articles, reviews, and opinionated information abound on the Internet. People commonly purchase products online and post their opinions about…

计算与语言 · 计算机科学 2014-04-09 Amani K Samha , Yuefeng Li , Jinglan Zhang

A bag-of-words based probabilistic classifier is trained using regularized logistic regression to detect vandalism in the English Wikipedia. Isotonic regression is used to calibrate the class membership probabilities. Learning curve,…

机器学习 · 计算机科学 2010-01-06 Amit Belani

A Wikipedia book (known as Wikibook) is a collection of Wikipedia articles on a particular theme that is organized as a book. We propose Wikibook-Bot, a machine-learning based technique for automatically generating high quality Wikibooks…

数字图书馆 · 计算机科学 2018-12-31 Shahar Admati , Lior Rokach , Bracha Shapira

Online encyclopedia such as Wikipedia has become one of the best sources of knowledge. Much effort has been devoted to expanding and enriching the structured data by automatic information extraction from unstructured text in Wikipedia.…

信息检索 · 计算机科学 2014-06-26 Kezun Zhang , Yanghua Xiao , Hanghang Tong , Haixun Wang , Wei Wang

We report on work in progress on extracting lexical simplifications (e.g., "collaborate" -> "work together"), focusing on utilizing edit histories in Simple English Wikipedia for this task. We consider two main approaches: (1) deriving…

计算与语言 · 计算机科学 2010-08-13 Mark Yatskar , Bo Pang , Cristian Danescu-Niculescu-Mizil , Lillian Lee

Wikipedia categories, a classification scheme built for organizing and describing Wikpedia articles, are being applied in computer science research. This paper adopts a systematic literature review approach, in order to identify different…

数字图书馆 · 计算机科学 2020-04-22 Jesús Tramullas , Piedad Garrido-Picazo , Ana I. Sánchez-Casabón
‹ 上一页 1 2 3 10 下一页 ›