中文
相关论文

相关论文: Product/Brand extraction from WikiPedia

200 篇论文

In the social sciences, there is a longstanding tension between data collection methods that facilitate quantification and those that are open to unanticipated information. Advances in technology now enable new, hybrid methods that combine…

应用统计 · 统计学 2014-10-03 Matthew J. Salganik , Karen E. C. Levy

In this paper, we introduce a new NLP task -- generating short factual articles with references for queries by mining supporting evidence from the Web. In this task, called WebBrain, the ultimate goal is to generate a fluent, informative,…

计算与语言 · 计算机科学 2023-04-11 Hongjing Qian , Yutao Zhu , Zhicheng Dou , Haoqi Gu , Xinyu Zhang , Zheng Liu , Ruofei Lai , Zhao Cao , Jian-Yun Nie , Ji-Rong Wen

In this project we propose a new approach for emotion recognition using web-based similarity (e.g. confidence, PMI and PMING). We aim to extract basic emotions from short sentences with emotional content (e.g. news titles, tweets,…

计算与语言 · 计算机科学 2017-01-12 Valentina Franzoni , Giulio Biondi , Alfredo Milani , Yuanxi Li

Dynamic Bayesian networks (DBNs) offer an elegant way to integrate various aspects of language in one model. Many existing algorithms developed for learning and inference in DBNs are applicable to probabilistic language modeling. To…

计算与语言 · 计算机科学 2007-05-23 Leonid Peshkin , Avi Pfeffer

The difficulties of automatic extraction of definitions and methods from scientific documents lie in two aspects: (1) the complexity and diversity of natural language texts, which requests an analysis method to support the discovery of…

计算与语言 · 计算机科学 2023-07-06 Yutian Sun , Hai Zhuge

In this paper we present a general method for information extraction that exploits the features of data compression techniques. We first define and focus our attention on the so-called "dictionary" of a sequence. Dictionaries are…

统计力学 · 物理学 2009-11-10 A. Baronchelli , E. Caglioti , V. Loreto , E. Pizzi

Several studies have used Wikipedia (WP) data-set to analyse worldwide human preferences by languages. However, those studies could suffer from bias related to exceptional social circumstances. Any massive event promoting the exceptional…

物理与社会 · 物理学 2022-05-17 Julien Assuied , Yérali Gandica

Event classification can add valuable information for semantic search and the increasingly important topic of fact validation in news. So far, only few approaches address image classification for newsworthy event types such as natural…

计算机视觉与模式识别 · 计算机科学 2020-11-11 Eric Müller-Budack , Matthias Springstein , Sherzod Hakimov , Kevin Mrutzek , Ralph Ewerth

In this paper we present a profile-based approach to information filtering by an analysis of the content of text documents. The Wikipedia index database is created and used to automatically generate the user profile from the user document…

信息检索 · 计算机科学 2008-05-08 A. V. Smirnov , A. A. Krizhanovsky

The frequency of a web search keyword generally reflects the degree of public interest in a particular subject matter. Search logs are therefore useful resources for trend analysis. However, access to search logs is typically restricted to…

社会与信息网络 · 计算机科学 2015-09-09 Mitsuo Yoshida , Yuki Arase , Takaaki Tsunoda , Mikio Yamamoto

Knowledge about entities and their interrelations is a crucial factor of success for tasks like question answering or text summarization. Publicly available knowledge graphs like Wikidata or DBpedia are, however, far from being complete. In…

信息检索 · 计算机科学 2021-02-16 Nicolas Heist , Heiko Paulheim

Researchers and scientists increasingly find themselves in the position of having to quickly understand large amounts of technical material. Our goal is to effectively serve this need by using bibliometric text mining and summarization…

The scientific publication output grows exponentially. Therefore, it is increasingly challenging to keep track of trends and changes. Understanding scientific documents is an important step in downstream tasks such as knowledge graph…

Template detection and content extraction are two of the main areas of information retrieval applied to the Web. They perform different analyses over the structure and content of webpages to extract some part of the document. However, their…

信息检索 · 计算机科学 2022-07-19 Julián Alarte , Josep Silva

Identification of new concepts in scientific literature can help power faceted search, scientific trend analysis, knowledge-base construction, and more, but current methods are lacking. Manual identification cannot keep up with the torrent…

信息检索 · 计算机科学 2021-03-24 Daniel King , Doug Downey , Daniel S. Weld

In this paper we present our web application SeRE designed to explore semantically related concepts. Wikipedia and DBpedia are rich data sources to extract related entities for a given topic, like in- and out-links, broader and narrower…

计算与语言 · 计算机科学 2015-04-28 Daniel Hienert , Dennis Wegener , Siegfried Schomisch

Among the manifold takes on world literature, it is our goal to contribute to the discussion from a digital point of view by analyzing the representation of world literature in Wikipedia with its millions of articles in hundreds of…

信息检索 · 计算机科学 2017-01-05 Christoph Hube , Frank Fischer , Robert Jäschke , Gerhard Lauer , Mads Rosendahl Thomsen

We propose a methodology for extracting concepts for a target domain from large-scale linked open data (LOD) to support the construction of domain ontologies providing field-specific knowledge and definitions. The proposed method defines…

信息检索 · 计算机科学 2022-01-31 Satoshi Kume , Kouji Kozaki

The quality and quantity of articles in each Wikipedia language varies greatly. Translating from another Wikipedia is a natural way to add more content, but the translation process is not properly supported in the software used by…

计算与语言 · 计算机科学 2015-06-08 Niklas Laxström , Pau Giner , Santhosh Thottingal

Boilerplate refers to unwanted and repeated parts of a webpage (such as ads or table of contents) that distracts the user from reading the core content of the webpage, such as a news article. Accurate detection and removal of boilerplate…

信息检索 · 计算机科学 2020-01-15 Joy Bose