中文
相关论文

相关论文: Semantic URL Analytics to Support Efficient Annota…

200 篇论文

Web archives are large longitudinal collections that store webpages from the past, which might be missing on the current live Web. Consequently, temporal search over such collections is essential for finding prominent missing webpages and…

信息检索 · 计算机科学 2017-02-07 Helge Holzmann , Wolfgang Nejdl , Avishek Anand

A large amount of data is present on the web. It contains huge number of web pages and to find suitable information from them is very cumbersome task. There is need to organize data in formal manner so that user can easily access and use…

信息检索 · 计算机科学 2014-03-28 Gagandeep Singh , Vishal Jain

Rapid progress in natural language processing has led to its utilization in a variety of industrial and enterprise settings, including in its use for information extraction, specifically named entity recognition and relation extraction,…

计算与语言 · 计算机科学 2021-08-13 Sharad Dixit , Varish Mulwad , Abhinav Saxena

Traditional information retrieval treats named entity recognition as a pre-indexing corpus annotation task, allowing entity tags to be indexed and used during search. Named entity taggers themselves are typically trained on thousands or…

信息检索 · 计算机科学 2018-06-14 John Foley , Sheikh Muhammad Sarwar , James Allan

Named entity recognition (NER) is highly sensitive to sentential syntactic and semantic properties where entities may be extracted according to how they are used and placed in the running text. To model such properties, one could rely on…

计算与语言 · 计算机科学 2020-10-30 Yuyang Nie , Yuanhe Tian , Yan Song , Xiang Ao , Xiang Wan

Recent advances of preservation technologies have led to an increasing number of Web archive systems and collections. These collections are valuable to explore the past of the Web, but their value can only be uncovered with effective access…

信息检索 · 计算机科学 2017-01-17 Khoi Duy Vo , Tuan Tran , Tu Ngoc Nguyen , Xiaofei Zhu , Wolfgang Nejdl

Archived collections of documents (like newspaper and web archives) serve as important information sources in a variety of disciplines, including Digital Humanities, Historical Science, and Journalism. However, the absence of efficient and…

信息检索 · 计算机科学 2021-07-30 Pavlos Fafalios , Vaibhav Kasturia , Wolfgang Nejdl

Many computer scientists use the aggregated answers of online workers to represent ground truth. Prior work has shown that aggregation methods such as majority voting are effective for measuring relatively objective features. For subjective…

计算与语言 · 计算机科学 2021-04-06 Jiele Wu , Chau-Wai Wong , Xinyan Zhao , Xianpeng Liu

Information Extraction is a well-researched area of Natural Language Processing with applications in web search and question answering concerned with identifying entities and relationships between them as expressed in a given context,…

信息检索 · 计算机科学 2020-11-17 Erin Macdonald , Denilson Barbosa

Both named entities and keywords are important in defining the content of a text in which they occur. In particular, people often use named entities in information search. However, named entities have ontological features, namely, their…

信息检索 · 计算机科学 2018-07-17 Tru H. Cao , Vuong M. Ngo

Web archives preserve unique and historically valuable information. They hold a record of past events and memories published by all kinds of people, such as journalists, politicians and ordinary people who have shared their testimony and…

数字图书馆 · 计算机科学 2021-08-04 Miguel Costa , Julien Masanès

Text extraction from web pages has many applications, including web crawling optimization and document clustering. Though much has been written about the acquisition of content from live web pages, content acquisition of archived web pages,…

数字图书馆 · 计算机科学 2016-02-24 Shawn M. Jones , Harihar Shankar

Named Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with 161K tokens…

计算与语言 · 计算机科学 2024-03-20 Siyao Peng , Zihang Sun , Huangyan Shan , Marie Kolm , Verena Blaschke , Ekaterina Artemova , Barbara Plank

Since the advent of the web, the amount of data on wen has been increased several million folds. In recent years web data generated is more than data stored for years. One important data format is text. To answer user queries over the…

信息检索 · 计算机科学 2018-11-19 Chandra Shekhar Yadav

Semantic technologies are designed to facilitate context-awareness for web content, enabling machines to understand and process them. However, this has been faced with several challenges, such as disparate nature of existing solutions and…

计算机与社会 · 计算机科学 2021-01-26 Oluwasegun Adedugbe , Elhadj Benkhelifa , Anoud Bani-Hani

In this work, we study how URL extraction results depend on input format. We compiled a pilot dataset by extracting URLs from 10 arXiv papers and used the same heuristic method to extract URLs from four formats derived from the PDF files or…

数字图书馆 · 计算机科学 2025-09-08 Rochana R. Obadage , Lamia Salsabil , Sawood Alam , Bipasha Banarjee , William A. Ingram , Edward A. Fox , Jian Wu

In this article, I present the questions that I seek to answer in my PhD research. I posit to analyze natural language text with the help of semantic annotations and mine important events for navigating large text corpora. Semantic…

信息检索 · 计算机科学 2016-03-02 Dhruv Gupta

The ability to efficiently search pictures with annotated semantics and emotion is an important problem for Human-Computer Interaction with considerable interdisciplinary significance. Accuracy and speed of the multimedia retrieval process…

人机交互 · 计算机科学 2017-07-03 Marko Horvat , Davor Kukolja , Dragutin Ivanec

With the fast growth of the Internet, more and more information is available on the Web. The Semantic Web has many features which cannot be handled by using the traditional search engines. It extracts metadata for each discovered Web…

人工智能 · 计算机科学 2011-11-30 Ahmed Tolba , Nabila Eladawi , Mohammed Elmogy

Properly annotated multimedia content is crucial for supporting advances in many Information Retrieval applications. It enables, for instance, the development of automatic tools for the annotation of large and diverse multimedia…

信息检索 · 计算机科学 2018-11-28 Xavier Favory , Eduardo Fonseca , Frederic Font , Xavier Serra