中文
相关论文

相关论文: A personalized web page content filtering model ba…

200 篇论文

Clickstreams on individual websites have been studied for decades to gain insights into user interests and to improve website experiences. This paper proposes and examines a novel sequence modeling approach for web clickstreams, that also…

人机交互 · 计算机科学 2021-03-09 Changkun Ou , Daniel Buschek , Malin Eiband , Andreas Butz

Topic segmentation is important in understanding scientific documents since it can not only provide better readability but also facilitate downstream tasks such as information retrieval and question answering by creating appropriate…

计算与语言 · 计算机科学 2023-01-06 Jeonghwan Lee , Jiyeong Han , Sunghoon Baek , Min Song

The current work is focusing on the implementation of a robust watermarking algorithm for digital images, which is based on an innovative spread spectrum analysis algorithm for watermark embedding and on a content-based image retrieval…

数据结构与算法 · 计算机科学 2009-09-29 Dimitrios K. Tsolis , Spyros Sioutas , Theodore S. Papatheodorou

With rapid increase in online information consumption, especially via social media sites, there have been concerns on whether people are getting selective exposure to a biased subset of the information space, where a user is receiving more…

社会与信息网络 · 计算机科学 2017-08-03 Abhijnan Chakraborty , Muhammad Ali , Saptarshi Ghosh , Niloy Ganguly , Krishna P. Gummadi

One of the first pre-processing steps for constructing web-scale LLM pretraining datasets involves extracting text from HTML. Despite the immense diversity of web content, existing open-source datasets predominantly apply a single fixed…

Modern social platforms are characterized by the presence of rich user-behavior data associated with the publication, sharing and consumption of textual content. Users interact with content and with each other in a complex and dynamic…

社会与信息网络 · 计算机科学 2019-02-08 Adit Krishnan , Ashish Sharma , Hari Sundaram

In this paper, we present a meta-analysis of several Web content extraction algorithms, and make recommendations for the future of content extraction on the Web. First, we find that nearly all Web content extractors do not consider a very…

信息检索 · 计算机科学 2015-08-19 Tim Weninger , Rodrigo Palacios , Valter Crescenzi , Thomas Gottron , Paolo Merialdo

Understanding the semantic meaning of content on the web through the lens of entities and concepts has many practical advantages. However, when building large-scale entity extraction systems, practitioners are facing unique challenges…

计算与语言 · 计算机科学 2021-10-04 Xuanting Cai , Quanbin Ma , Pan Li , Jianyu Liu , Qi Zeng , Zhengkan Yang , Pushkar Tripathi

Much of the existing approach to the digital divide suffers from an important limitation. It is based on a binary classification of Internet use by only considering whether someone is or is not an Internet user. To remedy this shortcoming,…

计算机与社会 · 计算机科学 2007-05-23 Eszter Hargittai

Personalization is pervasive in the online space as it leads to higher efficiency and revenue by allowing the most relevant content to be served to each user. However, recent studies suggest that personalization methods can propagate…

机器学习 · 计算机科学 2018-02-26 L. Elisa Celis , Sayash Kapoor , Farnood Salehi , Nisheeth K. Vishnoi

In recent years, we have witnessed the proliferation of large amounts of online content generated directly by users with virtually no form of external control, leading to the possible spread of misinformation. The search for effective…

信息检索 · 计算机科学 2024-07-12 Rishabh Upadhyay , Gabriella Pasi , Marco Viviani

Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model weaknesses -- and even expert-curated challenge sets quickly…

计算与语言 · 计算机科学 2026-05-27 Wenda Xu , Vilém Zouhar , Parker Riley , Mara Finkelstein , Markus Freitag , Daniel Deutsch

Pagination - the process of determining where to break an article across pages in a multi-article layout is a common layout challenge for most commercially printed newspapers and magazines. To date, no one has created an algorithm that…

计算与语言 · 计算机科学 2014-04-15 Joshua Hailpern , Niranjan Damera Venkata , Marina Danilevsky

Porous materials are widely used in different applications, in particular they are used to create various filters. Their quality depends on parameters that characterize the internal structure such as porosity, permeability and so on.…

计算机视觉与模式识别 · 计算机科学 2019-10-18 V. Kokhan , M. Grigoriev , A. Buzmakov , V. Uvarov , A. Ingacheva , E. Shvets , M. Chukalina

With the overwhelming online products available in recent years, there is an increasing need to filter and deliver relevant personalized advice for users. Recommender systems solve this problem by modeling and predicting individual…

机器学习 · 统计学 2020-02-11 Antonia Godoy-Lorite , Roger Guimera , Marta Sales-Pardo

Instance segmentation aims to delineate each individual object of interest in an image. State-of-the-art approaches achieve this goal by either partitioning semantic segmentations or refining coarse representations of detected objects. In…

计算机视觉与模式识别 · 计算机科学 2022-10-10 Long Chen , Yuli Wu , Dorit Merhof

Web attack detection is the first line of defense for securing web applications, designed to preemptively identify malicious activities. Deep learning-based approaches are increasingly popular for their advantages: automatically learning…

密码学与安全 · 计算机科学 2026-01-30 Kangqiang Luo , Yi Xie , Shiqian Zhao , Jing Pan

Majority of the computer or mobile phone enthusiasts make use of the web for searching activity. Web search engines are used for the searching; The results that the search engines get are provided to it by a software module known as the Web…

信息检索 · 计算机科学 2014-11-18 Prashant Dahiwale , M M Raghuwanshi , Latesh malik

The rise of ad-blockers is viewed as an economic threat by online publishers, especially those who primarily rely on ad- vertising to support their services. To address this threat, publishers have started retaliating by employing ad-block…

密码学与安全 · 计算机科学 2016-05-20 Muhammad Haris Mughees , Zhiyun Qian , Zubair Shafiq , Karishma Dash , Pan Hui

A large part of the hidden web resides in weblog servers. New content is produced in a daily basis and the work of traditional search engines turns to be insufficient due to the nature of weblogs. This work summarizes the structure of the…

信息检索 · 计算机科学 2009-03-25 A. Kritikopoulos , M. Sideri , I. Varlamis
‹ 上一页 1 8 9 10 下一页 ›