中文
相关论文

相关论文: Loklak - A Distributed Crawler and Data Harvester …

200 篇论文

Publicly available information contains valuable information for Cyber Threat Intelligence (CTI). This can be used to prevent attacks that have already taken place on other systems. Ideally, only the initial attack succeeds and all…

密码学与安全 · 计算机科学 2025-03-25 Philipp Kuehn , Mike Schmidt , Markus Bayer , Christian Reuter

Social science research increasingly demands data-driven insights, yet researchers often face barriers such as lack of technical expertise, inconsistent data formats, and limited access to reliable datasets.Social science research…

数据库 · 计算机科学 2025-12-03 Puneet Arya , Ojas Sahasrabudhe , Adwaiya Srivastav , Partha Pratim Das , Maya Ramanath

From many datasets gathered in online social networks, well defined community structures have been observed. A large number of users participate in these networks and the size of the resulting graphs poses computational challenges. There is…

社会与信息网络 · 计算机科学 2014-01-15 Alexander V. Mantzaris

Obtaining the desired dataset is still a prime challenge faced by researchers while analyzing Online Social Network (OSN) sites. Application Programming Interfaces (APIs) provided by OSN service providers for retrieving data impose several…

社会与信息网络 · 计算机科学 2019-11-27 Mudasir Ahmad Wani , Nancy Agarwal , Suraiya Jabin , Syed Zeeshan Hussai

Although web crawlers have been around for twenty years by now, there is virtually no freely available, opensource crawling software that guarantees high throughput, overcomes the limits of single-machine systems and at the same time scales…

信息检索 · 计算机科学 2016-01-27 Paolo Boldi , Andrea Marino , Massimo Santini , Sebastiano Vigna

We present an open-source interface for scientists to explore Twitter data through interactive network visualizations. Combining data collection, transformation and visualization in one easily accessible framework, the twitter explorer…

社会与信息网络 · 计算机科学 2021-04-08 Armin Pournaki , Felix Gaisbauer , Sven Banisch , Eckehard Olbrich

Twitter is a social network that offers a rich and interesting source of information challenging to retrieve and analyze. Twitter data can be accessed using a REST API. The available operations allow retrieving tweets on the basis of a set…

信息检索 · 计算机科学 2021-10-13 Ahmad Khazaie , Nacéra Bennacer Seghouani , Francesca Bugiotti

Scientists across disciplines often use data from the internet to conduct research, generating valuable insights about human behavior. However, as generative AI relying on massive text corpora becomes increasingly valuable, platforms have…

计算机与社会 · 计算机科学 2024-12-20 Megan A. Brown , Andrew Gruen , Gabe Maldoff , Solomon Messing , Zeve Sanderson , Michael Zimmer

Blockchain represents a technology for establishing a shared, immutable version of the truth between a network of participants that do not trust one another, and therefore has the potential to disrupt any financial or other industries that…

计算机与社会 · 计算机科学 2016-11-02 Marek Laskowski , Henry M. Kim

The success of generative AI relies heavily on training on data scraped through extensive crawling of the Internet, a practice that has raised significant copyright, privacy, and ethical concerns. While few measures are designed to resist a…

人机交互 · 计算机科学 2025-05-08 Enze Liu , Elisa Luo , Shawn Shan , Geoffrey M. Voelker , Ben Y. Zhao , Stefan Savage

Twitter data have become essential to Natural Language Processing (NLP) and social science research, driving various scientific discoveries in recent years. However, the textual data alone are often not enough to conduct studies: especially…

计算与语言 · 计算机科学 2022-01-27 Federico Bianchi , Vincenzo Cutrona , Dirk Hovy

Dark web crawling is a complex process that involves specific methodologies and techniques to navigate the Tor network and extract data from hidden services. This study proposes a general dark web crawler designed to extract pages handling…

密码学与安全 · 计算机科学 2024-05-13 Daniel De Pascale , Giuseppe Cascavilla , Damian A. Tamburri , Willem-Jan Van Den Heuvel

Researchers in the Digital Humanities and journalists need to monitor, collect and analyze fresh online content regarding current events such as the Ebola outbreak or the Ukraine crisis on demand. However, existing focused crawling…

数字图书馆 · 计算机科学 2017-07-31 Gerhard Gossen , Elena Demidova , Thomas Risse

Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex,…

计算与语言 · 计算机科学 2025-08-12 Jialong Wu , Wenbiao Yin , Yong Jiang , Zhenglin Wang , Zekun Xi , Runnan Fang , Linhai Zhang , Yulan He , Deyu Zhou , Pengjun Xie , Fei Huang

This paper presents twAwler, a lightweight twitter crawler that targets language-specific communities of users. twAwler takes advantage of multiple endpoints of the twitter API to explore user relations and quickly recognize users belonging…

社会与信息网络 · 计算机科学 2018-04-23 Polyvios Pratikakis

Nowadays, the size of the Internet is experiencing rapid growth. As of December 2014, the number of global Internet websites has more than 1 billion and all kinds of information resources are integrated together on the Internet, however,the…

分布式、并行与集群计算 · 计算机科学 2015-06-02 Qingpei Guo , Chao Xu , Yang Song

[WhitePaper] The TelegramScrap tool provides a robust and versatile solution for extracting and analyzing data from Telegram channels and groups, addressing the increasing demand for efficient methods to study digital ecosystems. This white…

计算机与社会 · 计算机科学 2024-12-24 Ergon Cugler de Moraes Silva

Mining the silent members of an online community, also called lurkers, has been recognized as an important problem that accompanies the extensive use of online social networks (OSNs). Existing solutions to the ranking of lurkers can aid…

社会与信息网络 · 计算机科学 2015-09-08 Andrea Tagarelli , Roberto Interdonato

In this paper, we introduce a novel, general purpose, technique for faster sampling of nodes over an online social network. Specifically, unlike traditional random walk which wait for the convergence of sampling distribution to a…

社会与信息网络 · 计算机科学 2014-11-04 Azade Nazi , Zhuojie Zhou , Saravanan Thirumuruganathan , Nan Zhang , Gautam Das

We describe a novel, "focusable", scalable, distributed web crawler based on GNU/Linux and PostgreSQL that we designed to be easily extendible and which we have released under a GNU public licence. We also report a first use case related to…

信息检索 · 计算机科学 2012-12-27 Pierre Joulin , Romain Deveaud , Eric SanJuan-Ibekwe , Jean-Marc Francony , Françoise Para