中文
相关论文

相关论文: The Many Shapes of Archive-It

200 篇论文

The Web today has millions of datasets, and the number of datasets continues to grow at a rapid pace. These datasets are not standalone entities; rather, they are intricately connected through complex relationships. Semantic relationships…

信息检索 · 计算机科学 2024-08-28 Kate Lin , Tarfah Alrashed , Natasha Noy

Research into socio-technical systems like Wikipedia has overlooked important structural patterns in the coordination of distributed work. This paper argues for a conceptual reorientation towards sequences as a fundamental unit of analysis…

社会与信息网络 · 计算机科学 2015-08-31 Brian C. Keegan , Shakked Lev , Ofer Arazy

Music engagement spans diverse interactions with music, from selection and emotional response to its impact on behavior, identity, and social connections. Social media platforms provide spaces where such engagement can be observed in…

信息检索 · 计算机科学 2025-09-25 Jatin Agarwala , George Paul , Nemani Harsha Vardhan , Vinoo Alluri

Although user access patterns on the live web are well-understood, there has been no corresponding study of how users, both humans and robots, access web archives. Based on samples from the Internet Archive's public Wayback Machine, we…

数字图书馆 · 计算机科学 2013-09-17 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

We consider the problem of classifying business process instances based on structural features derived from event logs. The main motivation is to provide machine learning based techniques with quick response times for interactive computer…

机器学习 · 计算机科学 2018-05-18 Markku Hinkka , Teemu Lehto , Keijo Heljanko , Alexander Jung

Assessing where and how information is stored in biological networks (such as neuronal and genetic networks) is a central task both in neuroscience and in molecular genetics, but most available tools focus on the network's structure as…

信息论 · 计算机科学 2023-02-24 Clifford Bohm , Douglas Kirkpatrick , Victoria Cao , Christoph Adami

Decision tree classifiers are a widely used tool in data stream mining. The use of confidence intervals to estimate the gain associated with each split leads to very effective methods, like the popular Hoeffding tree algorithm. From a…

机器学习 · 统计学 2016-04-13 Rocco De Rosa

Subject classification schemes are foundational to the organization, evaluation, and navigation of scientific knowledge. While expert-curated systems like Scopus provide widely used taxonomies, they often suffer from coarse granularity,…

数字图书馆 · 计算机科学 2025-12-30 Zhuoqi Lyu , Qing Ke

The web is often treated as a durable record of institutional and social life, yet in practice it is fragile, revisable, and frequently ephemeral. Domains change, redesigns erase earlier material, institutions relocate, maintainers…

数字图书馆 · 计算机科学 2026-05-22 Meliksah Yorulmazlar

The importance of repetitions in music is well-known. In this paper, we study music repetitions in the context of effective and efficient automatic genre classification in large-scale music-databases. We aim at enhancing the access and…

信息检索 · 计算机科学 2019-10-22 Andres Ferraro , Kjell Lemström

As the information contained within the web is increasing day by day, organizing this information could be a necessary requirement.The data mining process is to extract information from a data set and transform it into an understandable…

信息检索 · 计算机科学 2014-05-22 Prabhjot Kaur

Most archived HTML pages embed other web resources, such as images and stylesheets. Playback of the archived web pages typically provides only the capture date (or Memento-Datetime) of the root resource and not the Memento-Datetime of the…

数字图书馆 · 计算机科学 2014-10-07 Scott G. Ainsworth , Michael L. Nelson , Herbert Van de Sompel

Getting informed of what is registered in the Web space on time, can greatly help the psychologists, marketers and political analysts to familiarize, analyse, make decision and act correctly based on the society`s different needs. The great…

信息检索 · 计算机科学 2012-02-10 Mehdi Naghavi , Mohsen Sharifi

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

综合经济学 · 经济学 2023-08-07 Jens Foerderer

Document networks are found in various collections of real-world data, such as citation networks, hyperlinked web pages, and online social networks. A large number of generative models have been proposed because they offer intuitive and…

物理与社会 · 物理学 2020-01-22 Takafumi J. Suzuki

Ontologies have become the effective modeling for various applications and significantly in the semantic web. The difficulty of extracting information from the web, which was created mainly for visualising information, has driven the birth…

信息检索 · 计算机科学 2013-04-10 B. Kamala , J. M. Nandhini

This paper is a survey discussing Information Retrieval concepts, methods, and applications. It goes deep into the document and query modelling involved in IR systems, in addition to pre-processing operations such as removing stop words and…

信息检索 · 计算机科学 2012-12-11 Youssef Bassil

When searching for information, a human reader first glances over a document, spots relevant sections and then focuses on a few sentences for resolving her intention. However, the high variance of document structure complicates to identify…

计算与语言 · 计算机科学 2019-02-14 Sebastian Arnold , Rudolf Schneider , Philippe Cudré-Mauroux , Felix A. Gers , Alexander Löser

We have developed a set of Python applications that use large language models to identify and analyze data from social media platforms relevant to a population of interest. Our pipeline begins with using OpenAI's GPT-3 to generate potential…

人机交互 · 计算机科学 2023-01-16 Philip Feldman , Shimei Pan , James R. Foulds

Deep clustering uncovers hidden patterns and groups in complex time series data, yet its opaque decision-making limits use in safety-critical settings. This survey offers a structured overview of explainable deep clustering for time series,…

机器学习 · 计算机科学 2025-10-21 Udo Schlegel , Gabriel Marques Tavares , Thomas Seidl