中文
相关论文

相关论文: Popularity Driven Data Integration

200 篇论文

As the amount of scientific data continues to grow at ever faster rates, the research community is increasingly in need of flexible computational infrastructure that can support the entirety of the data science lifecycle, including…

计算机与社会 · 计算机科学 2016-04-12 Robert L. Grossman , Allison Heath , Mark Murphy , Maria Patterson , Walt Wells

Current research on Internet of Things (IoT) mainly focuses on how to enable general objects to see, hear, and smell the physical world for themselves, and make them connected to share the observations. In this paper, we argue that only…

分布式、并行与集群计算 · 计算机科学 2013-11-19 Guoru Ding , Long Wang , Qihui Wu

The Internet of Things (IoT) envisions a world-wide, interconnected network of smart physical entities. These physical entities generate a large amount of data in operation and as the IoT gains momentum in terms of deployment, the combined…

数据库 · 计算机科学 2018-07-04 Eugene Siow , Thanassis Tiropanis , Wendy Hall

In a world of information overload, understanding how we can most effectively manage information is crucial to success. We set out to understand how people view deletion, the removal of material no longer needed: does it help by reducing…

人机交互 · 计算机科学 2026-01-01 Paul Englefield , Russell Beale

Widespread use of the Internet and social networks invokes the generation of big data, which is proving to be useful in a number of applications. To deal with explosively growing amounts of data, data analytics has emerged as a critical…

信息论 · 计算机科学 2016-11-17 Kwang-Cheng Chen , Shao-Lun Huang , Lizhong Zheng , H. Vincent Poor

Multiple imputation is a highly recommended technique to deal with missing data, but the application to longitudinal datasets can be done in multiple ways. When a new wave of longitudinal data arrives, we can treat the combined data of…

统计方法学 · 统计学 2026-05-18 X. M. Kavelaars , S. van Buuren , J. R. van Ginkel

Crowdsourcing allows running simple human intelligence tasks on a large crowd of workers, enabling solving problems for which it is difficult to formulate an algorithm or train a machine learning model in reasonable time. One of such…

人机交互 · 计算机科学 2023-06-05 Daniil Likhobaba , Daniil Fedulov , Dmitry Ustalov

This paper explores a prevailing trend in the industry: migrating data-intensive analytics applications from on-premises to cloud-native environments. We find that the unique cost models associated with cloud-based storage necessitate a…

分布式、并行与集群计算 · 计算机科学 2023-11-02 Chunxu Tang , Yi Wang , Bin Fan , Beinan Wang , Shouwei Chen , Ziyue Qiu , Chen Liang , Jing Zhao , Yu Zhu , Mingmin Chen , Zhongting Hu

Scientific research requires access, analysis, and sharing of data that is distributed across various heterogeneous data sources at the scale of the Internet. An eager ETL process constructs an integrated data repository as its first step,…

数据库 · 计算机科学 2018-04-25 Pradeeban Kathiravelu , Ashish Sharma , Helena Galhardas , Peter Van Roy , Luıs Veiga

In the field of machine learning there is a growing interest towards more robust and generalizable algorithms. This is for example important to bridge the gap between the environment in which the training data was collected and the…

机器学习 · 计算机科学 2020-10-08 Wim Casteels , Peter Hellinckx

Collaborative tagging has recently attracted the attention of both industry and academia due to the popularity of content-sharing systems such as CiteULike, del.icio.us, and Flickr. These systems give users the opportunity to add data items…

数字图书馆 · 计算机科学 2007-06-24 Elizeu Santos-Neto , Matei Ripeanu , Adriana Iamnitchi

Data Lake (DL) is a Big Data analysis solution which ingests raw data in their native format and allows users to process these data upon usage. Data ingestion is not a simple copy and paste of data, it is a complicated and important phase…

数据库 · 计算机科学 2021-07-08 Yan Zhao , Imen Megdiche , Franck Ravat

To evaluate Information Retrieval (IR) effectiveness, a possible approach is to use test collections, which are composed of a collection of documents, a set of description of information needs (called topics), and a set of relevant…

信息检索 · 计算机科学 2020-11-03 Kevin Roitero

Product bundling aims to organize a set of thematically related items into a combined bundle for shipment facilitation and item promotion. To increase the exposure of fresh or overstocked products, sellers typically bundle these items with…

信息检索 · 计算机科学 2024-12-02 Shuo Xu , Haokai Ma , Yunshan Ma , Xiaohao Liu , Lei Meng , Xiangxu Meng , Tat-Seng Chua

Clustering is a widely used technique in data mining applications for discovering patterns in underlying data. Most traditional clustering algorithms are limited to handling datasets that contain either numeric or categorical attributes.…

人工智能 · 计算机科学 2007-05-23 Zengyou He , Xiaofei Xu , Shengchun Deng

Clustering is often used for discovering structure in data. Clustering systems differ in the objective function used to evaluate clustering quality and the control strategy used to search the space of clusterings. Ideally, the search…

人工智能 · 计算机科学 2014-11-17 D. Fisher

With the increasing use of big data and business analytics, data storytelling has gained popularity as an effective means of communicating analytical insights to audiences to support decision making and improve business performance.…

计算机与社会 · 计算机科学 2024-02-08 Valeria Zitz , Patrick Baier

Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density,…

In this study, we conducted semi-structured interviews with 21 IIR researchers to investigate their data reuse practices. This study aims to expand upon current findings by exploring IIR researchers' information-obtaining behaviors…

信息检索 · 计算机科学 2025-12-23 Tianji Jiang , Wenqi Li , Jiqun Liu

Generative retrieval for search and recommendation is a promising paradigm for retrieving items, offering an alternative to traditional methods that depend on external indexes and nearest-neighbor searches. Instead, generative models…

信息检索 · 计算机科学 2024-10-23 Gustavo Penha , Ali Vardasbi , Enrico Palumbo , Marco de Nadai , Hugues Bouchard