中文
相关论文

相关论文: Benchmarking the Open Science Data Federation serv…

200 篇论文

The adoption of heterogeneous computing systems based on diverse architectures to achieve exascale computing power has worsened the performance portability problem of scientific applications that were designed to run on these platforms. To…

分布式、并行与集群计算 · 计算机科学 2023-10-17 Ami Marowka

Research Data Management (RDM) is essential in handling and organizing data in the research field. The Berlin Open Science Platform (BOP) serves as a case study that exemplifies the significance of standardization within the Berlin…

数字图书馆 · 计算机科学 2024-05-24 Sefika Efeoglu , Zongxiong Chen , Sonja Schimmler , Bianca Wentzel

While federated learning (FL) is a widely popular distributed machine learning (ML) strategy that protects data privacy, time-varying wireless network parameters and heterogeneous configurations of the wireless devices pose significant…

机器学习 · 计算机科学 2025-08-28 Ferdous Pervej , Minseok Choi , Andreas F. Molisch

RDF streaming has been explored by the Semantic Web community from many angles, resulting in multiple task formulations and streaming methods. However, for many existing formulations of the problem, reliably benchmarking streaming solutions…

数据库 · 计算机科学 2023-11-28 Piotr Sowinski , Maria Ganzha , Marcin Paprzycki

Input data for applications that run in cloud computing centres can be stored at distant repositories, often with multiple copies of the popular data stored at many sites. Locating and retrieving the remote data can be challenging, and we…

Remote sensing semantic segmentation (RSS) is an essential technology in earth observation missions. Due to concerns over geographic information security, data privacy, storage bottleneck and industry competition, high-quality annotated…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Jieyi Tan , Yansheng Li , Sergey A. Bartalev , Shinkarenko Stanislav , Bo Dang , Yongjun Zhang , Liangqi Yuan , Wei Chen

Scientific data management is at a critical juncture, driven by exponential data growth, increasing cross-domain dependencies, and a severe reproducibility crisis in modern research. Traditional centralized data management approaches are…

Research challenges such as climate change and the search for habitable planets increasingly use academic and commercial computing resources distributed across different institutions and physical sites. Furthermore, such analyses often…

密码学与安全 · 计算机科学 2023-05-16 Richard Cardone , Smruti Padhy , Steven Black , Sean Cleveland , Joe Stubbs

In dataspaces, federation services facilitate key functions such as enabling participating organizations to establish mutual trust and assisting them in discovering data and services available for consumption. Discovery is enabled by a…

Decentralized federated learning (DFL), inherited from distributed optimization, is an emerging paradigm to leverage the explosively growing data from wireless devices in a fully distributed manner.DFL enables joint training of machine…

信号处理 · 电气工程与系统科学 2023-10-10 Zhiyuan Zhai , Xiaojun Yuan , Xin Wang

Machine learning relies on the availability of a vast amount of data for training. However, in reality, most data are scattered across different organizations and cannot be easily integrated under many legal and practical constraints. In…

机器学习 · 计算机科学 2020-06-25 Yang Liu , Yan Kang , Chaoping Xing , Tianjian Chen , Qiang Yang

The Open Science Grid (OSG) includes work to enable new science, new scientists, and new modalities in support of computationally based research. There are frequently significant sociological and organizational changes required in…

The conjunction of edge intelligence and the ever-growing Internet-of-Things (IoT) network heralds a new era of collaborative machine learning, with federated learning (FL) emerging as the most prominent paradigm. With the growing interest…

机器学习 · 计算机科学 2024-11-25 Nizar Masmoudi , Wael Jaafar

Data science pipelines commonly utilize dataframe and array operations for tasks such as data preprocessing, analysis, and machine learning. The most popular tools for these tasks are pandas and NumPy. However, these tools are limited to…

分布式、并行与集群计算 · 计算机科学 2024-03-20 Weizheng Lu , Kaisheng He , Xuye Qin , Chengjie Li , Zhong Wang , Tao Yuan , Xia Liao , Feng Zhang , Yueguo Chen , Xiaoyong Du

Combining the results of different search engines in order to improve upon their performance has been the subject of many research papers. This has become known as the "Data Fusion" task, and has great promise in dealing with the vast…

信息检索 · 计算机科学 2018-02-13 Weinan Huang , Junyi Chen , Lei Meng , David Lillis

Different from the traditional benchmarking methodology that creates a new benchmark or proxy for every possible workload, this paper presents a scalable big data benchmarking methodology. Among a wide variety of big data analytics…

硬件体系结构 · 计算机科学 2017-11-10 Wanling Gao , Lei Wang , Jianfeng Zhan , Chunjie Luo , Daoyi Zheng , Zhen Jia , Biwei Xie , Chen Zheng , Qiang Yang , Haibin Wang

In Federated Learning (FL) with over-the-air aggregation, the quality of the signal received at the server critically depends on the receive scaling factors. While a larger scaling factor can reduce the effective noise power and improve…

信息论 · 计算机科学 2025-10-07 Faeze Moradi Kalarde , Ben Liang , Min Dong , Yahia A. Eldemerdash Ahmed , Ho Ting Cheng

Today's big data science communities manage their data publication and replication at the application layer. These communities utilize myriad mechanisms to publish, discover, and retrieve datasets - the result is an ecosystem of either…

网络与互联网体系结构 · 计算机科学 2022-11-03 Justin Presley , Xi Wang , Tym Brandel , Xusheng Ai , Proyash Podder , Tianyuan Yu , Varun Patil , Lixia Zhang , Alex Afanasyev , F. Alex Feltus , Susmit Shannigrahi

We document the data transfer workflow, data transfer performance, and other aspects of staging approximately 56 terabytes of climate model output data from the distributed Coupled Model Intercomparison Project (CMIP5) archive to the…

分布式、并行与集群计算 · 计算机科学 2017-09-28 Eli Dart , Michael F. Wehner , Prabhat

Scientific research increasingly relies on distributed computational resources, storage systems, networks, and instruments, ranging from HPC and cloud systems to edge devices. Event-driven architecture (EDA) benefits applications targeting…

分布式、并行与集群计算 · 计算机科学 2024-10-01 Haochen Pan , Ryan Chard , Sicheng Zhou , Alok Kamatar , Rafael Vescovi , Valérie Hayot-Sasson , André Bauer , Maxime Gonthier , Kyle Chard , Ian Foster