中文
相关论文

相关论文: HEAL: Online Incremental Recovery for Leaderless D…

200 篇论文

Training LLMs on decentralized nodes or on-spot instances, lowers the training cost and enables model democratization. The inevitable challenge here is the transient churns of nodes due to failures and the operator's scheduling policies,…

分布式、并行与集群计算 · 计算机科学 2026-04-07 Nikolay Blagoev , Oğuzhan Ersoy , Lydia Yiyu Chen

Due to the system scaling, transient errors caused by external noises, e.g., heat fluxes and particle strikes, have become a growing concern for the current and upcoming extreme-scale high-performance-computing (HPC) systems. However, since…

分布式、并行与集群计算 · 计算机科学 2021-03-10 Chao Chen , Greg Eisenhauer , Santosh Pande

Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitation of rejection sampling. Standard methods treat the teacher as a static filter, discarding complex…

人工智能 · 计算机科学 2026-03-12 Wenjing Zhang , Jiangze Yan , Jieyun Huang , Yi Shen , Shuming Shi , Ping Chen , Ning Wang , Zhaoxiang Liu , Kai Wang , Shiguo Lian

Extreme weather events have a significant impact on the aging and outdated power distribution infrastructures. These high-impact low-probability (HILP) events often result in extended outages and loss of critical services, thus, severely…

系统与控制 · 电气工程与系统科学 2019-12-11 Shiva Poudel , Anamika Dubey

In energy constrained wireless sensor networks, it is significant to make full use of the limited energy and maximize the network lifetime even when facing some unexpected situation. In this paper, all sensor nodes are grouped into…

网络与互联网体系结构 · 计算机科学 2012-01-04 Tie Qiu , Wei Wang , Feng Xia , Guowei Wu , Yu Zhou

Edge inference has become more widespread, as its diverse applications range from retail to wearable technology. Clusters of networked resource-constrained edge devices are becoming common, yet no system exists to split a DNN across these…

网络与互联网体系结构 · 计算机科学 2023-04-25 Arjun Parthasarathy , Bhaskar Krishnamachari

Edge computing faces unprecedented resource orchestration challenges from multi-dimensional heterogeneity across device architectures, diverse task requirements in CPU-intensive, GPU-intensive, I/O-intensive, and dynamic network conditions.…

分布式、并行与集群计算 · 计算机科学 2026-05-12 Jianyong Zhu , Hao Chen , Juan Zhang , Fangda Guo , Albert Y. Zomaya , Renyu Yang

Generation-driven world models create immersive virtual environments but suffer slow inference due to the iterative nature of diffusion models. While recent advances have improved diffusion model efficiency, directly applying these…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Quanjian Song , Xinyu Wang , Donghao Zhou , Jingyu Lin , Cunjian Chen , Yue Ma

Supercomputers getting ever larger and energy-efficient is at odds with the reliability of the used hardware. Thus, the time intervals between component failures are decreasing. Contrarily, the latencies for individual operations of…

分布式、并行与集群计算 · 计算机科学 2024-11-26 Demian Hespe , Lukas Hübner , Charel Mercatoris , Peter Sanders

Modern distributed systems employ aggressive optimization strategies that create latent risks - hidden vulnerabilities where exceptional performance masks catastrophic fragility when optimizations fail. Cache layers achieving 99% hit rates…

软件工程 · 计算机科学 2025-10-24 Jahidul Arafat , Kh. M. Moniruzzaman , Shamim Hossain , Fariha Tasmin

With the increasing number of compute components, failures in future exa-scale computer systems are expected to become more frequent. This motivates the study of novel resilience techniques. Here, we extend a recently proposed…

数学软件 · 计算机科学 2018-04-18 Markus Huber , Ulrich Rüde , Barbara Wohlmuth

In cloud storage systems with a large number of servers, files are typically not stored in single servers. Instead, they are split, replicated (to ensure reliability in case of server malfunction) and stored in different servers. We analyze…

分布式、并行与集群计算 · 计算机科学 2018-07-09 Avishek Ghosh , Kannan Ramchandran

Indexes are critical for efficient data retrieval and updates in modern databases. Recent advances in machine learning have led to the development of learned indexes, which model the cumulative distribution function of data to predict…

数据库 · 计算机科学 2026-04-27 Xinyi Zhang , Liang Liang , Anastasia Ailamaki , Jianliang Xu

In condition-based maintenance, real-time observations are crucial for on-line health assessment. When the monitoring system is a wireless sensor network, data loss becomes highly probable and this affects the quality of the remaining…

分布式、并行与集群计算 · 计算机科学 2016-08-23 Jacques Bahi , Wiem Elghazel , Christophe Guyeux , Mohammed Haddad , Mourad Hakem , Kamal Medjaher , Nourredine Zerhouni

With the increasing importance of data in the modern business environment, effective data man-agement and protection strategies are gaining increasing research attention. Data protection in a cloud environment is crucial for safeguarding…

分布式、并行与集群计算 · 计算机科学 2024-02-06 Ji-Beom Kim , Je-Bum Choi , Eun-Sung Jung

The idle time of personal computers has increased steadily due to the generalization of computer usage and cloud computing. Clustering research aims at utilizing idle computer resources for processing a variable workload on a large number…

分布式、并行与集群计算 · 计算机科学 2021-01-25 Geunsik Lim , Minho Lee , R. J. W. E. Lahaye , Young Ik Eom

Modern machine learning (ML) has grown into a tightly coupled, full-stack ecosystem that combines hardware, software, network, and applications. Many users rely on cloud providers for elastic, isolated, and cost-efficient resources.…

性能 · 计算机科学 2025-11-03 Ziji Chen , Steven W. D. Chien , Peng Qian , Noa Zilberman

Distributed storage systems provide large-scale reliable data storage services by spreading redundancy across a large group of storage nodes. In such a large system, node failures take place on a regular basis. When a storage node breaks…

分布式、并行与集群计算 · 计算机科学 2016-03-17 Yan Wang , Xunrui Yin , Dongsheng Wei , Xin Wang , Yucheng He

Self-healing systems depend on following a set of predefined instructions to recover from a known failure state. Failure states are generally detected based on domain specific specialized metrics. Failure fixes are applied at predefined…

分布式、并行与集群计算 · 计算机科学 2024-01-24 Mateo Sanabria , Ivana Dusparic , Nicolas Cardozo

With the increasing complexity of computing systems, complete hardware reliability can no longer be guaranteed. We need, however, to ensure overall system reliability. One of the most important features of artificial neural networks is…

神经与进化计算 · 计算机科学 2015-10-07 Anton Kulakov , Mark Zwolinski , Jeff Reeve