中文
相关论文

相关论文: Tail Index for a Distributed Storage System with P…

200 篇论文

Performance of distributed optimization and learning systems is bottlenecked by "straggler" nodes and slow communication links, which significantly delay computation. We propose a distributed optimization framework where the dataset is…

机器学习 · 统计学 2018-03-15 Can Karakus , Yifan Sun , Suhas Diggavi , Wotao Yin

This paper considers learning deep features from long-tailed data. We observe that in the deep feature space, the head classes and the tail classes present different distribution patterns. The head classes have a relatively large spatial…

计算机视觉与模式识别 · 计算机科学 2020-04-14 Jialun Liu , Yifan Sun , Chuchu Han , Zhaopeng Dou , Wenhui Li

We analyze the (computational) complexity distribution of sphere-decoding (SD) for random infinite lattices. In particular, we show that under fairly general assumptions on the statistics of the lattice basis matrix, the tail behavior of…

信息论 · 计算机科学 2016-11-15 Dominik Seethaler , Joakim Jaldén , Christoph Studer , Helmut Bölcskei

We introduce a new actuarial tail-shape index, the $\theta$-index, based on a probability equal level relationship between Value at Risk and Expected Shortfall. The index is defined at each tail probability level as the parameter value for…

风险管理 · 定量金融 2026-01-29 Georgios I. Papayiannis , Georgios Psarrakos

Erasure coding has been recognized as a powerful method to mitigate delays due to slow or straggling nodes in distributed systems. This work shows that erasure coding of data objects can flexibly handle skews in the request rates. Coding…

信息论 · 计算机科学 2021-06-29 Mehmet Aktas , Gauri Joshi , Swanand Kadhe , Fatemeh Kazemi , Emina Soljanin

One potential solution to combat the scarcity of tail observations in extreme value analysis is to integrate information from multiple datasets sharing similar tail properties, for instance, a common extreme value index. In other words, for…

统计方法学 · 统计学 2025-06-25 Liujun Chen , Marco Oesting , Chen Zhou

This paper proposes a Mixture Density Network specifically designed for forecasting time series that exhibit locally explosive behavior. By incorporating skewed t-distributions as mixture components, our approach offers enhanced flexibility…

统计方法学 · 统计学 2026-02-11 Elena Dumitrescu , Julien Peignon , Arthur Thomas

Large Language Model-based Recommender Systems (LRSs) have recently emerged as a new paradigm in sequential recommendation by directly adopting LLMs as backbones. While LRSs demonstrate strong knowledge utilization and instruction-following…

信息检索 · 计算机科学 2026-03-16 Jiaming Zhang , Yuyuan Li , Xiaohua Feng , Li Zhang , Longfei Li , Jun Zhou , Chaochao Chen

Scientific computing workflows generate enormous distributed data that is short-lived, yet critical for job completion time. This class of data is called intermediate data. A common way to achieve high data availability is to replicate…

分布式、并行与集群计算 · 计算机科学 2020-04-14 Zhe Zhang , Brian Bockelman , Derek Weitzel , David Swanson

This paper introduces a new classification scheme - head/tail breaks - in order to find groupings or hierarchy for data with a heavy-tailed distribution. The heavy-tailed distributions are heavily right skewed, with a minority of large…

数据分析、统计与概率 · 物理学 2013-10-22 Bin Jiang

Peer-to-peer distributed storage systems provide reliable access to data through redundancy spread over nodes across the Internet. A key goal is to minimize the amount of bandwidth used to maintain that redundancy. Storing a file using an…

Overload-induced cascading failures can cause extreme disruptions in a wide range of networked systems, such as power grids, transportation networks, or financial systems. Empirical studies across domains report that the size of such…

物理与社会 · 物理学 2026-01-06 Agnieszka Janicka , Fiona Sloothaak , Maria Vlasiou , Bert Zwart

The performance of processing search queries depends heavily on the stored index size. Accordingly, considerable research efforts have been devoted to the development of efficient compression techniques for inverted indexes. Roughly, index…

信息检索 · 计算机科学 2011-07-29 M. Feldman , R. Lempel , O. Somekh , K. Vornovitsky

This paper studies the fundamental problem of data persistency for a general family of redundancy schemes in distributed storage systems, called replicated erasure codes. Namely, we analyze two strategies of replicated erasure codes…

分布式、并行与集群计算 · 计算机科学 2021-07-28 Roy Friedman , Rafał Kapelko , Karol Marchwicki

Edge networks are promising to provide better services to users by provisioning computing and storage resources at the edge of networks. However, due to the uncertainty and diversity of user interests, content popularity, distributed…

网络与互联网体系结构 · 计算机科学 2020-03-16 Nitish K. Panigrahy , Jian Li , Faheem Zafari , Don Towsley , Paul Yu

In the study of large scale stochastic networks with resource management, differential equations and mean-field limits are two key techniques. Recent research shows that the expected fraction vector (that is, the tailed probability vector)…

概率论 · 数学 2013-05-27 Quan-Lin Li

This paper introduces the concept of size-aware sharding to improve tail latencies for in-memory key-value stores, and describes its implementation in the Minos key-value store. Tail latencies are crucial in distributed applications with…

数据库 · 计算机科学 2018-02-05 Diego Didona , Willy Zwaenepoel

We consider a distributed storage system which stores several hot (popular) and cold (less popular) data files across multiple nodes or servers. Hot files are stored using repetition codes while cold files are stored using erasure codes.…

性能 · 计算机科学 2021-05-10 Tim Hellemans , Arti Yardi , Tejas Bodas

The rapid growth in the size of large language models has necessitated the partitioning of computational workloads across accelerators such as GPUs, TPUs, and NPUs. However, these parallelization strategies incur substantial data…

机器学习 · 计算机科学 2026-05-11 Rezaul Karim , Austin Wen , Wang Zongzuo , Weiwei Zhang , Yang Liu , Walid Ahmed

Edge computing is a distributed computing paradigm that brings computation and data storage closer to the user's geographical location to improve response times and save bandwidth. It also helps to power a variety of applications requiring…

分布式、并行与集群计算 · 计算机科学 2025-04-30 Ravi Shankar , Aryabartta Sahu