English

Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm

Distributed, Parallel, and Cluster Computing 2008-11-26 v1

Abstract

When thousands of processors are involved in performing event filtering on a trigger farm, there is likely to be a large number of failures within the software and hardware systems. BTeV, a proton/antiproton collider experiment at Fermi National Accelerator Laboratory, has designed a trigger, which includes several thousand processors. If fault conditions are not given proper treatment, it is conceivable that this trigger system will experience failures at a high enough rate to have a negative impact on its effectiveness. The RTES (Real Time Embedded Systems) collaboration is a group of physicists, engineers, and computer scientists working to address the problem of reliability in large-scale clusters with real-time constraints such as this. Resulting infrastructure must be highly scalable, verifiable, extensible by users, and dynamically changeable.

Keywords

Cite

@article{arxiv.cs/0306074,
  title  = {Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm},
  author = {Jim Kowalkowski},
  journal= {arXiv preprint arXiv:cs/0306074},
  year   = {2008}
}

Comments

Paper for the 2003 Computing in High Energy and Nuclear Physics (CHEP03), La Jolla, Ca, USA, March 2003. PSN THGT001

R2 v1 2026-07-22T12:21:01.898Z