Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm
Abstract
When thousands of processors are involved in performing event filtering on a trigger farm, there is likely to be a large number of failures within the software and hardware systems. BTeV, a proton/antiproton collider experiment at Fermi National Accelerator Laboratory, has designed a trigger, which includes several thousand processors. If fault conditions are not given proper treatment, it is conceivable that this trigger system will experience failures at a high enough rate to have a negative impact on its effectiveness. The RTES (Real Time Embedded Systems) collaboration is a group of physicists, engineers, and computer scientists working to address the problem of reliability in large-scale clusters with real-time constraints such as this. Resulting infrastructure must be highly scalable, verifiable, extensible by users, and dynamically changeable.
Keywords
Cite
@article{arxiv.cs/0306074,
title = {Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm},
author = {Jim Kowalkowski},
journal= {arXiv preprint arXiv:cs/0306074},
year = {2008}
}
Comments
Paper for the 2003 Computing in High Energy and Nuclear Physics (CHEP03), La Jolla, Ca, USA, March 2003. PSN THGT001