English
Related papers

Related papers: The apeNEXT project (Status report)

200 papers

Efficiently serving Large Language Models (LLMs) requires selecting an optimal parallel execution plan, balancing computation, memory, and communication overhead. However, determining the best strategy is challenging due to varying…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-01 Yi-Chien Lin , Woosuk Kwon , Ronald Pineda , Fanny Nina Paravecino

This work presents CascadeCNN, an automated toolflow that pushes the quantisation limits of any given CNN model, to perform high-throughput inference by exploiting the computation time-accuracy trade-off. Without the need for retraining, a…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Alexandros Kouris , Stylianos I. Venieris , Christos-Savvas Bouganis

Customized hardware accelerators have been developed to provide improved performance and efficiency for DNN inference and training. However, the existing hardware accelerators may not always be suitable for handling various DNN models as…

Hardware Architecture · Computer Science 2021-04-07 Xiaofan Zhang , Hanchen Ye , Deming Chen

Recently, special-purpose computers have surpassed general-purpose computers in the speed with which large-scale stellar dynamics simulations can be performed. Speeds up to a Teraflops are now available, for simulations in a variety of…

Astrophysics · Physics 2007-05-23 Piet Hut

The design of a parallel computing system using several thousands or even up to a million processors asks for processing units that are simple and thus small in space, to make as many processing units as possible fit on a single die. The…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-12-21 Oskar Schirmer

We study the current state of the Quantum Software Engineering (QSE) ecosystem, focusing on the achievements, activities, and engagements from academia and industry, with a special focus on successful entrepreneurial endeavors in this…

Software Engineering · Computer Science 2026-01-07 Nazanin Siavash , Armin Moin

The status of the multi-purpose event generator HELAC is briefly presented. The aim of this tool is the full simulation of events within the Standard Model at current and future high energy experiments, in particular the LHC. Some results…

High Energy Physics - Phenomenology · Physics 2017-08-23 C. G. Papadopoulos , M. Worek

The prospects of quantum computing have driven efforts to realize fully functional quantum processing units (QPUs). Recent success in developing proof-of-principle QPUs has prompted the question of how to integrate these emerging processors…

Emerging Technologies · Computer Science 2015-12-10 Keith A. Britt , Travis S. Humble

The objective of our research is to demonstrate the practical usage and orders of magnitude speedup of real-world applications by using alternative technologies to support high performance computing. Currently, the main barrier to the…

Astrophysics · Physics 2007-11-22 Robert J. Brunner , Volodymyr V. Kindratenko , Adam D. Myers

Automated Feature Engineering (AFE) refers to automatically generate and select optimal feature sets for downstream tasks, which has achieved great success in real-world applications. Current AFE methods mainly focus on improving the…

Machine Learning · Computer Science 2022-12-27 Kafeng Wang , Pengyang Wang , Chengzhong xu

Project X is a multi-megawatt proton facility being developed to support a world-leading program in Intensity Frontier physics at Fermilab. The facility will support programs in elementary particle and nuclear physics, with the potential…

Accelerator Physics · Physics 2014-09-23 S. D. Holmes , M. Kaducak , R. Kephart , I. Kourbanis , V. Lebedev , S. Mishra , S. Nagaitsev , N. Solyak , R. Tschirhart

We define a benchmark suite for lattice QCD and report on benchmark results from several computer platforms. The platforms considered are apeNEXT, CRAY T3E, Hitachi SR8000, IBM p690, PC-Clusters, and QCDOC.

High Energy Physics - Lattice · Physics 2009-11-10 M. Hasenbusch , K. Jansen , D. Pleiter , H. St"uben , P. Wegner , T. Wettig , H. Wittig

There are many different models of concurrent processes. The goal of this work is to introduce a common formalized framework for current research in this area and to eliminate shortcomings of existing models of concurrency. Following up the…

Logic in Computer Science · Computer Science 2008-03-24 Mark Burgin a , Marc L. Smith

Edge computing is deemed a promising technique to execute latency-sensitive applications by offloading computation-intensive tasks to edge servers. Extensive research has been conducted in the field of end-device to edge server task…

Networking and Internet Architecture · Computer Science 2024-09-18 Xiang Li , Mustafa Abdallah , Yuan-Yao Lou , Mung Chiang , Kwang Taik Kim , Saurabh Bagchi

Over the past several years, new machine learning accelerators were being announced and released every month for a variety of applications from speech recognition, video object detection, assisted driving, and many data center applications.…

Hardware Architecture · Computer Science 2021-12-13 Albert Reuther , Peter Michaleas , Michael Jones , Vijay Gadepally , Siddharth Samsi , Jeremy Kepner

The past decade has seen a remarkable series of advances in machine learning, and in particular deep learning approaches based on artificial neural networks, to improve our abilities to build more accurate systems across a broad range of…

Machine Learning · Computer Science 2019-11-14 Jeffrey Dean

This work is for designing one-stage lightweight detectors which perform well in terms of mAP and latency. With baseline models each of which targets on GPU and CPU respectively, various operations are applied instead of the main operations…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Deokki Hong

Modernizing the Fermilab accelerator control system is essential to future operations of the laboratory's accelerator complex. The existing control system has evolved over four decades and uses hardware that is no longer available and…

Accelerator Physics · Physics 2024-01-05 D. Finstrom , E. Gottschalk

In the field of High Performance Computing, communications among processes represent a typical bottleneck for massively parallel scientific applications. Object of this research is the development of a network interface card with specific…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-07 Roberto Ammendola
‹ Prev 1 3 4 5 6 7 10 Next ›