English
Related papers

Related papers: Cascade Pipeline for Leading-Order Matrix Element …

200 papers

Transformer-based models are becoming more and more intelligent and are revolutionizing a wide range of human tasks. To support their deployment, AI labs offer inference services that consume hundreds of GWh of energy annually and charge…

Systems and Control · Electrical Eng. & Systems 2025-08-29 Ching-Yi Lin , Sahil Shah

Methods of Machine and Deep Learning are gradually being integrated into industrial operations, albeit at different speeds for different types of industries. The aerospace and aeronautical industries have recently developed a roadmap for…

Automated data preparation pipeline construction is critical for machine learning success, yet existing methods suffer from two fundamental limitations: they treat pipeline construction as black-box optimization without quantifying…

Databases · Computer Science 2025-11-03 Jing Chang , Chang Liu , Jinbin Huang , Shuyuan Zheng , Rui Mao , Jianbin Qin

Quantum computing promises to revolutionize various fields, yet the execution of quantum programs necessitates an effective compilation process. This involves strategically mapping quantum circuits onto the physical qubits of a quantum…

Quantum Physics · Physics 2024-12-19 Tian Li , Xiao-Yue Xu , Chen Ding , Tian-Ci Tian , Wei-You Liao , Shuo Zhang , He-Liang Huang

As deep learning techniques advance more than ever, hyper-parameter optimization is the new major workload in deep learning clusters. Although hyper-parameter optimization is crucial in training deep learning models for high model…

Machine Learning · Computer Science 2019-11-26 Ahnjae Shin , Dong-Jin Shin , Sungwoo Cho , Do Yoon Kim , Eunji Jeong , Gyeong-In Yu , Byung-Gon Chun

We design and implement parallel prefix sum (scan) algorithms using Ascend AI accelerators. Ascend accelerators feature specialized computing units: the cube units for efficient matrix multiplication and the vector units for optimized…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-05 Bartłomiej Wróblewski , Gioele Gottardo , Anastasios Zouzias

Automated Machine Learning (AutoML) is a promising direction for democratizing AI by automatically deploying Machine Learning systems with minimal human expertise. The core technical challenge behind AutoML is optimizing the pipelines of…

Machine Learning · Computer Science 2023-05-26 Sebastian Pineda Arango , Josif Grabocka

The application of Transformer-based large models has achieved numerous success in recent years. However, the exponential growth in the parameters of large models introduces formidable memory challenge for edge deployment. Prior works to…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-09-11 Xueyuan Han , Zinuo Cai , Yichu Zhang , Chongxin Fan , Junhan Liu , Ruhui Ma , Rajkumar Buyya

To address increasing compute demand from recent multi-model workloads with heavy models like large language models, we propose to deploy heterogeneous chiplet-based multi-chip module (MCM)-based accelerators. We develop an advanced…

Hardware Architecture · Computer Science 2023-12-18 Mohanad Odema , Hyoukjun Kwon , Mohammad Abdullah Al Faruque

In this paper, an optimized efficient VLSI architecture of a pipeline Fast Fourier transform (FFT) processor capable of producing the reverse output order sequence is presented. Paper presents Radix-2 multipath delay architecture for FFT…

Hardware Architecture · Computer Science 2017-07-07 Tanaji U. Kamble , B. G. Patil , Rakhee S. Bhojakar

A common method to define a parallel solution for a computational problem consists in finding a way to use the Divide and Conquer paradigm in order to have processors acting on its own data and scheduled in a parallel fashion. MapReduce is…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-01-13 Edelmira Pasarella , Maria-Esther Vidal , Cristina Zoltan

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB/s of raw detector data. This…

Instrumentation and Detectors · Physics 2026-03-18 J. Alme , T. Alt , C. Andrei , V. Anguelov , H. Appelshäuser , M. Arslandok , R. Averbeck , M. Ball , G. G. Barnaföldi , P. Becht , R. Bellwied , A. Berdnikova , B. Blidaru , L. Boldizsár , L. Bratrud , P. Braun-Munzinger , M. Bregant , C. L. Britton , H. Büsching , H. Caines , P. Chatzidaki , P. Christiansen , T. M. Cormier , L. Döpper , R. Ehlers , L. Fabbietti , F. Flor , J. J. Gaardhøje , M. G. Munhoz , C. Garabatos , P. Gasik , Á. Gera , P. Glässel , N. Grünwald , T. Gündem , T. Gunji , H. Hamagaki , J. W. Harris , P. Hauer , E. Hellbär , H. Helstrup , A. Herghelegiu , H. D. Hernandez Herrera , Y. Hou , C. Hughes , M. Ivanov , J. Jäger , Y. Ji , J. Jung , M. Jung , B. Ketzer , S. Kirsch , M. Kleiner , A. G. Knospe , M. Korwieser , M. Kowalski , L. Lautner , M. Lesch , C. Lippmann , G. Mantzaridis , R. D. Majka , A. Marin , C. Markert , S. Masciocchi , A. Matyja , M. Meres , D. L. Mihaylov , D. Miśkowiec , R. H. Munzer , H. Murakami , K. Münning , A. Nassirpour , C. Nattrass , B. S. Nielsen , W. A. V. Noije , A. C. Oliveira Da Silva , A. Oskarsson , K. Oyama , L. Österman , Y. Pachmayer , G. Paić , M. Petris , M. Petrovici , M. Planinic , J. Rasson , K. F. Read , A. Rehman , R. Renfordt , A. Riedel , K. Røed , D. Röhrich , E. Rubio , A. Rusu , S. Sadhu , B. C. S. Sanches , J. Schambach , A. Schmah , C. Schmidt , A. Schmier , K. Schweda , D. Sekihata , D. Silvermyr , B. Sitar , N. Smirnov , H. K. Soltveit , C. Sonnabend , S. P. Sorensen , J. Stachel , L. Šerkšnytė , G. Tambave , K. Ullaland , B. Ulukutlu , D. Varga , O. Vazquez Rueda , B. Voss , J. Wiechula , B. Windelband , J. Wilkinson , J. Witte , A. Yadav , F. Zanone , S. Zhu

Data pre-processing pipelines are the bread and butter of any successful AI project. We introduce a novel programming model for pipelines in a data lakehouse, allowing users to interact declaratively with assets in object storage. Motivated…

Databases · Computer Science 2024-11-14 Jacopo Tagliabue , Ryan Curtin , Ciro Greco

Modern deployment often requires trading accuracy for efficiency under tight CPU and memory constraints, yet common compression proxies such as parameter count or FLOPs do not reliably predict wall-clock inference time. In particular,…

Machine Learning · Computer Science 2026-04-10 Longsheng Zhou , Yu Shen

Computation of a signal's estimated covariance matrix is an important building block in signal processing, e.g., for spectral estimation. Each matrix element is a sum of products of elements in the input matrix taken over a sliding window.…

Data Structures and Algorithms · Computer Science 2013-03-12 Oded Green , Lior David , Ami Galperin , Yitzhak Birk

A new parallel algorithm utilizing partitioned global address space (PGAS) programming model to achieve high scalability is reported for particle tracking in direct numerical simulations of turbulent flow. The work is motivated by the…

Computational Physics · Physics 2020-05-28 Dhawal Buaria , P. K. Yeung

AI systems comprise a range of interactions across the technical and organisational components of a range of actors. These components work together to provide the systems' functionality. This socio-technical assemblage is increasingly…

Computers and Society · Computer Science 2026-03-03 Anna Neumann , Jatinder Singh

The convex hull is a fundamental geometrical structure for many applications where groups of points must be enclosed or represented by a convex polygon. Although efficient sequential convex hull algorithms exist, and are constantly being…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-27 Alan Keith , Héctor Ferrada , Cristóbal A. Navarro

The development of an LHC physics analysis involves numerous investigations that require the repeated processing of terabytes of data. Thus, a rapid completion of each of these analysis cycles is central to mastering the science project. We…

Data Analysis, Statistics and Probability · Physics 2022-07-19 Niclas Eich , Martin Erdmann , Peter Fackeldey , Benjamin Fischer , Dennis Noll , Yannik Rath

Both IP lookup and packet classification in IP routers can be implemented by some form of tree traversal. SRAM-based Pipelining can improve the throughput dramatically. However, previous pipelining schemes result in unbalanced memory…

Networking and Internet Architecture · Computer Science 2011-07-28 Weirong Jiang , Hoang Le , Viktor K. Prasanna