English
Related papers

Related papers: End-to-End Throughput Benchmarking of Portable Det…

200 papers

A neural network architecture is presented that exploits the multilevel properties of high-dimensional parameter-dependent partial differential equations, enabling an efficient approximation of parameter-to-solution maps, rivaling…

Machine Learning · Computer Science 2024-08-21 Janina Enrica Schütte , Martin Eigel

High-throughput physics experiments require efficient and increasingly complex real-time processing. This paper presents a modular, software-defined platform combining high-bandwidth PCIe digitizers with consumer GPUs to achieve continuous,…

Instrumentation and Detectors · Physics 2026-05-12 Toma-Stefan Cezar , Marios Maroudas , Dieter Horns

Deep Convolutional Neural Networks (CNNs) are the state-of-the-art in image classification. Since CNN feed forward propagation involves highly regular parallel computation, it benefits from a significant speed-up when running on fine grain…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-12-13 Kamel Abdelouahab , Maxime Pelcat , Jocelyn Sérot , Cédric Bourrasset , François Berry , Jocelyn Serot

Deep networks are now able to achieve human-level performance on a broad spectrum of recognition tasks. Independently, neuromorphic computing has now demonstrated unprecedented energy-efficiency through a new chip architecture based on…

The predictive power of Convolutional Neural Networks (CNNs) has been an integral factor for emerging latency-sensitive applications, such as autonomous drones and vehicles. Such systems employ multiple CNNs, each one trained for a…

Computer Vision and Pattern Recognition · Computer Science 2021-06-09 Stylianos I. Venieris , Christos-Savvas Bouganis

Deep convolutional neural networks (CNNs) obtain outstanding results in tasks that require human-level understanding of data, like image or speech recognition. However, their computational load is significant, motivating the development of…

Neural and Evolutionary Computing · Computer Science 2019-11-28 Paolo Meloni , Alessandro Capotondi , Gianfranco Deriu , Michele Brian , Francesco Conti , Davide Rossi , Luigi Raffo , Luca Benini

Implementing Deep Neural Networks (DNNs) on resource-constrained edge devices is a challenging task that requires tailored hardware accelerator architectures and a clear understanding of their performance characteristics when executing the…

Recent advances in probabilistic modelling have led to a large number of simulation-based inference algorithms which do not require numerical evaluation of likelihoods. However, a public benchmark with appropriate performance metrics for…

Machine Learning · Statistics 2021-04-12 Jan-Matthis Lueckmann , Jan Boelts , David S. Greenberg , Pedro J. Gonçalves , Jakob H. Macke

Modern CNN are typically based on floating point linear algebra based implementations. Recently, reduced precision NN have been gaining popularity as they require significantly less memory and computational resources compared to floating…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Jiang Su , Nicholas J. Fraser , Giulio Gambardella , Michaela Blott , Gianluca Durelli , David B. Thomas , Philip Leong , Peter Y. K. Cheung

Automated design methods for convolutional neural networks (CNNs) have recently been developed in order to increase the design productivity. We propose a neuroevolution method capable of evolving and optimizing CNNs with respect to the…

Neural and Evolutionary Computing · Computer Science 2019-10-16 Filip Badan , Lukas Sekanina

Convolutional Neural Networks (CNNs) serve various applications with diverse performance and resource requirements. Model-aware CNN accelerators best address these diverse requirements. These accelerators usually combine multiple dedicated…

Hardware Architecture · Computer Science 2025-04-08 Fareed Qararyah , Mohammad Ali Maleki , Pedro Trancoso

For LHC Run 3, the ALICE Time Projection Chamber was upgraded to operate in continuous readout mode. Interaction rates of up to 50 kHz in Pb-Pb collisions require real-time processing of more than 3 TB/s of raw detector data. This…

Instrumentation and Detectors · Physics 2026-03-18 J. Alme , T. Alt , C. Andrei , V. Anguelov , H. Appelshäuser , M. Arslandok , R. Averbeck , M. Ball , G. G. Barnaföldi , P. Becht , R. Bellwied , A. Berdnikova , B. Blidaru , L. Boldizsár , L. Bratrud , P. Braun-Munzinger , M. Bregant , C. L. Britton , H. Büsching , H. Caines , P. Chatzidaki , P. Christiansen , T. M. Cormier , L. Döpper , R. Ehlers , L. Fabbietti , F. Flor , J. J. Gaardhøje , M. G. Munhoz , C. Garabatos , P. Gasik , Á. Gera , P. Glässel , N. Grünwald , T. Gündem , T. Gunji , H. Hamagaki , J. W. Harris , P. Hauer , E. Hellbär , H. Helstrup , A. Herghelegiu , H. D. Hernandez Herrera , Y. Hou , C. Hughes , M. Ivanov , J. Jäger , Y. Ji , J. Jung , M. Jung , B. Ketzer , S. Kirsch , M. Kleiner , A. G. Knospe , M. Korwieser , M. Kowalski , L. Lautner , M. Lesch , C. Lippmann , G. Mantzaridis , R. D. Majka , A. Marin , C. Markert , S. Masciocchi , A. Matyja , M. Meres , D. L. Mihaylov , D. Miśkowiec , R. H. Munzer , H. Murakami , K. Münning , A. Nassirpour , C. Nattrass , B. S. Nielsen , W. A. V. Noije , A. C. Oliveira Da Silva , A. Oskarsson , K. Oyama , L. Österman , Y. Pachmayer , G. Paić , M. Petris , M. Petrovici , M. Planinic , J. Rasson , K. F. Read , A. Rehman , R. Renfordt , A. Riedel , K. Røed , D. Röhrich , E. Rubio , A. Rusu , S. Sadhu , B. C. S. Sanches , J. Schambach , A. Schmah , C. Schmidt , A. Schmier , K. Schweda , D. Sekihata , D. Silvermyr , B. Sitar , N. Smirnov , H. K. Soltveit , C. Sonnabend , S. P. Sorensen , J. Stachel , L. Šerkšnytė , G. Tambave , K. Ullaland , B. Ulukutlu , D. Varga , O. Vazquez Rueda , B. Voss , J. Wiechula , B. Windelband , J. Wilkinson , J. Witte , A. Yadav , F. Zanone , S. Zhu

FPGA-based hardware accelerators for convolutional neural networks (CNNs) have obtained great attentions due to their higher energy efficiency than GPUs. However, it is challenging for FPGA-based solutions to achieve a higher throughput…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-06-09 Yixing Li , Zichuan Liu , Kai Xu , Hao Yu , Fengbo Ren

Generative model based image lossless compression algorithms have seen a great success in improving compression ratio. However, the throughput for most of them is less than 1 MB/s even with the most advanced AI accelerated chips, preventing…

Image and Video Processing · Electrical Eng. & Systems 2022-06-14 Ning Kang , Shanzhao Qiu , Shifeng Zhang , Zhenguo Li , Shutao Xia

The inherent diversity of computation types within the deep neural network (DNN) models often requires a variety of specialized units in hardware processors, which limits computational efficiency, increasing both inference latency and power…

Machine Learning · Computer Science 2024-08-21 Ruiqi Sun , Siwei Ye , Jie Zhao , Xin He , Jianzhe Lin , Yiran Li , An Zou

Neural Network designs are quite diverse, from VGG-style to ResNet-style, and from Convolutional Neural Networks to Transformers. Towards the design of efficient accelerators, many works have adopted a dataflow-based, inter-layer pipelined…

Machine Learning · Computer Science 2023-06-23 Zhewen Yu , Christos-Savvas Bouganis

The training process of Deep Neural Network (DNN) is compute-intensive, often taking days to weeks to train a DNN model. Therefore, parallel execution of DNN training on GPUs is a widely adopted approach to speed up the process nowadays.…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-10-29 Chi-Chung Chen , Chia-Lin Yang , Hsiang-Yun Cheng

This paper presents the Container Profiler, a software tool that measures and records the resource usage of any containerized task. Our tool profiles the CPU, memory, disk, and network utilization of containerized tasks collecting over…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-02-08 Varik Hoang , Ling-Hong Hung , David Perez , Huazeng Deng , Raymond Schooley , Niharika Arumilli , Ka Yee Yeung , Wes Lloyd

Graph neural networks (GNNs) have recently empowered various novel computer vision (CV) tasks. In GNN-based CV tasks, a combination of CNN layers and GNN layers or only GNN layers are employed. This paper introduces GCV-Turbo, a…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-04-11 Bingyi Zhang , Rajgopal Kannan , Carl Busart , Viktor Prasanna

Growing deployment of power and energy efficient throughput accelerators (GPU) in data centers demands enhancement of power-performance co-optimization capabilities of GPUs. Realization of exascale computing using accelerators requires…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-11-06 Nilanjan Goswami , Amer Qouneh , Chao Li , Tao Li
‹ Prev 1 8 9 10 Next ›