English
Related papers

Related papers: NaNet:a low-latency NIC enabling GPU-based, real-t…

200 papers

The dominant approaches for named entity recognition (NER) mostly adopt complex recurrent neural networks (RNN), e.g., long-short-term-memory (LSTM). However, RNNs are limited by their recurrent nature in terms of computational efficiency.…

Computation and Language · Computer Science 2019-07-22 Hui Chen , Zijia Lin , Guiguang Ding , Jianguang Lou , Yusen Zhang , Borje Karlsson

In cloud machine learning (ML) inference systems, providing low latency to end-users is of utmost importance. However, maximizing server utilization and system throughput is also crucial for ML service providers as it helps lower the…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-01 Yunseong Kim , Yujeong Choi , Minsoo Rhu

The LHCb experiment at CERN is undergoing an upgrade in preparation for the Run 3 data taking period of the LHC. As part of this upgrade the trigger is moving to a fully software implementation operating at the LHC bunch crossing rate. We…

Instrumentation and Detectors · Physics 2022-01-06 R. Aaij , M. Adinolfi , S. Aiola , S. Akar , J. Albrecht , M. Alexander , S. Amato , Y. Amhis , F. Archilli , M. Bala , G. Bassi , L. Bian , M. P. Blago , T. Boettcher , A. Boldyrev , S. Borghi , A. Brea Rodriguez , L. Calefice , M. Calvo Gomez , D. H. Cámpora Pérez , A. Cardini , M. Cattaneo , V. Chobanova , G. Ciezarek , X. Cid Vidal , J. L. Cobbledick , J. A. B. Coelho , T. Colombo , A. Contu , B. Couturier , D. C. Craik , R. Currie , P. d'Argent , M. De Cian , D. Derkach , F. Dordei , M. Dorigo , L. Dufour , P. Durante , A. Dziurda , A. Dzyuba , S. Easo , S. Esen , P. Fernandez Declara , S. Filippov , C. Fitzpatrick , M. Frank , P. Gandini , V. V. Gligorov , E. Golobardes , G. Graziani , L. Grillo , P. A. Günther , S. Hansmann-Menzemer , A. M. Hennequin , L. Henry , D. Hill , S. E. Hollitt , J. Hu , W. Hulsbergen , R. J. Hunter , M. Hushchyn , B. K. Jashal , C. R. Jones , S. Klaver , K. Klimaszewski , R. Kopecna , W. Krzemien , M. Kucharczyk , R. Lane , F. Lazzari , R. Le Gac , P. Li , J. H. Lopes , M. Lucio Martinez , A. Lupato , O. Lupton , X. Lyu , F. Machefert , O. Madejczyk , S. Malde , J. F. Marchand , S. Mariani , C. Marin Benito , D. Martinez Santos , F. Martinez Vidal , R. Matev , M. Mazurek , B. Mitreska , D. S. Mitzel , M. J. Morello , H. Mu , P. Muzzetto , P. Naik , M. Needham , N. Neri , N. Neufeld , N. S. Nolte , D. O'Hanlon , A. Oyanguren , M. Pepe Altarelli , S. Petrucci , M. Petruzzo , L. Pica , F. Pisani , A. Piucci , F. Polci , A. Poluektov , E. Polycarpo , C. Prouve , G. Punzi , R. Quagliani , R. I. Rabadan Trejo , M. Ramos Pernas , M. S. Rangel , F. Ratnikov , G. Raven , F. Reiss , V. Renaudin , P. Robbe , A. Ryzhikov , M. Santimaria , M. Saur , M. Schiller , R. Schwemmer , B. Sciascia , A. Solomin , F. Suljik , N. Skidmore , M. D. Sokoloff , P. Spradlin , M. Stahl , S. Stahl , H. Stevens , L. Sun , A. Szabelski , T. Szumlak , M. Szymanski , D. Y. Tou , G. Tuci , A. Usachov , N. Valls Canudas , R. Vazquez Gomez , S. Vecchi , M. Vesterinen , X. Vilasis-Cardona , D. Vom Bruch , Z. Wang , T. Wojton , M. Whitehead , M. Williams , M. Witek , Y. Xie , A. Xu , H. Yin , M. Zdybal , O. Zenaiev , D. Zhang , L. Zhang , X. Zhu

As the particle physics community needs higher and higher precisions in order to test our current model of the subatomic world, larger and larger datasets are necessary. With upgrades scheduled for the detectors of colliding-beam…

Data Analysis, Statistics and Probability · Physics 2025-09-09 Fotis I. Giasemis

As large language models (LLMs) continue to advance, retrieval-augmented generation (RAG) has become the key mechanism for expanding model knowledge and reducing hallucinations. Central to RAG is approximate nearest neighbor search (ANNS),…

Hardware Architecture · Computer Science 2026-05-22 Cheng Zou , Shuo Yang , Chen Nie , Yu Zou , Yu He , Chao Jiang , Limin Xiao , Weifeng Zhang , Zhezhi He

Systems for serving inference requests on graph neural networks (GNN) must combine low latency with high throughout, but they face irregular computation due to skew in the number of sampled graph nodes and aggregated GNN features. This…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-05-19 Zeyuan Tan , Xiulong Yuan , Congjie He , Man-Kit Sit , Guo Li , Xiaoze Liu , Baole Ai , Kai Zeng , Peter Pietzuch , Luo Mai

While GPU clusters are the de facto choice for training large deep neural network (DNN) models today, several reasons including ease of workflow, security and cost have led to efforts investigating whether CPUs may be viable for inference…

Machine Learning · Computer Science 2024-03-13 Zhanpeng Zeng , Michael Davies , Pranav Pulijala , Karthikeyan Sankaralingam , Vikas Singh

PEGs are a formal grammar foundation for describing syntax, and are not hard to generate parsers with a plain recursive decent parsing. However, the large amount of C-stack consumption in the recursive parsing is not acceptable especially…

Programming Languages · Computer Science 2015-11-12 Shun Honda , Kimio Kuramitsu

We propose an analytical device model for a graphene nanoribbon field-effect transistor (GNR-FET). The GNR-FET under consideration is based on a heterostructure which consists of an array of nanoribbons clad between the highly conducting…

Mesoscale and Nanoscale Physics · Physics 2009-11-13 M. Ryzhii , A. Satou , V. Ryzhii , T. Otsuji

We present and evaluate the ExaNeSt Prototype, a liquid-cooled rack prototype consisting of 256 Xilinx ZU9EG MPSoCs, 4 TBytes of DRAM, 16 TBytes of SSD, and configurable interconnection 10-Gbps hardware. We developed this testbed in…

In this paper, we provide a fine-grain machine learning-based method, PerfNetV2, which improves the accuracy of our previous work for modeling the neural network performance on a variety of GPU accelerators. Given an application, the…

Machine Learning · Computer Science 2020-12-02 Chuan-Chi Wang , Ying-Chiao Liao , Ming-Chang Kao , Wen-Yew Liang , Shih-Hao Hung

The 5th-generation wireless networks (5G) technologies and mobile edge computing (MEC) provide great promises of enabling new capabilities for the industrial Internet of Things. However, the solutions enabled by the 5G ultra-reliable…

Networking and Internet Architecture · Computer Science 2022-02-16 Peng Hu , Jinhuan Zhang

This paper introduces NEST (Network-Enforced Session Types), a runtime verification framework that moves application-level protocol monitoring into the network fabric. Unlike prior work that instruments or wraps application code, we…

Programming Languages · Computer Science 2026-04-24 Jens Kanstrup Larsen , Alceste Scalas , Guy Amir , Jules Jacobs , Jana Wagemaker , Nate Foster

We present an overview of the Graphics Processing Unit (GPU) based spatial processing system created for the Canadian Hydrogen Intensity Mapping Experiment (CHIME). The design employs AMD S9300x2 GPUs and readily-available commercial…

Instrumentation and Methods for Astrophysics · Physics 2020-10-19 Nolan Denman , Andre Renard , Keith Vanderlinde , Philippe Berger , Kiyoshi Masui , Ian Tretyakov , the CHIME Collaboration

In-memory computing is an emerging computing paradigm that could enable deeplearning inference at significantly higher energy efficiency and reduced latency. The essential idea is to map the synaptic weights corresponding to each layer to…

Machine Learning · Computer Science 2019-06-11 Martino Dazzi , Abu Sebastian , Pier Andrea Francese , Thomas Parnell , Luca Benini , Evangelos Eleftheriou

Compute nodes on modern heterogeneous supercomputing systems comprise CPUs, GPUs, and high-speed network interconnects (NICs). Parallelization is identified as a technique for effectively utilizing these systems to execute scalable…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-01 Naveen Namashivayam

Hybrid light fidelity (LiFi) and wireless fidelity (WiFi) networks are a promising paradigm of heterogeneous network (HetNet), attributed to the complementary physical properties of optical spectra and radio frequency. However, the current…

Machine Learning · Computer Science 2025-09-09 Han Ji , Xiping Wu , Zhihong Zeng , Chen Chen

We implemented the pressure-implicit with splitting of operators (PISO) and semi-implicit method for pressure-linked equations (SIMPLE) solvers of the Navier-Stokes equations on Fermi-class graphics processing units (GPUs) using the CUDA…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-10-01 Tadeusz Tomczak , Katarzyna Zadarnowska , Zbigniew Koza , Maciej Matyka , Łukasz Mirosław

Though CNNs are highly parallel workloads, in the absence of efficient on-chip memory reuse techniques, an accelerator for them quickly becomes memory bound. In this paper, we propose a CNN accelerator design for inference that is able to…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-08-26 Kingshuk Majumder , Shubham Nema , Uday Bondhugula

Large-scale integration of converter-based renewable energy sources (RESs) into the power system will lead to a higher risk of frequency nadir limit violation and even frequency instability after the large power disturbance. Therefore, it…

Systems and Control · Electrical Eng. & Systems 2021-10-27 Likai Liu , Zechun Hu , Nikhil Pathak , Haocheng Luo
‹ Prev 1 8 9 10 Next ›