中文
相关论文

相关论文: NaNet:a low-latency NIC enabling GPU-based, real-t…

200 篇论文

The dominant approaches for named entity recognition (NER) mostly adopt complex recurrent neural networks (RNN), e.g., long-short-term-memory (LSTM). However, RNNs are limited by their recurrent nature in terms of computational efficiency.…

计算与语言 · 计算机科学 2019-07-22 Hui Chen , Zijia Lin , Guiguang Ding , Jianguang Lou , Yusen Zhang , Borje Karlsson

In cloud machine learning (ML) inference systems, providing low latency to end-users is of utmost importance. However, maximizing server utilization and system throughput is also crucial for ML service providers as it helps lower the…

分布式、并行与集群计算 · 计算机科学 2022-03-01 Yunseong Kim , Yujeong Choi , Minsoo Rhu

The LHCb experiment at CERN is undergoing an upgrade in preparation for the Run 3 data taking period of the LHC. As part of this upgrade the trigger is moving to a fully software implementation operating at the LHC bunch crossing rate. We…

仪器与探测器 · 物理学 2022-01-06 R. Aaij , M. Adinolfi , S. Aiola , S. Akar , J. Albrecht , M. Alexander , S. Amato , Y. Amhis , F. Archilli , M. Bala , G. Bassi , L. Bian , M. P. Blago , T. Boettcher , A. Boldyrev , S. Borghi , A. Brea Rodriguez , L. Calefice , M. Calvo Gomez , D. H. Cámpora Pérez , A. Cardini , M. Cattaneo , V. Chobanova , G. Ciezarek , X. Cid Vidal , J. L. Cobbledick , J. A. B. Coelho , T. Colombo , A. Contu , B. Couturier , D. C. Craik , R. Currie , P. d'Argent , M. De Cian , D. Derkach , F. Dordei , M. Dorigo , L. Dufour , P. Durante , A. Dziurda , A. Dzyuba , S. Easo , S. Esen , P. Fernandez Declara , S. Filippov , C. Fitzpatrick , M. Frank , P. Gandini , V. V. Gligorov , E. Golobardes , G. Graziani , L. Grillo , P. A. Günther , S. Hansmann-Menzemer , A. M. Hennequin , L. Henry , D. Hill , S. E. Hollitt , J. Hu , W. Hulsbergen , R. J. Hunter , M. Hushchyn , B. K. Jashal , C. R. Jones , S. Klaver , K. Klimaszewski , R. Kopecna , W. Krzemien , M. Kucharczyk , R. Lane , F. Lazzari , R. Le Gac , P. Li , J. H. Lopes , M. Lucio Martinez , A. Lupato , O. Lupton , X. Lyu , F. Machefert , O. Madejczyk , S. Malde , J. F. Marchand , S. Mariani , C. Marin Benito , D. Martinez Santos , F. Martinez Vidal , R. Matev , M. Mazurek , B. Mitreska , D. S. Mitzel , M. J. Morello , H. Mu , P. Muzzetto , P. Naik , M. Needham , N. Neri , N. Neufeld , N. S. Nolte , D. O'Hanlon , A. Oyanguren , M. Pepe Altarelli , S. Petrucci , M. Petruzzo , L. Pica , F. Pisani , A. Piucci , F. Polci , A. Poluektov , E. Polycarpo , C. Prouve , G. Punzi , R. Quagliani , R. I. Rabadan Trejo , M. Ramos Pernas , M. S. Rangel , F. Ratnikov , G. Raven , F. Reiss , V. Renaudin , P. Robbe , A. Ryzhikov , M. Santimaria , M. Saur , M. Schiller , R. Schwemmer , B. Sciascia , A. Solomin , F. Suljik , N. Skidmore , M. D. Sokoloff , P. Spradlin , M. Stahl , S. Stahl , H. Stevens , L. Sun , A. Szabelski , T. Szumlak , M. Szymanski , D. Y. Tou , G. Tuci , A. Usachov , N. Valls Canudas , R. Vazquez Gomez , S. Vecchi , M. Vesterinen , X. Vilasis-Cardona , D. Vom Bruch , Z. Wang , T. Wojton , M. Whitehead , M. Williams , M. Witek , Y. Xie , A. Xu , H. Yin , M. Zdybal , O. Zenaiev , D. Zhang , L. Zhang , X. Zhu

As the particle physics community needs higher and higher precisions in order to test our current model of the subatomic world, larger and larger datasets are necessary. With upgrades scheduled for the detectors of colliding-beam…

数据分析、统计与概率 · 物理学 2025-09-09 Fotis I. Giasemis

As large language models (LLMs) continue to advance, retrieval-augmented generation (RAG) has become the key mechanism for expanding model knowledge and reducing hallucinations. Central to RAG is approximate nearest neighbor search (ANNS),…

硬件体系结构 · 计算机科学 2026-05-22 Cheng Zou , Shuo Yang , Chen Nie , Yu Zou , Yu He , Chao Jiang , Limin Xiao , Weifeng Zhang , Zhezhi He

Systems for serving inference requests on graph neural networks (GNN) must combine low latency with high throughout, but they face irregular computation due to skew in the number of sampled graph nodes and aggregated GNN features. This…

分布式、并行与集群计算 · 计算机科学 2023-05-19 Zeyuan Tan , Xiulong Yuan , Congjie He , Man-Kit Sit , Guo Li , Xiaoze Liu , Baole Ai , Kai Zeng , Peter Pietzuch , Luo Mai

While GPU clusters are the de facto choice for training large deep neural network (DNN) models today, several reasons including ease of workflow, security and cost have led to efforts investigating whether CPUs may be viable for inference…

机器学习 · 计算机科学 2024-03-13 Zhanpeng Zeng , Michael Davies , Pranav Pulijala , Karthikeyan Sankaralingam , Vikas Singh

PEGs are a formal grammar foundation for describing syntax, and are not hard to generate parsers with a plain recursive decent parsing. However, the large amount of C-stack consumption in the recursive parsing is not acceptable especially…

编程语言 · 计算机科学 2015-11-12 Shun Honda , Kimio Kuramitsu

We propose an analytical device model for a graphene nanoribbon field-effect transistor (GNR-FET). The GNR-FET under consideration is based on a heterostructure which consists of an array of nanoribbons clad between the highly conducting…

介观与纳米尺度物理 · 物理学 2009-11-13 M. Ryzhii , A. Satou , V. Ryzhii , T. Otsuji

We present and evaluate the ExaNeSt Prototype, a liquid-cooled rack prototype consisting of 256 Xilinx ZU9EG MPSoCs, 4 TBytes of DRAM, 16 TBytes of SSD, and configurable interconnection 10-Gbps hardware. We developed this testbed in…

In this paper, we provide a fine-grain machine learning-based method, PerfNetV2, which improves the accuracy of our previous work for modeling the neural network performance on a variety of GPU accelerators. Given an application, the…

机器学习 · 计算机科学 2020-12-02 Chuan-Chi Wang , Ying-Chiao Liao , Ming-Chang Kao , Wen-Yew Liang , Shih-Hao Hung

The 5th-generation wireless networks (5G) technologies and mobile edge computing (MEC) provide great promises of enabling new capabilities for the industrial Internet of Things. However, the solutions enabled by the 5G ultra-reliable…

网络与互联网体系结构 · 计算机科学 2022-02-16 Peng Hu , Jinhuan Zhang

This paper introduces NEST (Network-Enforced Session Types), a runtime verification framework that moves application-level protocol monitoring into the network fabric. Unlike prior work that instruments or wraps application code, we…

编程语言 · 计算机科学 2026-04-24 Jens Kanstrup Larsen , Alceste Scalas , Guy Amir , Jules Jacobs , Jana Wagemaker , Nate Foster

We present an overview of the Graphics Processing Unit (GPU) based spatial processing system created for the Canadian Hydrogen Intensity Mapping Experiment (CHIME). The design employs AMD S9300x2 GPUs and readily-available commercial…

天体物理仪器与方法 · 物理学 2020-10-19 Nolan Denman , Andre Renard , Keith Vanderlinde , Philippe Berger , Kiyoshi Masui , Ian Tretyakov , the CHIME Collaboration

In-memory computing is an emerging computing paradigm that could enable deeplearning inference at significantly higher energy efficiency and reduced latency. The essential idea is to map the synaptic weights corresponding to each layer to…

Compute nodes on modern heterogeneous supercomputing systems comprise CPUs, GPUs, and high-speed network interconnects (NICs). Parallelization is identified as a technique for effectively utilizing these systems to execute scalable…

分布式、并行与集群计算 · 计算机科学 2025-04-01 Naveen Namashivayam

Hybrid light fidelity (LiFi) and wireless fidelity (WiFi) networks are a promising paradigm of heterogeneous network (HetNet), attributed to the complementary physical properties of optical spectra and radio frequency. However, the current…

机器学习 · 计算机科学 2025-09-09 Han Ji , Xiping Wu , Zhihong Zeng , Chen Chen

We implemented the pressure-implicit with splitting of operators (PISO) and semi-implicit method for pressure-linked equations (SIMPLE) solvers of the Navier-Stokes equations on Fermi-class graphics processing units (GPUs) using the CUDA…

分布式、并行与集群计算 · 计算机科学 2012-10-01 Tadeusz Tomczak , Katarzyna Zadarnowska , Zbigniew Koza , Maciej Matyka , Łukasz Mirosław

Though CNNs are highly parallel workloads, in the absence of efficient on-chip memory reuse techniques, an accelerator for them quickly becomes memory bound. In this paper, we propose a CNN accelerator design for inference that is able to…

分布式、并行与集群计算 · 计算机科学 2025-08-26 Kingshuk Majumder , Shubham Nema , Uday Bondhugula

Large-scale integration of converter-based renewable energy sources (RESs) into the power system will lead to a higher risk of frequency nadir limit violation and even frequency instability after the large power disturbance. Therefore, it…

系统与控制 · 电气工程与系统科学 2021-10-27 Likai Liu , Zechun Hu , Nikhil Pathak , Haocheng Luo
‹ 上一页 1 8 9 10 下一页 ›