中文
相关论文

相关论文: A systolic update scheme to overcome memory bandwi…

200 篇论文

While real-time image generation using diffusion models has advanced rapidly on NVIDIA GPUs, systematic optimization research on non-CUDA platforms such as Apple Silicon remains extremely limited. In this study, we conducted comprehensive…

机器学习 · 计算机科学 2026-05-19 Yoichi Ochiai

Recent advancements in neural network-based optical flow estimation often come with prohibitively high computational and memory requirements, presenting challenges in their model adaptation for mobile and low-power use cases. In this paper,…

计算机视觉与模式识别 · 计算机科学 2023-06-12 Risheek Garrepalli , Jisoo Jeong , Rajeswaran C Ravindran , Jamie Menjay Lin , Fatih Porikli

A symplectic pseudospectral time-domain (SPSTD) scheme is developed to solve Schrodinger equation. Instead of spatial finite differences in conventional finite-difference time-domain (FDTD) method, the fast Fourier transform is used to…

计算物理 · 物理学 2018-05-09 Jing Shen , Wei E. I. Sha , Xiaojing Kuang , Jinhua Hu , Zhixiang Huang , Xianliang Wu

Discrete cosine transform (DCT) and other Fourier-related transforms have broad applications in scientific computing. However, off-the-shelf high-performance multi-dimensional DCT (MD DCT) libraries are not readily available in parallel…

分布式、并行与集群计算 · 计算机科学 2021-10-05 Zixuan Jiang , Jiaqi Gu , David Z. Pan

Over the past few years, silicon photonics-based computing has emerged as a promising alternative to CMOS-based computing for Deep Neural Networks (DNN). Unfortunately, the non-linear operations and the high-precision requirements of DNNs…

This paper proposes a versatile high-performance execution model, inspired by systolic arrays, for memory-bound regular kernels running on CUDA-enabled GPUs. We formulate a systolic model that shifts partial sums by CUDA warp primitives for…

分布式、并行与集群计算 · 计算机科学 2019-09-09 Peng Chen , Mohamed Wahib , Shinichiro Takizawa , Ryousei Takano , Satoshi Matsuoka

Modern problems in high-performance computing, ranging from training and inferencing deep learning models in computer vision and language models to simulating complex physical systems with nonlinearly-coupled equations, require exponential…

The research interest in specialized hardware accelerators for deep neural networks (DNN) spikes recently owing to their superior performance and efficiency. However, today's DNN accelerators primarily focus on accelerating specific…

分布式、并行与集群计算 · 计算机科学 2020-06-11 Cong Guo , Yangjie Zhou , Jingwen Leng , Yuhao Zhu , Zidong Du , Quan Chen , Chao Li , Bin Yao , Minyi Guo

Dynamic programming (DP) is a cornerstone of combinatorial optimization, yet its inherently sequential structure has long limited its scalability in scenario-based stochastic programming (SP). This paper introduces a GPU-accelerated…

最优化与控制 · 数学 2025-11-25 Jingyi Zhao , Linxin Yang , Haohua Zhang , Tian Ding

In this paper, we describe the algorithms we implemented in FDPS to make efficient use of accelerator hardware such as GPGPUs. We have developed FDPS to make it possible for many researchers to develop their own high-performance parallel…

天体物理仪器与方法 · 物理学 2020-02-12 Masaki Iwasawa , Daisuke Namekata , Keigo Nitadori , Kentaro Nomura , Long Wang , Miyuki Tsubouchi , Junichiro Makino

In terms of 3D imaging speed and system cost, the single-camera system projecting single-frequency patterns is the ideal option among all proposed Fringe Projection Profilometry (FPP) systems. This system necessitates a robust spatial phase…

图像与视频处理 · 电气工程与系统科学 2023-03-14 Xiaolong Luo , Wanzhong Song , Songlin Bai , Yu Li , Zhihe Zhao

This paper presents the generalized formulations of fundamental schemes for efficient unconditionally stable implicit finite-difference time-domain (FDTD) methods. The fundamental schemes constitute a family of implicit schemes that feature…

数值分析 · 数学 2020-12-01 Eng Leong Tan

Ultra-fast electronic phenomena originating from finite temperature, such as nonlinear optical excitation, can be simulated with high fidelity via real-time time dependent density functional theory (rt-TDDFT) calculations with hybrid…

材料科学 · 物理学 2025-01-07 Rongrong Liu , Zhuoqiang Guo , Qiuchen Sha , Tong Zhao , Haibo Li , Wei Hu , Lijun Liu , Guangming Tan , Weile Jia

This paper reports large-scale direct numerical simulations of homogeneous-isotropic fluid turbulence, achieving sustained performance of 1.08 petaflop/s on gpu hardware using single precision. The simulations use a vortex particle method…

数值分析 · 计算机科学 2012-10-30 R. Yokota , L. A. Barba , T. Narumi , K. Yasuoka

This work presents an optical neuromorphic imaging and processing cytometry system that integrates an excitable VCSEL-based time-delayed (TD) extreme learning machine with an event-based 2D camera. The proposed system is designed for the…

光学 · 物理学 2025-05-20 M. Skontranis , G. Moustakas , A. Bogris , C. Mesaritakis

The rapid growth of deep neural networks (DNNs) has exposed fundamental limitations in electronic accelerators, where data movement dominates energy consumption, commonly referred to as the memory wall. Photonic accelerators offer a…

硬件体系结构 · 计算机科学 2026-05-01 Belal Jahannia , Abdolah Amirany , Hamed Dalir

Solving compressible flows containing both smooth and discontinuous flow structures remains a significant challenge for finite volume methods. Godunov-type finite volume methods are commonly used for numerical simulations of compressible…

流体动力学 · 物理学 2025-02-07 Minsheng Huang , Lidong Cheng , Wenjun Ying , Xi Deng , Feng Xiao

The integration of computing with memory is essential for distributed, massively parallel, and adaptive architectures such as neural networks in artificial intelligence (AI). Accelerating AI can be achieved through photonic computing, but…

Matrix-accelerated stencil computation is a hot research topic, yet its application to three-dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence of matrix units on multicore CPUs, we analyze matrix-based…

分布式、并行与集群计算 · 计算机科学 2025-07-16 Yinuo Wang , Tianqi Mao , Lin Gan , Wubing Wan , Zeyu Song , Jiayu Fu , Lanke He , Wenqiang Wang , Zekun Yin , Wei Xue , Guangwen Yang

Photonic in-memory computing is a high-speed, low-energy alternative to traditional transistor-based digital computing that utilizes high photonic operating frequencies and bandwidths. In this work, we develop a comprehensive system-level…

分布式、并行与集群计算 · 计算机科学 2026-02-03 Jebacyril Arockiaraj , Sasindu Wijeratne , Sugeet Sunder , Md Abdullah-Al Kaiser , Akhilesh Jaiswal , Ajey P. Jacob , Viktor Prasanna