English
Related papers

Related papers: Fine Grain 3D Integration for Microarchitecture De…

200 papers

We describe a modified SIMD architecture suitable for single-chip integration of a large number of processing elements, such as 1,000 or more. Important differences from traditional SIMD designs are: a) The size of the memory per processing…

Astrophysics · Physics 2007-05-23 Junichiro Makino

Emerging interconnects, such as CXL and NVLink, have been integrated into the intra-host topology to scale more accelerators and facilitate efficient communication between them, such as GPUs. To keep pace with the accelerator's growing…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-19 Xu Zhang , Ke Liu , Yisong Chang , Ke Zhang , Mingyu Chen

Design space exploration is commonly performed in embedded system, where the architecture is a complicated piece of engineering. With the current trend of many-core systems, design space exploration in general-purpose computers can no…

Hardware Architecture · Computer Science 2013-09-24 Irfan Uddin

FPGAs are well established in the signal processing domain, where their fine-grained programmable nature allows the inherent parallelism in these applications to be exploited for enhanced performance. As architectures have evolved, FPGA…

Hardware Architecture · Computer Science 2017-10-18 Abdullah Al-Dujaili , Suhaib A. Fahmy

Recent nano-technological advances enable the Monolithic 3D (M3D) integration of multiple memory and logic layers in a single chip, allowing for fine-grained connections between layers and significantly alleviating main memory bottlenecks.…

Two-dimensional (2D) materials have attracted much recent attention because they exhibit various distinct intrinsic properties/functionalities, which are, however, usually not interchangeable. Interestingly, here we propose a generic…

Materials Science · Physics 2019-12-16 Lei Gao , Jia-Tao Sun , Gurjyot Sethi , Yu-Yang Zhang , Shixuan Du , Feng Liu

From AlexNet to Inception, autoencoders to diffusion models, the development of novel and powerful deep learning models and learning algorithms has proceeded at breakneck speeds. In part, we believe that rapid iteration of model…

Computational Engineering, Finance, and Science · Computer Science 2022-11-17 Shehtab Zaman , Ethan Ferguson , Cecile Pereira , Denis Akhiyarov , Mauricio Araya-Polo , Kenneth Chiu

Computing is bottlenecked by data. Large amounts of application data overwhelm storage capability, communication capability, and computation capability of the modern machines we design today. We argue that an intelligent architecture should…

Hardware Architecture · Computer Science 2020-12-24 Onur Mutlu

This work introduces Open3DBench, an open-source 3D-IC backend implementation benchmark built upon the OpenROAD-flow-scripts framework, enabling comprehensive evaluation of power, performance, area, and thermal metrics. Our proposed flow…

Hardware Architecture · Computer Science 2026-04-07 Yunqi Shi , Chengrui Gao , Wanqi Ren , Peng Xie , Siyuan Xu , Ke Xue , Mingxuan Yuan , Chao Qian , Zhi-Hua Zhou

Three-dimensional (3D) photonic integration offers a pathway to overcome the fundamental scaling limitations of planar platforms by enabling enhanced routing flexibility for compact, low-loss, and highly interconnected photonic circuits. In…

Optics · Physics 2026-05-26 Kanhaya Sharma , Adrià Grabulosa , Erik Jung , Daniel Brunner

We present a single-node, multi-GPU programmable graph processing library that allows programmers to easily extend single-GPU graph algorithms to achieve scalable performance on large graphs with billions of edges. Directly using the…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-03-02 Yuechao Pan , Yangzihao Wang , Yuduo Wu , Carl Yang , John D. Owens

Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural…

Machine Learning · Computer Science 2018-02-20 Yanzhi Wang , Caiwen Ding , Zhe Li , Geng Yuan , Siyu Liao , Xiaolong Ma , Bo Yuan , Xuehai Qian , Jian Tang , Qinru Qiu , Xue Lin

The performance bottleneck of deep-learning-based recommender systems resides in their backbone Deep Neural Networks. By integrating Processing-In-Memory~(PIM) architectures, researchers can reduce data movement and enhance energy…

Hardware Architecture · Computer Science 2025-05-19 Feng Cheng , Tunhou Zhang , Junyao Zhang , Jonathan Hao-Cheng Ku , Yitu Wang , Xiaoxuan Yang , Hai , Li , Yiran Chen

In embedded vision systems, parallel computation of the integral image presents several design challenges in terms of hardware resources, speed and power consumption. Although recursive equations significantly reduce the number of…

Computer Vision and Pattern Recognition · Computer Science 2015-10-20 Shoaib Ehsan , Adrian F. Clark , Wah M. Cheung , Arjunsingh M. Bais , Bayar I. Menzat , Nadia Kanwal , Klaus D. McDonald-Maier

To reach large-scale quantum computing, three-dimensional integration of scalable qubit arrays and their control electronics in multi-chip assemblies is promising. Within these assemblies, the use of superconducting interconnections, as…

The massive amounts of data generated by camera sensors motivate data processing inside pixel arrays, i.e., at the extreme-edge. Several critical developments have fueled recent interest in the processing-in-pixel-in-memory paradigm for a…

Image and Video Processing · Electrical Eng. & Systems 2024-02-26 Md Abdullah-Al Kaiser , Gourav Datta , Sreetama Sarkar , Souvik Kundu , Zihan Yin , Manas Garg , Ajey P. Jacob , Peter A. Beerel , Akhilesh R. Jaiswal

Photonics has been one of the primary beneficiaries of advanced silicon manufacturing. By leveraging on mature complementary metal-oxide-semiconductor (CMOS) process nodes, unprecedented device uniformities and scalability have been…

Optics · Physics 2023-12-13 Luigi Ranno , Jia Xu Brian Sia , Khoi Phuong Dao , Juejun Hu

Inverse design coupled with adjoint optimization is a powerful method to design on-chip nanophotonic devices with multi-wavelength and multi-mode optical functionalities. Although only two simulations are required in each iteration of this…

Applied Physics · Physics 2023-07-12 Ahmet Onur Dasdemir , Victor Minden , Emir Salih Magden

In the next decade, the demands for computing in large scientific experiments are expected to grow tremendously. During the same time period, CPU performance increases will be limited. At the CERN Large Hadron Collider (LHC), these two…

Printed Electronics (PE) exhibits on-demand, extremely low-cost hardware due to its additive manufacturing process, enabling machine learning (ML) applications for domains that feature ultra-low cost, conformity, and non-toxicity…

Machine Learning · Computer Science 2023-03-07 Giorgos Armeniakos , Georgios Zervakis , Dimitrios Soudris , Mehdi B. Tahoori , Jörg Henkel
‹ Prev 1 8 9 10 Next ›