中文
相关论文

相关论文: Faster Energy Efficient Dadda Based Baugh-Wooley M…

200 篇论文

In this work novel results concerning Network-on-Chip-based turbo decoder architectures are presented. Stemming from previous publications, this work concentrates first on improving the throughput by exploiting adaptive-bandwidth reduction…

硬件体系结构 · 计算机科学 2011-05-06 Maurizio Martina , Guido Masera

We present a new method that efficiently solves TO problems and provides a practical pathway to leverage quantum computing to exploit potential quantum advantages. This work targets on large-scale, multi-material TO challenges for…

计算物理 · 物理学 2026-01-16 Zisheng Ye , Wenxiao Pan

Slow working nodes, known as stragglers, can greatly reduce the speed of distributed computation. Coded matrix multiplication is a recently introduced technique that enables straggler-resistant distributed multiplication of large matrices.…

信息论 · 计算机科学 2019-07-23 Shahrzad Kiani , Nuwan Ferdinand , Stark C. Draper

Approximate computing is emerging as an alternative to accurate computing due to its potential for realizing digital circuits and systems with low power dissipation, less critical path delay, and less area occupancy for an acceptable…

硬件体系结构 · 计算机科学 2018-01-19 P Balasubramanian

In this emerging world of connected devices, the need for more computing devices with a focus on delay-sensitive application is critical. In this paper, we propose a priority-queue based Fog computing architecture combined with dynamic…

网络与互联网体系结构 · 计算机科学 2021-07-20 Saksham Bhushan , Maode Ma

The rapidly increasing number of cores available in multicore processors does not necessarily lead directly to a commensurate increase in performance: programs written in conventional languages, such as C, need careful restructuring,…

编程语言 · 计算机科学 2015-01-28 Esraa Alwan , John Fitch , Julian Padget

As IoT and edge inference proliferate,there is a growing need to simultaneously optimize area and delay in lookup-table (LUT)-based multipliers that implement large numbers of low-bitwidth operations in parallel. This paper proposes a…

硬件体系结构 · 计算机科学 2025-10-27 Misaki Kida , Shimpei Sato

We propose an extremely energy-efficient mixed-signal approach for performing vector-by-matrix multiplication in a time domain. In such implementation, multi-bit values of the input and output vector elements are represented with…

硬件体系结构 · 计算机科学 2017-11-30 Mohammad Bavandpour , Mohammad Reza Mahmoodi , Dmitri B. Strukov

We consider the fundamental problem of constructing fast circuits for the carry bit computation in binary addition. Up to a small additive constant, the carry bit computation reduces to computing an \aop, i.e., a formula of type $t_0 \land…

数据结构与算法 · 计算机科学 2019-10-28 Ulrich Brenner , Anna Hermann

We present new algorithms for computing the low $n$ bits or the high $n$ bits of the product of two $n$-bit integers. We show that these problems may be solved in asymptotically 75% of the time required to compute the full $2n$-bit product,…

符号计算 · 计算机科学 2023-08-03 David Harvey

Approximate multipliers are widely being advocated for energy-efficient computing in applications that exhibit an inherent tolerance to inaccuracy. However, the inclusion of accuracy as a key design parameter, besides the performance, area…

新兴技术 · 计算机科学 2018-03-20 Mahmoud Masadeh , Osman Hasan , Sofiene Tahar

Fast binary compressors are the main components of many basic digital calculation units. In this paper, a high-speed (7,2) compressor with a fast carry-generation logic is proposed. The carry-generation logic is based on the sorting…

硬件体系结构 · 计算机科学 2023-09-08 Wenbo Guo

There is a recent trend in artificial intelligence (AI) inference towards lower precision data formats down to 8 bits and less. As multiplication is the most complex operation in typical inference tasks, there is a large demand for…

硬件体系结构 · 计算机科学 2024-05-06 Andreas Böttcher , Martin Kumm

We present a new implementation of the Floyd-Warshall All-Pairs Shortest Paths algorithm on CUDA. Our algorithm runs approximately 5 times faster than the previously best reported algorithm. In order to achieve this speedup, we applied a…

分布式、并行与集群计算 · 计算机科学 2010-02-25 Ben Lund , Justin W Smith

This paper investigates novel techniques to solve prime factorization by quantum annealing (QA). Our contribution is twofold. First, we present a novel and very compact modular encoding of a binary multiplier circuit into the Pegasus…

量子物理 · 物理学 2023-10-27 Jingwen Ding , Giuseppe Spallitta , Roberto Sebastiani

Conventional wireless power transfer systems are linear and time-invariant, which sets fundamental limitations on their performance, including a tradeoff between transfer efficiency and the level of transferred power. In this paper, we…

应用物理 · 物理学 2024-02-26 X. Wang , I. Krois , N. Ha-Van , M. S. Mirmoosa , P. Jayathurathnage , S. Hrabar , S. A. Tretyakov

We propose Booster, a novel accelerator for gradient boosting trees based on the unique characteristics of gradient boosting models. We observe that the dominant steps of gradient boosting training (accounting for 90-98% of training time)…

硬件体系结构 · 计算机科学 2020-11-06 Mingxuan He , T. N. Vijaykumar , Mithuna Thottethodi

This paper focuses on the design and implementing of GPU-accelerated Adaptive Inverse Distance Weighting (AIDW) interpolation algorithm. The AIDW is an improved version of the standard IDW, which can adaptively determine the power parameter…

分布式、并行与集群计算 · 计算机科学 2017-09-21 Gang Mei , Liangliang Xu , Nengxiong Xu

The rapid advancements in machine learning techniques have led to significant achievements in various real-world robotic tasks. These tasks heavily rely on fast and energy-efficient inference of deep neural network (DNN) models when…

机器人学 · 计算机科学 2024-05-30 Zekai Sun , Xiuxian Guan , Junming Wang , Haoze Song , Yuhao Qing , Tianxiang Shen , Dong Huang , Fangming Liu , Heming Cui

In this paper, an optimized efficient VLSI architecture of a pipeline Fast Fourier transform (FFT) processor capable of producing the reverse output order sequence is presented. Paper presents Radix-2 multipath delay architecture for FFT…

硬件体系结构 · 计算机科学 2017-07-07 Tanaji U. Kamble , B. G. Patil , Rakhee S. Bhojakar