English
Related papers

Related papers: Hardware Complexity Aware Design Strategy for a Fu…

200 papers

When serving a single base LLM with several different LoRA adapters simultaneously, the adapters cannot simply be merged with the base model's weights as the adapter swapping would create overhead and requests using different adapters could…

Machine Learning · Computer Science 2026-01-07 Xinyu Wang , Jonas M. Kübler , Kailash Budhathoki , Yida Wang , Matthäus Kleindessner

Transformer encoders contextualize token representations by attending to all other tokens at each layer, leading to quadratic increase in compute effort with the input length. In practice, however, the input text of many NLP tasks can be…

Computation and Language · Computer Science 2023-06-01 Jeremiah Milbauer , Annie Louis , Mohammad Javad Hosseini , Alex Fabrikant , Donald Metzler , Tal Schuster

We calculate the complete double logarithmic contribution to cross sections for semi-inclusive hadron production in the modified minimal-subtraction scheme by applying dimensional regularization to the double logarithm approximation. The…

High Energy Physics - Phenomenology · Physics 2011-07-06 S. Albino , P. Bolzoni , B. A. Kniehl , A. Kotikov

Looped transformers apply a shared block multiple times and have emerged as a parameter-efficient route to scaling compute in language models. However, at fixed FLOPs a looped model has strictly less capacity than a baseline transformer. We…

Computation and Language · Computer Science 2026-05-29 Markus Frey , Behzad Shomali , Joachim Koehler , Mehdi Ali

While Clifford operations are relatively easy to implement in fault-tolerant quantum computers,continuous rotation gates remain a significant bottleneck in typical quantum algorithms. In this work, we ask the question: "What is the most…

Quantum Physics · Physics 2026-05-06 Zhu Sun , Balint Koczor

The transformer architecture has catalyzed revolutionary advances in language modeling. However, recent architectural recipes, such as state-space models, have bridged the performance gap. Motivated by this, we examine the benefits of…

Machine Learning · Computer Science 2024-07-09 Mingchen Li , Xuechen Zhang , Yixiao Huang , Samet Oymak

This article presents two area/latency optimized gate level asynchronous full adder designs which correspond to early output logic. The proposed full adders are constructed using the delay-insensitive dual-rail code and adhere to the…

Hardware Architecture · Computer Science 2016-04-15 P Balasubramanian , S Yamashita

In this paper, we introduce the gauge-fixed QLDPC surgery scheme, an improved logical measurement scheme based on the construction of Cohen et al. (Sci. Adv. 8, eabn1717). Our scheme leverages expansion properties of the Tanner graph to…

Quantum Physics · Physics 2025-10-29 Andrew W. Cross , Zhiyang He , Patrick J. Rall , Theodore J. Yoder

Crosstalk computing, involving engineered interference between nanoscale metal lines, offers a fresh perspective to scaling through co-existence with CMOS. Through capacitive manipulations and innovative circuit style, not only primitive…

Emerging Technologies · Computer Science 2019-04-09 Md Arif Iqbal , Naveen Kumar Macha , Bhavana Tejaswini Repalle , Mostafizur Rahman

Reversible logic allows low power dissipating circuit design and founds its application in cryptography, digital signal processing, quantum and optical information processing. This paper presents a novel quantum cost efficient reversible…

Hardware Architecture · Computer Science 2012-05-04 Md. Saiful Islam , Mohd. Zulfiquar Hafiz , Zerina Begum

Modern key-value stores rely heavily on Log-Structured Merge (LSM) trees for write optimization, but this design introduces significant read amplification. Auxiliary structures like Bloom filters help, but impose memory costs that scale…

Data Structures and Algorithms · Computer Science 2025-08-05 Nicholas Fidalgo , Puyuan Ye

We present a theoretical analysis and empirical evaluations of a novel set of techniques for computational cost reduction of classifiers that are based on learned transform and soft-threshold. By modifying optimization procedures for…

Specialized accelerators have recently garnered attention as a method to reduce the power consumption of neural network inference. A promising category of accelerators utilizes nonvolatile memory arrays to both store weights and perform…

We consider a low-complexity version of the Compute and Forward scheme that involves only scaling, offset (dithering removal) and scalar quantization at the relays. The proposed scheme is suited for the uplink of a distributed antenna…

Information Theory · Computer Science 2011-09-06 Song-Nam Hong , Giuseppe Caire

Analog multiplexing appears to be a promising solution for modern transmitters, where speed is the primary limitation. The objective is the development of a low-cost solution to compare different digital to analog (DAC) schemes. In…

Signal Processing · Electrical Eng. & Systems 2026-02-03 Alfredo Pérez Vega-Leal , Manuel G. Satué

As IoT and edge inference proliferate,there is a growing need to simultaneously optimize area and delay in lookup-table (LUT)-based multipliers that implement large numbers of low-bitwidth operations in parallel. This paper proposes a…

Hardware Architecture · Computer Science 2025-10-27 Misaki Kida , Shimpei Sato

Transformers have achieved remarkable success in sequence modeling and beyond but suffer from quadratic computational and memory complexities with respect to the length of the input sequence. Leveraging techniques include sparse and linear…

Machine Learning · Computer Science 2022-08-02 Tan Nguyen , Richard G. Baraniuk , Robert M. Kirby , Stanley J. Osher , Bao Wang

In the paper, the family of conformal four-point ladder diagrams in arbitrary space-time dimensions is considered. We use the representation obtained via explicit calculation using the operator approach and conformal quantum mechanics to…

High Energy Physics - Theory · Physics 2026-01-22 S. E. Derkachov , A. P. Isaev , L. A. Shumilov

Liquid Argon Time Projection Chamber (LArTPC) technology is commonly utilized in neutrino detector designs. It enables detailed reconstruction of neutrino events with high spatial precision and low energy threshold. Its field response (FR)…

Instrumentation and Detectors · Physics 2023-05-03 S. Martynenko , F. Pietropaolo , B. Viren , X. Qian , H. Chen , S. Gao , W. Gu , J. Jo , S. Kettell , Y. Li , H. Liu , N. Nayak , B. Yu , H. Yu , C. Zhang , U. Kose , F. Resnati , S. Tufanli , F. Boran , F. Dolek

For decoding non-binary low-density parity check (LDPC) codes, logarithm-domain sum-product (Log-SP) algorithms were proposed for reducing quantization effects of SP algorithm in conjunction with FFT. Since FFT is not applicable in the…

Information Theory · Computer Science 2015-05-19 Kenta Kasai , Kohichi Sakaniwa
‹ Prev 1 4 5 6 7 8 10 Next ›