English
Related papers

Related papers: Range, Not Precision: Block-Floating-Point Half-Pr…

200 papers

We present the first kernel-fused SAR Range Doppler pipeline on any GPU platform. By fusing FFT, matched-filter multiply, and IFFT into a single Metal compute dispatch -- keeping all intermediate data in 32\,KiB on-chip memory -- we process…

Performance · Computer Science 2026-04-07 Mohamed Amine Bergach

We investigate the use of half-precision floating-point numbers (FP16) in mixed-precision linear solvers for lattice QCD simulations. Since the emergence of GPUs for general-purpose, mixed-precision algorithms that combine single-precision…

High Energy Physics - Lattice · Physics 2026-02-17 Issaku Kanamori , Hideo Matsufuru , Tatsumi Aoyama , Kazuyuki Kanaya , Yusuke Namekawa , Hidekatsu Nemura , Keigo Nitadori

Fourier ptychography (FP) is a promising computational imaging technique that overcomes the physical space-bandwidth product (SBP) limit of a conventional microscope by applying angular diversity illuminations. However, to date, the…

Optics · Physics 2018-10-17 An Pan , Yan Zhang , Kai Wen , Maosen Li , Meiling Zhou , Junwei Min , Ming Lei , Baoli Yao

This research work focuses on the design of a high-resolution fast Fourier transform (FFT) /inverse fast Fourier transform (IFFT) processors for constraints analysis purpose. Amongst the major setbacks associated with such high resolution,…

Signal Processing · Electrical Eng. & Systems 2018-06-13 Rozita Teymourzadeh , Mometo Jim Abigo , Mok Vee Hoong

We present an optimized Fast Fourier Transform (FFT) implementation for Apple Silicon GPUs, achieving 138.45~GFLOPS for $N\!=\!4096$ complex single-precision transforms -- a 29\% improvement over Apple's highly optimized vDSP/Accelerate…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-03-31 Mohamed Amine Bergach

Single-precision floating point (FP32) data format, defined by the IEEE 754 standard, is widely employed in scientific computing, signal processing, and deep learning training, where precision is critical. However, FP32 multiplication is…

Hardware Architecture · Computer Science 2025-10-09 Bindu G Gowda , Yogesh Goyal , Yash Gupta , Madhav Rao

The precise analysis and accurate measurement of harmonic provides a reliable scientific industrial application. However, the high-performance DSP processor is the important method of electrical harmonic analysis. Hence, in this research…

Signal Processing · Electrical Eng. & Systems 2018-06-13 Rozita Teymourzadeh , Memtode Jim , Mok Vee hong

We introduce a dual-wavelength Fourier ptychographic topography (FPT) method that extends the lambda/2 height-range limit of single-wavelength FPT. By reconstructing complex fields at two illumination wavelengths and exploiting their phase…

Tensor Core is a mixed-precision matrix-matrix multiplication unit on NVIDIA GPUs with a theoretical peak performance of more than 300 TFlop/s on Ampere architectures. Tensor Cores were developed in response to the high demand of dense…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-10-19 Hiroyuki Ootomo , Rio Yokota

In this paper, we combine the design of band-pass frequency selective surfaces (FSSs) with polarization converters to realize a broadband frequency-selective polarization converter (FSPC) with lowbackward scattering, which consists of the…

Applied Physics · Physics 2019-11-25 Lingling Wang , Shaobin Liu , Xiangkun Kong , Haifeng Zhang , Yongdiao Wen , Qiming Yu

Contemporary field-programmable gate arrays (FPGAs) are predestined for the application of finite impulse response (FIR) filters. Their embedded digital signal processing (DSP) blocks for multiply-accumulate operations enable efficient…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-13 Philipp Födisch , Artsiom Bryksa , Bert Lange , Wolfgang Enghardt , Peter Kaever

Synthetic aperture radar (SAR) is a well established technology in the field of Earth remote sensing. Over the years, the resolution of SAR images has been steadily improving and the pixel count increasing as a result of advances in the…

Quantum Physics · Physics 2025-11-03 Alessandro Giovagnoli , Sigurd Huber , Gerhard Krieger

The fast Fourier transform, FFT, is a useful and prevalent algorithm in signal processing. It characterizes the spectral components of a signal, or is used in combination with other operations to perform more complex computations such as…

Signal Processing · Electrical Eng. & Systems 2017-11-08 Hani Nejadriahi , David HillerKuss , Jonathan K. George , Volker J. Sorger

We describe the performance of our latest generations of sensitive wide-band high-resolution digital Fast Fourier Transform Spectrometer (FFTS). Their design, optimized for a wide range of radio astronomical applications, is presented.…

Instrumentation and Methods for Astrophysics · Physics 2015-06-04 Bernd Klein , Stefan Hochgürtel , Ingo Krämer , Andreas Bell , Klaus Meyer , Rolf Güsten

This paper reports an 11.7 GHz compact 50 ohm ladder filter based on single layer Scandium Aluminum Nitride (ScAlN) film bulk acoustic resonators (FBARs) with platinum (Pt) electrodes, and uses it as a quantitative case study of the limits…

The stretch processing architecture is commonly used for frequency modulated continuous wave (FMCW) radar due to its inexpensive hardware, low sampling rate, and simple architecture. However, the stretch processing architecture is not able…

Signal Processing · Electrical Eng. & Systems 2021-11-02 Moein Movafagh , Avik Santra , Daniel Oloumi

Real-time frequency measurement for non-repetitive and statistically rare signals are challenging problems in the electronic measurement area, which places high demands on the bandwidth, sampling rate, data processing and transmission…

Signal Processing · Electrical Eng. & Systems 2023-08-21 Ruiyuan Ming , Peng Ye , Kuojun Yang , Zhixiang Pan , ChenYang Li , Chuang Huang

In this paper, we propose a mixed-precision convolution unit architecture which supports different integer and floating point (FP) precisions. The proposed architecture is based on low-bit inner product units and realizes higher precision…

Hardware Architecture · Computer Science 2021-01-29 Hamzah Abdel-Aziz , Ali Shafiee , Jong Hoon Shin , Ardavan Pedram , Joseph H. Hassoun

Millimeter wave communications require multibeam beamforming in order to utilize wireless channels that suffer from obstructions, path loss, and multi-path effects. Digital multibeam beamforming has maximum degrees of freedom compared to…

Training Deep Neural Networks (DNNs) can be computationally demanding, particularly when dealing with large models. Recent work has aimed to mitigate this computational challenge by introducing 8-bit floating-point (FP8) formats for…

Hardware Architecture · Computer Science 2024-09-27 Sami Ben Ali , Silviu-Ioan Filip , Olivier Sentieys
‹ Prev 1 2 3 10 Next ›