Related papers: Planck 2015 results. XII. Full Focal Plane simulat…
The Planck High Frequency Instrument (HFI) has observed the full sky at six frequencies (100, 143, 217, 353, 545, and 857 GHz) in intensity and at four frequencies in linear polarization (100, 143, 217, and 353 GHz). In order to obtain sky…
The massive computational costs associated with large language model (LLM) pretraining have spurred great interest in reduced-precision floating-point representations to accelerate the process. As a result, the BrainFloat16 (BF16) precision…
Multiplication is a core operation in modern neural network (NN) computations, contributing significantly to energy consumption. The linear-complexity multiplication (L-Mul) algorithm is specifically proposed as an approximate…
Recent advances in deep learning methods such as LLMs and Diffusion models have created a need for improved quantization methods that can meet the computational demands of these modern architectures while maintaining accuracy. Towards this…
The Planck satellite will map the full sky at nine frequencies from 30 to 857 GHz. The CMB intensity and polarization that are its prime targets are contaminated by foreground emission. The goal of this paper is to compare proposed methods…
The Planck Catalogue of Compact Sources (PCCS) is the catalogue of sources detected in the first 15 months of Planck operations, the "nominal" mission. It consists of nine single-frequency catalogues of compact sources, both Galactic and…
We present a data analysis pipeline for CMB polarization experiments, running from multi-frequency maps to the power spectra. We focus mainly on component separation and, for the first time, we work out the covariance matrix accounting for…
To asses stability against 1/f noise, the Low Frequency Instrument (LFI) onboard the Planck mission will acquire data at a rate much higher than the data rate allowed by its telemetry bandwith of 35.5 kbps. The data are processed by an…
The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promising solution with significant theoretical efficiency gains, its widespread adoption has been…
Neural network quantization is widely used to reduce model inference complexity in real-world deployments. However, traditional integer quantization suffers from accuracy degradation when adapting to various dynamic ranges. Recent research…
We train, for the first time, large language models using FP8 precision on datasets up to 2 trillion tokens -- a 20-fold increase over previous limits. Through these extended training runs, we uncover critical instabilities in FP8 training…
The predictive power of Convolutional Neural Networks (CNNs) has been an integral factor for emerging latency-sensitive applications, such as autonomous drones and vehicles. Such systems employ multiple CNNs, each one trained for a…
Modern deep neural network (DNN) models generally require a huge amount of weight and activation values to achieve good inference outcomes. Those data inevitably demand a massive off-chip memory capacity/bandwidth, and the situation gets…
We present simulations of observations with the 143 GHz channel of the Planck High Frequency Instrument (HFI). These simulations are performed over the entire sky, using the true angular resolution of this channel: 8 arcmin FWHM, 3.5 arcmin…
We introduce MF-Box, an extended version of MFEmulator, designed as a fast surrogate for power spectra, trained using N-body simulation suites from various box sizes and particle loads. To demonstrate MF-Box's effectiveness, we design…
As fusion energy devices advance, plasma simulations are crucial for reactor design. Our work extends BIT1 hybrid parallelization by integrating MPI with OpenMP and OpenACC, focusing on asynchronous multi-GPU programming. Results show…
This paper presents the characterization of the in-flight beams, the beam window functions, and the associated uncertainties for the Planck Low Frequency Instrument (LFI). The structure of the paper is similar to that presented in the 2013…
To efficiently support Large Language Models (LLMs), modern GPGPU architectures have introduced new features and programming paradigms, such as warp specialization. These features enable temporal overlap between the producer and consumer,…
FP8 low-precision formats have gained significant adoption in Transformer inference and training. However, existing digital compute-in-memory (DCIM) architectures face challenges in supporting variable FP8 aligned-mantissa bitwidths, as…
We use an iterative generalized least squares map-making algorithm, in conjunction with Monte Carlo techniques, to obtain estimates of the angular power spectrum from cosmic microwave background (CMB) maps. This is achieved by…