English
Related papers

Related papers: From Merging Frameworks to Merging Stars: Experien…

200 papers

We review the recent optimizations of gravitational $N$-body kernels for running them on graphics processing units (GPUs), on single hosts and massive parallel platforms. For each of the two main $N$-body techniques, direct summation and…

Instrumentation and Methods for Astrophysics · Physics 2014-09-22 Simon Portegies Zwart , Jeroen Bédorf

We present a GPU-accelerated backend for QOCO, a C-based solver for quadratic objective second-order cone programs (SOCPs) based on a primal-dual interior point method. Our backend uses NVIDIA's cuDSS library to perform a direct sparse LDL…

Optimization and Control · Mathematics 2026-04-01 Govind M. Chari , Behçet Açıkmeşe

We present details of our implementation of the Wuppertal adaptive algebraic multigrid code DD-$\alpha$AMG on SIMD architectures, with particular emphasis on the Intel Xeon Phi processor (KNC) used in QPACE 2. As a smoother, the algorithm…

Computational Physics · Physics 2015-12-15 Simon Heybrock , Matthias Rottmann , Peter Georg , Tilo Wettig

Evolutionary model merging enables the creation of high-performing multi-task models but remains computationally prohibitive for consumer hardware. We introduce MERGE$^3$, an efficient framework that makes evolutionary merging feasible on a…

Neural and Evolutionary Computing · Computer Science 2025-05-12 Tommaso Mencattini , Adrian Robert Minut , Donato Crisostomi , Andrea Santilli , Emanuele Rodolà

We present efficient implementations of atom reconfiguration algorithms for both CPUs and GPUs, along with a batching routine to merge displacement operations for parallel execution. Leveraging graph-theoretic methods, our approach derives…

We employ pressure point analysis and roofline modeling to identify performance bottlenecks and determine an upper bound on the performance of the Canonical Polyadic Alternating Poisson Regression Multiplicative Update (CP-APR MU) algorithm…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-07-10 S. Isaac Geronimo Anderson , Keita Teranishi , Daniel M. Dunlavy , Jee Choi

Compression algorithms are important for data oriented tasks, especially in the era of Big Data. Modern processors equipped with powerful SIMD instruction sets, provide us an opportunity for achieving better compression performance.…

Information Retrieval · Computer Science 2015-04-15 Wayne Xin Zhao , Xudong Zhang , Daniel Lemire , Dongdong Shan , Jian-Yun Nie , Hongfei Yan , Ji-Rong Wen

We describe a modified SIMD architecture suitable for single-chip integration of a large number of processing elements, such as 1,000 or more. Important differences from traditional SIMD designs are: a) The size of the memory per processing…

Astrophysics · Physics 2007-05-23 Junichiro Makino

Planning under uncertainty for real-world robotics tasks, such as autonomous driving, requires reasoning in enormous high-dimensional belief spaces, rendering the problem computationally intensive. While parallelization offers scalability,…

Robotics · Computer Science 2026-02-10 Xuanjin Jin , Yanxin Dong , Bin Sun , Huan Xu , Zhihui Hao , XianPeng Lang , Panpan Cai

We investigate the achievable rate (AR) of a stacked intelligent metasurface (SIM)-aided holographic multiple-input multiple-output (HMIMO) system by jointly optimizing the SIM phase shifts and power allocation. Contrary to earlier studies…

Signal Processing · Electrical Eng. & Systems 2025-08-14 Eduard E. Bahingayi , Nemanja Stefan Perović , Le-Nam Tran

Galaxy interactions are a common phenomenon in clusters of galaxies. Especially major mergers are of particular importance, because they can change the morphological type of galaxies. They have an impact on the mass function of galaxies and…

Astrophysics of Galaxies · Physics 2015-05-14 J. Weniger , Ch. Theis , S. Harfst

AMD Xilinx's new Versal Adaptive Compute Acceleration Platform (ACAP) is an FPGA architecture combining reconfigurable fabric with other on-chip hardened compute resources. AI engines are one of these and, by operating in a highly…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-01-31 Nick Brown

We describe recent work focused towards a better understanding of red supergiant stars using 3D radiative-hydrodynamics (RHD) simulations with CO5BOLD. A small number of simulations now exist that span up to seven years of stellar time, at…

Solar and Stellar Astrophysics · Physics 2013-05-30 B. Plez , A. Chiavassa

This paper presents VoxelMap++: a voxel mapping method with plane merging which can effectively improve the accuracy and efficiency of LiDAR(-inertial) based simultaneous localization and mapping (SLAM). This map is a collection of voxels…

Robotics · Computer Science 2023-08-08 Yifei Yuan , Chang Wu , Yuan You , Xiaotong Kong , Ying Zhang , Qiyan Li

Recent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be transformed to matrix multiplications. However, this…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-27 Qiqi GU , Chenpeng Wu , Heng Shi , Jianguo Yao

Thermal density and hot spots limit three-dimensional (3D) implementation of massively-parallel SIMD processors and prohibit stacking DRAM dies above them. This study proposes replacing SIMD by an Associative Processor (AP). AP exhibits…

Hardware Architecture · Computer Science 2013-07-16 Leonid Yavits , Amir Morad , Ran Ginosar

Sparse Matrix-Matrix multiplication is a key kernel that has applications in several domains such as scientific computing and graph analysis. Several algorithms have been studied in the past for this foundational kernel. In this paper, we…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-01-10 Mehmet Deveci , Christian Trott , Sivasankaran Rajamanickam

An Eulerian TVD code and a Lagrangian SPH code are used to simulate the off-axis collision of equal-mass main sequence stars in order to address the question of whether stellar mergers can produce a remnant star where the interior has been…

Astrophysics · Physics 2008-11-26 Hy Trac , Alison Sills , Ue-Li Pen

The challenge to fully exploit the potential of existing and upcoming scientific instruments like large single-dish radio telescopes is to process the collected massive data effectively and efficiently. As a "quasi 2D stencil computation"…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-07-12 Hao Wang , Ce Yu , Jian Xiao , Shanjiang Tang , Min Long , Ming Zhu

At least 25 kinds of detector-like devices need to be integrated in Phase I of the High Energy Photon Source (HEPS), and the work needs to be carefully planned to maximise productivity with highly limited human resources. After a systematic…

Instrumentation and Detectors · Physics 2024-11-06 Qun Zhang , Peng-Cheng Li , Ling-Zhu Bian , Chun Li , Zong-Yang Yue , Cheng-Long Zhang , Zhuo-Feng Zhao , Yi Zhang , Gang Li , Ai-Yu Zhou , Yu Liu
‹ Prev 1 4 5 6 7 8 10 Next ›