中文
相关论文

相关论文: From Merging Frameworks to Merging Stars: Experien…

200 篇论文

We review the recent optimizations of gravitational $N$-body kernels for running them on graphics processing units (GPUs), on single hosts and massive parallel platforms. For each of the two main $N$-body techniques, direct summation and…

天体物理仪器与方法 · 物理学 2014-09-22 Simon Portegies Zwart , Jeroen Bédorf

We present a GPU-accelerated backend for QOCO, a C-based solver for quadratic objective second-order cone programs (SOCPs) based on a primal-dual interior point method. Our backend uses NVIDIA's cuDSS library to perform a direct sparse LDL…

最优化与控制 · 数学 2026-04-01 Govind M. Chari , Behçet Açıkmeşe

We present details of our implementation of the Wuppertal adaptive algebraic multigrid code DD-$\alpha$AMG on SIMD architectures, with particular emphasis on the Intel Xeon Phi processor (KNC) used in QPACE 2. As a smoother, the algorithm…

计算物理 · 物理学 2015-12-15 Simon Heybrock , Matthias Rottmann , Peter Georg , Tilo Wettig

Evolutionary model merging enables the creation of high-performing multi-task models but remains computationally prohibitive for consumer hardware. We introduce MERGE$^3$, an efficient framework that makes evolutionary merging feasible on a…

神经与进化计算 · 计算机科学 2025-05-12 Tommaso Mencattini , Adrian Robert Minut , Donato Crisostomi , Andrea Santilli , Emanuele Rodolà

We present efficient implementations of atom reconfiguration algorithms for both CPUs and GPUs, along with a batching routine to merge displacement operations for parallel execution. Leveraging graph-theoretic methods, our approach derives…

We employ pressure point analysis and roofline modeling to identify performance bottlenecks and determine an upper bound on the performance of the Canonical Polyadic Alternating Poisson Regression Multiplicative Update (CP-APR MU) algorithm…

分布式、并行与集群计算 · 计算机科学 2023-07-10 S. Isaac Geronimo Anderson , Keita Teranishi , Daniel M. Dunlavy , Jee Choi

Compression algorithms are important for data oriented tasks, especially in the era of Big Data. Modern processors equipped with powerful SIMD instruction sets, provide us an opportunity for achieving better compression performance.…

信息检索 · 计算机科学 2015-04-15 Wayne Xin Zhao , Xudong Zhang , Daniel Lemire , Dongdong Shan , Jian-Yun Nie , Hongfei Yan , Ji-Rong Wen

We describe a modified SIMD architecture suitable for single-chip integration of a large number of processing elements, such as 1,000 or more. Important differences from traditional SIMD designs are: a) The size of the memory per processing…

天体物理学 · 物理学 2007-05-23 Junichiro Makino

Planning under uncertainty for real-world robotics tasks, such as autonomous driving, requires reasoning in enormous high-dimensional belief spaces, rendering the problem computationally intensive. While parallelization offers scalability,…

机器人学 · 计算机科学 2026-02-10 Xuanjin Jin , Yanxin Dong , Bin Sun , Huan Xu , Zhihui Hao , XianPeng Lang , Panpan Cai

We investigate the achievable rate (AR) of a stacked intelligent metasurface (SIM)-aided holographic multiple-input multiple-output (HMIMO) system by jointly optimizing the SIM phase shifts and power allocation. Contrary to earlier studies…

信号处理 · 电气工程与系统科学 2025-08-14 Eduard E. Bahingayi , Nemanja Stefan Perović , Le-Nam Tran

Galaxy interactions are a common phenomenon in clusters of galaxies. Especially major mergers are of particular importance, because they can change the morphological type of galaxies. They have an impact on the mass function of galaxies and…

星系天体物理 · 物理学 2015-05-14 J. Weniger , Ch. Theis , S. Harfst

AMD Xilinx's new Versal Adaptive Compute Acceleration Platform (ACAP) is an FPGA architecture combining reconfigurable fabric with other on-chip hardened compute resources. AI engines are one of these and, by operating in a highly…

分布式、并行与集群计算 · 计算机科学 2023-01-31 Nick Brown

We describe recent work focused towards a better understanding of red supergiant stars using 3D radiative-hydrodynamics (RHD) simulations with CO5BOLD. A small number of simulations now exist that span up to seven years of stellar time, at…

太阳与恒星天体物理 · 物理学 2013-05-30 B. Plez , A. Chiavassa

This paper presents VoxelMap++: a voxel mapping method with plane merging which can effectively improve the accuracy and efficiency of LiDAR(-inertial) based simultaneous localization and mapping (SLAM). This map is a collection of voxels…

机器人学 · 计算机科学 2023-08-08 Yifei Yuan , Chang Wu , Yuan You , Xiaotong Kong , Ying Zhang , Qiyan Li

Recent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be transformed to matrix multiplications. However, this…

分布式、并行与集群计算 · 计算机科学 2026-01-27 Qiqi GU , Chenpeng Wu , Heng Shi , Jianguo Yao

Thermal density and hot spots limit three-dimensional (3D) implementation of massively-parallel SIMD processors and prohibit stacking DRAM dies above them. This study proposes replacing SIMD by an Associative Processor (AP). AP exhibits…

硬件体系结构 · 计算机科学 2013-07-16 Leonid Yavits , Amir Morad , Ran Ginosar

Sparse Matrix-Matrix multiplication is a key kernel that has applications in several domains such as scientific computing and graph analysis. Several algorithms have been studied in the past for this foundational kernel. In this paper, we…

分布式、并行与集群计算 · 计算机科学 2018-01-10 Mehmet Deveci , Christian Trott , Sivasankaran Rajamanickam

An Eulerian TVD code and a Lagrangian SPH code are used to simulate the off-axis collision of equal-mass main sequence stars in order to address the question of whether stellar mergers can produce a remnant star where the interior has been…

天体物理学 · 物理学 2008-11-26 Hy Trac , Alison Sills , Ue-Li Pen

The challenge to fully exploit the potential of existing and upcoming scientific instruments like large single-dish radio telescopes is to process the collected massive data effectively and efficiently. As a "quasi 2D stencil computation"…

分布式、并行与集群计算 · 计算机科学 2022-07-12 Hao Wang , Ce Yu , Jian Xiao , Shanjiang Tang , Min Long , Ming Zhu

At least 25 kinds of detector-like devices need to be integrated in Phase I of the High Energy Photon Source (HEPS), and the work needs to be carefully planned to maximise productivity with highly limited human resources. After a systematic…