English
Related papers

Related papers: Architecture-aware $h$-to-$p$ optimisation: spectr…

200 papers

The $hp$-adaptive finite element method (FEM) - where one independently chooses the mesh size ($h$) and polynomial degree ($p$) to be used on each cell - has long been known to have better theoretical convergence properties than either $h$-…

Numerical Analysis · Mathematics 2023-09-14 Marc Fehling , Wolfgang Bangerth

High Performance Computing (HPC) platforms allow scientists to model computationally intensive algorithms. HPC clusters increasingly use General-Purpose Graphics Processing Units (GPGPUs) as accelerators; FPGAs provide an attractive…

Hardware Architecture · Computer Science 2015-04-20 Syed Waqar Nabi , Saji N. Hameed , Wim Vanderbauwhede

This paper deals with the applications of stochastic spectral methods for structural topology optimization in the presence of uncertainties. A non-intrusive polynomial chaos expansion is integrated into a topology optimization algorithm to…

Computational Engineering, Finance, and Science · Computer Science 2021-08-04 Nilton Cuellar , Anderson Pereira , Ivan F. M. Menezes , Americo Cunha

Sparse Matricized Tensor Times Khatri-Rao Product (spMTTKRP) is the bottleneck kernel of sparse tensor decomposition. In this work, we propose a GPU-based algorithm design to address the key challenges in accelerating spMTTKRP computation,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-05-15 Sasindu Wijeratne , Rajgopal Kannan , Viktor Prasanna

We present and analyze a new finite element method for solving interface problems on a triangular grid. The method locally modifies a given triangulation such that the interfaces are accurately resolved and the maximal angle condition…

Numerical Analysis · Mathematics 2026-04-02 Peter Gangl , Ulrich Langer

Accurate hardware performance models are critical to efficient code generation. They can be used by compilers to make heuristic decisions, by superoptimizers as a minimization objective, or by autotuners to find an optimal configuration for…

In this article, a new generic higher-order finite-element framework for massively parallel simulations is presented. The modular software architecture is carefully designed to exploit the resources of modern and future supercomputers.…

Mathematical Software · Computer Science 2018-05-28 Nils Kohl , Dominik Thönnes , Daniel Drzisga , Dominik Bartuschat , Ulrich Rüde

This work presents Squeeze, an efficient compact fractal processing scheme for tensor core GPUs. By combining discrete-space transformations between compact and expanded forms, one can do data-parallel computation on a fractal with…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-01-04 Felipe A. Quezada , Cristóbal A. Navarro , Nancy Hitschfeld , Benjamin Bustos

In high-order finite element analysis for elasticity, matrix-free (PA) methods are a key technology for overcoming the memory bottleneck of traditional Full Assembly (FA). However, existing implementations fail to fully exploit the special…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-14 Dali Chang , Chong Zhang , Kaiqi Zhang , Mingguan Yang , Huiyuan Li , Weiqiang Kong

We describe a framework for controlling and improving the quality of high-order finite element meshes based on extensions of the Target-Matrix Optimization Paradigm (TMOP) of Knupp. This approach allows high-order applications to have a…

Numerical Analysis · Mathematics 2018-07-27 Veselin Dobrev , Patrick Knupp , Tzanio Kolev , Ketan Mittal , Vladimir Tomov

To solve large-scale or high-resolution topology optimization problem, a novel algorithm is developed based on modified bi-directional evolutionary structure optimization (BESO) and extended finite element method (XFEM). Within XFEM, a set…

Applied Physics · Physics 2026-04-07 Hongxin Wang , Jie Liu , Guilin Wen

Bloom filters are a fundamental data structure for approximate membership queries, with applications ranging from data analytics to databases and genomics. Several variants have been proposed to accommodate parallel architectures. GPUs,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-18 Daniel Jünger , Kevin Kristensen , Yunsong Wang , Xiangyao Yu , Bertil Schmidt

We generalize the two dimensional mixed finite elements of Arbogast and Correa [T. Arbogast and M. R. Correa, SIAM J. Numer. Anal., 54 (2016), pp. 3332--3356] defined on quadrilaterals to three dimensional cuboidal hexahedra. The…

Numerical Analysis · Mathematics 2018-11-06 Todd Arbogast , Zhen Tao

Numerical methods such as the Finite Element Method (FEM) have been successfully adapted to utilize the computational power of GPU accelerators. However, much of the effort around applying FEM to GPU's has been focused on high-order FEM due…

In the context of adaptive remeshing, the virtual element method provides significant advantages over the finite element method. The attractive features of the virtual element method, such as the permission of arbitrary element geometries,…

Numerical Analysis · Mathematics 2023-08-16 Daniel van Huyssteen , Felipe Lopez Rivarola , Guillermo Etse , Paul Steinmann

In this paper the hp-version of the boundary element method is applied to the electric field integral equation on a piecewise plane (open or closed) Lipschitz surface. The underlying meshes are supposed to be quasi-uniform. We use…

Numerical Analysis · Mathematics 2008-10-21 Alexei Bespalov , Norbert Heuer

Transformers have revolutionized deep learning and generative modeling to enable unprecedented advancements in natural language processing tasks and beyond. However, designing hardware accelerators for executing transformer models is…

Hardware Architecture · Computer Science 2024-08-08 Pratyush Dhingra , Janardhan Rao Doppa , Partha Pratim Pande

Numerical integration of the stiffness matrix in higher order finite element (FE) methods is recognized as one of the heaviest computational tasks in a FE solver. The problem becomes even more relevant when computing the Gram matrix in the…

Numerical Analysis · Mathematics 2017-11-06 Jaime Mora , Leszek Demkowicz

As GPU architectures rapidly evolve to meet the growing demands of exascale computing and machine learning, the performance implications of architectural innovations remain poorly understood across diverse workloads. NVIDIA Blackwell (B200)…

Hardware Architecture · Computer Science 2026-03-04 Aaron Jarmusch , Sunita Chandrasekaran

This paper presents the development of a complete CAD-compatible framework for structural shape optimization in 3D. The boundaries of the domain are described using NURBS while the interior is discretized with B\'ezier tetrahedra. The…

Computational Engineering, Finance, and Science · Computer Science 2024-01-01 Jorge López , Cosmin Anitescu , Timon Rabczuk