Mathematical Software · Computer Science
Accelerating High-Order Mesh Optimization Using Finite Element Partial Assembly on GPUs
Jean-Sylvain Camier, Veselin Dobrev, Patrick Knupp, Tzanio Kolev +3
2022-12-28
Numerical Analysis · Mathematics
Architecture-aware $h$-to-$p$ optimisation: spectral/$hp$ element operators for mixed-element meshes
Jacques Y. Xing, Boyang Xia, Diego Renner, Chris D. Cantwell +3
2026-04-07
Distributed, Parallel, and Cluster Computing · Computer Science
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
Hangda Liu, Boyu Diao, Yu Yang, Wenxin Chen +2
2025-02-18
Distributed, Parallel, and Cluster Computing · Computer Science
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
Jiqun Tu, Ian Karlin, John Camier, Veselin Dobrev +3
2026-04-13
Numerical Analysis · Mathematics
Generalized L-product for hight order tensors with applications using GPU computation
Abdeslem Hafid Ben Tbib, Mouad Elalj, Anas EL Hachimi, Khalide Jbilou +1
2022-02-08
Numerical Analysis · Mathematics
An implementation of tensor product patch smoothers on GPU
Cu Cui, Paul Grosse-Bley, Guido Kanschat, Robert Strzodka
2025-05-07
Numerical Analysis · Mathematics
Performant low-order matrix-free finite element kernels on GPU architectures
Randolph R. Settgast, Yohann Dudouit, Nicola Castelletto, William R. Tobin +2
2023-08-24
Distributed, Parallel, and Cluster Computing · Computer Science
Analyzing GPU Tensor Core Potential for Fast Reductions
Roberto Carrasco, Raimundo Vega, Cristóbal A. Navarro
2019-03-12
Distributed, Parallel, and Cluster Computing · Computer Science
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
Daniel Nichols, Konstantinos Parasyris, Charles Jekel, Abhinav Bhatele +1
2025-10-21
Machine Learning · Computer Science
Kernel methods through the roof: handling billions of points efficiently
Giacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro Rudi
2020-11-30
Numerical Analysis · Mathematics
Performance of linear solvers in tensor-train format on current multicore architectures
Melven Röhrig-Zöllner, Manuel Joey Becklas, Jonas Thies, Achim Basermann
2024-10-25
Machine Learning · Computer Science
Learning to Optimize Tensor Programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang +4
2019-01-10
Performance · Computer Science
A Learned Performance Model for Tensor Processing Units
Samuel J. Kaufman, Phitchaya Mangpo Phothilimthana, Yanqi Zhou, Charith Mendis +3
2021-03-19
Distributed, Parallel, and Cluster Computing · Computer Science
Optimization of Tensor-product Operations in Nekbone on GPUs
Martin Karp, Niclas Jansson, Artur Podobas, Philipp Schlatter +1
2020-05-28
Machine Learning · Computer Science
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
Qitong Sun, Jun Han, Tianlin Li, Zhe Tang +5
2026-03-12
Numerical Analysis · Mathematics
Optimization of Functions Given in the Tensor Train Format
Andrei Chertkov, Gleb Ryzhakov, Georgii Novikov, Ivan Oseledets
2022-09-30
Distributed, Parallel, and Cluster Computing · Computer Science
Benchmarking optimization algorithms for auto-tuning GPU kernels
Richard Schoonhoven, Ben van Werkhoven, Kees Joost Batenburg
2022-10-05
Software Engineering · Computer Science
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
Haonan Li, Keyu Man, Partha Kanuparthy, Hanning Chen +5
2025-12-16
Distributed, Parallel, and Cluster Computing · Computer Science
Low-ordered Orthogonal Voxel Finite Element with INT8 Tensor Cores for GPU-based Explicit Elastic Wave Propagation Analysis
Tsuyoshi Ichimura, Kohei Fujita, Muneo Hori, Maddegedara Lalith
2024-07-02