2025/09/23 by B. Ferrari, Cyrill Püntener, Ferrari, Marcel +5
Computer Science · Engineering · #65F08 #65N22 #65N55 #76M20 #Advanced Numerical Methods in Computational Mathematics #C.1.4 #Computational Fluid Dynamics and Aerodynamics #Computational Physics (physics.comp-ph) #D.1.3 #F.2.1 #FOS: Mathematics #FOS: Physical sciences #G.1.8 #Matrix Theory and Algorithms #Numerical Analysis (math.NA)
paper · pdf · doi:10.48550/arxiv.2509.19061
openalex publication_date 2025/09/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present the design, implementation, and evaluation of optimized matrix-free stencil kernels for multigrid smoothing in the incompressible Stokes equations with variable viscosity, motivated by geophysical flow problems. We investigate five smoother variants derived from different optimisation strategies: Red-Black Gauss-Seidel, Jacobi, fused Jacobi, blocked fused Jacobi, and a novel Jacobi smoother with RAS-type temporal blocking, a strategy that applies local iterations on overlapping tiles to improve cache reuse. To ensure correctness, we introduce an energy-based residual norm that balances velocity and pressure contributions, and validate all implementations using a high-contrast sinker benchmark representative of realistic geodynamic numerical models. Our performance study on NVIDIA GH200 Grace Hopper nodes of the ALPS supercomputer demonstrates that all smoothers scale well within a single NUMA domain, but the RAS-Jacobi smoother consistently achieves the best performance at higher core counts. It sustains over 90% weak-scaling efficiency up to 64 cores and delivers up to a threefold speedup compared to the C++ Jacobi baseline, owing to improved cache reuse and reduced memory traffic. These results show that temporal blocking, already employed in distributed-memory solvers to reduce communication, can also provide substantial benefits at the socket and NUMA level. This work highlights the importance of cache-aware stencil design for harnessing modern heterogeneous architectures and lays the groundwork for extending RAS-type temporal blocking strategies to three-dimensional problems and GPU accelerators.