vix.ing · top · new · best · stats

Massively Parallel Algorithms for the Lattice Boltzmann Method on NonUniform Grids

2015/08/31 by Florian Schornbaum, Ulrich Rüde · 117 citations
Computer Science · Engineering · Mathematics · #Adaptive mesh refinement #Advanced Data Storage Technologies #Aerosol Filtration and Electrostatic Precipitation #Algorithm #Computational science #Computer science #Distributed computing #Grid #Lattice Boltzmann Simulation Studies #Lattice Boltzmann methods #Massively parallel #Mathematics #Operating system #Parallel computing #Petascale computing #Physics #Scalability #Scaling #Supercomputer #cs.CE #cs.DC

paper · pdf · doi:10.1137/15m1035240

published in SIAM Journal on Scientific Computing 38(2), C96-C126 (Society for Industrial and Applied Mathematics) · 32 pages, 20 figures, 4 tables

openalex publication_date 2016/01/01 · arxiv created 2016/01/21 · arxiv updated 2016/05/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

The lattice Boltzmann method exhibits excellent scalability on current supercomputing systems and has thus increasingly become an alternative method for large-scale nonstationary flow simulations, reaching up to a trillion (1012) grid nodes. Additionally, grid refinement can lead to substantial savings in memory and compute time. These savings, however, come at the cost of much more complex data structures and algorithms. In particular, the interface between subdomains with different grid sizes must receive special treatment. In this article, we present parallel algorithms, distributed data structures, and communication routines that are implemented in the software framework waLBerla in order to support large-scale, massively parallel lattice Boltzmann-based simulations on nonuniform grids. Additionally, we evaluate the performance of our approach on two current petascale supercomputers. On an IBM Blue Gene/Q system, the largest weak scaling benchmarks with refined grids are executed with almost 2 million threads, demonstrating not only near-perfect scalability but also an absolute performance of close to a trillion lattice Boltzmann cell updates per second. On an Intel-based system, the strong scaling of a simulation with refined grids and a total of more than 8.5 million cells is demonstrated to reach a performance of less than 1 millisecond per time step. This enables simulations with complex, nonuniform grids and 4 million time steps per hour compute time.

Citations

Cited by