vix.ing · top · new · best · stats

Efficient implementation of the overlap operator on multi-GPUs

2011/06/24 by Andrei Alexandru, Alexandru, Andrei, Michael Lujan +8 · 5 citations
Computer Science · Physics and Astronomy · #Advanced Data Storage Technologies #CUDA #Central processing unit #Cluster (spacecraft) #Computational science #Computer science #FOS: Computer and information sciences #FOS: Physical sciences #GPU cluster #High Energy Physics - Lattice (hep-lat) #Memory footprint #Operating system #Operator (biology) #Parallel computing #Particle physics theoretical and experimental studies #Performance (cs.PF) #Quantum Chromodynamics and Particle Interactions #Supercomputer #cs.PF #hep-lat

paper · pdf · doi:10.48550/arxiv.1106.4964

published in arXiv (Cornell University) (Cornell University) · 8 pages with 10 figures; accepted for presentation at the 2011 Symposium on Application Accelerators in High Performance Computing (Knoxville, July 19-20, 2011)

arxiv created 2011/06/24 · openalex publication_date 2011/06/24 · arxiv updated 2011/06/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Lattice QCD calculations were one of the first applications to show the potential of GPUs in the area of high performance computing. Our interest is to find ways to effectively use GPUs for lattice calculations using the overlap operator. The large memory footprint of these codes requires the use of multiple GPUs in parallel. In this paper we show the methods we used to implement this operator efficiently. We run our codes both on a GPU cluster and a CPU cluster with similar interconnects. We find that to match performance the CPU cluster requires 20-30 times more CPU cores than GPUs.

Citations

Related