vix.ing · top · new · best · stats · spec

Efficient Algorithms for Monte Carlo Particle Transport on AI Accelerator Hardware

2023/11/03 by John Tramm, Tramm, John, Bryce Allen +7 · 3 citations
Computer Science · #Advanced Neural Network Applications #D.1.3 #Distributed #FOS: Computer and information sciences #J.2 #Parallel #Parallel Computing and Optimization Techniques #Performance (cs.PF) #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)

paper · pdf · doi:10.48550/arxiv.2311.01739

openalex publication_date 2023/11/03 · openalex created_date 2023/11/07 · openalex updated_date 2026/08/01

Abstract

The recent trend toward deep learning has led to the development of a variety of highly innovative AI accelerator architectures. One such architecture, the Cerebras Wafer-Scale Engine 2 (WSE-2), features 40 GB of on-chip SRAM, making it a potentially attractive platform for latency- or bandwidth-bound HPC simulation workloads. In this study, we examine the feasibility of performing continuous energy Monte Carlo (MC) particle transport on the WSE-2 by porting a key kernel from the MC transport algorithm to Cerebras's CSL programming model. New algorithms for minimizing communication costs and for handling load balancing are developed and tested. The WSE-2 is found to run 130 times faster than a highly optimized CUDA version of the kernel run on an NVIDIA A100 GPU -- significantly outpacing the expected performance increase given the difference in transistor counts between the architectures.

Cited by

Related