Skip to main navigation Skip to search Skip to main content

Vectorization and minimization of memory footprint for linear high-order discontinuous galerkin schemes

  • Jean Matthieu Gallard
  • , Leonhard Rannabauer
  • , Anne Reinarz
  • , Michael Bader
  • Technical University of Munich

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

4 Scopus citations

Abstract

We present a sequence of optimizations to the performance-critical compute kernels of the high-order discontinuous Galerkin solver of the hyperbolic PDE engine ExaHyPE - successively tackling bottlenecks due to SIMD operations, cache hierarchies and restrictions in the software design.Starting from a generic scalar implementation of the numerical scheme, our first optimized variant applies state-of the-art optimization techniques by vectorizing loops, improving the data layout and using Loop-over-GEMM to perform tensor contractions via highly optimized matrix multiplication functions provided by the LIBXSMM library. We show that memory stalls due to a memory footprint exceeding our L2 cache size hindered the vectorization gains. We therefore introduce a new kernel that applies a sum factorization approach to reduce the kernel's memory footprint and improve its cache locality. With the L2 cache bottleneck removed, we were able to exploit additional vectorization opportunities, by introducing a hybrid Array-of-Structure-of-Array data layout that solves the data layout conflict between matrix multiplications kernels and the point-wise functions to implement PDE-specific terms.With this last kernel, evaluated in a benchmark simulation at high polynomial order, only 2% of the floating point operations are still performed using scalar instructions and 22.5% of the available performance is achieved.

Original languageEnglish
Title of host publicationProceedings - 2020 IEEE 34th International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2020
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages711-720
Number of pages10
ISBN (Electronic)9781728174457
DOIs
StatePublished - May 2020
Event34th IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2020 - Virtual, Online, United States
Duration: 18 May 202022 May 2020

Publication series

NameProceedings - 2020 IEEE 34th International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2020

Conference

Conference34th IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2020
Country/TerritoryUnited States
CityVirtual, Online
Period18/05/2022/05/20

Keywords

  • ADER
  • Array-of-Struct-of-Array
  • Code Generation
  • ExaHyPE
  • High-Order Discontinuous Galerkin
  • Hyperbolic PDE Systems
  • Vectorization

Fingerprint

Dive into the research topics of 'Vectorization and minimization of memory footprint for linear high-order discontinuous galerkin schemes'. Together they form a unique fingerprint.

Cite this