Lightweight GPU-Accelerated Parallel Processing of the SCHISM Model Using CUDA Fortran
The SCHISM model is widely used for ocean numerical simulations, but its computational efficiency is constrained by the substantial resources it requires.
Hongchun Zhang +7 more
doaj +1 more source
Performance Portability Strategies for Grid C++ Expression Templates
One of the key requirements for the Lattice QCD Application Development as part of the US Exascale Computing Project is performance portability across multiple architectures.
Boyle Peter A. +5 more
doaj +1 more source
Portable multi- and many-core performance for finite-difference or finite-element codes – application to the free-surface component of NEMO (NEMOLite2D 1.0) [PDF]
We present an approach which we call PSyKAl that is designed to achieve portable performance for parallel finite-difference, finite-volume, and finite-element earth-system models.
A. R. Porter +6 more
doaj +1 more source
Toward exascale climate modelling: a python DSL approach to ICON's (icosahedral non-hydrostatic) dynamical core (icon-exclaim v0.2.0) [PDF]
A refactored atmospheric dynamical core of the ICON model implemented in GT4Py, a Python-based domain-specific language designed for performance portability across heterogeneous CPU-GPU architectures, is presented.
A. Dipankar +28 more
doaj +1 more source
Lattice Simulations using OpenACC compilers [PDF]
OpenACC compilers allow one to use Graphics Processing Units without having to write explicit CUDA codes. Programs can be modified incrementally using OpenMP like directives which causes the compiler to generate CUDA kernels to be run on the GPUs. In this article we look at the performance gain in lattice simulations with dynamical fermions using ...
openaire +2 more sources
XcalableACC: An Integration of XcalableMP and OpenACC [PDF]
AbstractXcalableACC (XACC) is an extension of XcalableMP for accelerated clusters. It is defined as a diagonal integration of XcalableMP and OpenACC, which is another directive-based language designed to program heterogeneous CPU/accelerator systems. XACC has features for handling distributed-memory parallelism, inherited from XMP, offloading tasks to ...
Akihiro Tabuchi +4 more
openaire +1 more source
Offline Atmospheric Transport on a Global Mesh of Hexagons
Abstract We present a new version of the offline transport model from the LMDz atmospheric general circulation model. It paves the globe with hexagons and 12 pentagons of similar surface areas rather than with regular longitude‐latitude rectangles. It is available in a complete, nonlinear version and in a linearized configuration for use in variational
Frédéric Chevallier +4 more
wiley +1 more source
Multi-GPU programming with OpenACC and MPI [PDF]
International audienceIn this session you will learn how to program multi GPU systems or GPU clusters using the Message Passing Interface (MPI) and OpenACC.
Kraus, Jiri, Etancelin, Jean-Matthieu
core +2 more sources
Operational numerical weather prediction with ICON on GPUs (version 2024.10) [PDF]
Numerical weather prediction and climate models require continuous adaptation to take advantage of advances in high-performance computing hardware. This paper presents the port of the ICON model to GPUs using OpenACC compiler directives for numerical ...
X. Lapillonne +33 more
doaj +1 more source
We present the details of implementing a highly efficient and scalable Cosmological N‐body simulation framework on the heterogeneous many‐core supercomputer Sunway TaihuLight. We manage to conduct cosmological simulations which contain up to 1.6 trillion particles, obtaining a sustained performance of 56.3 PFlops with a weak‐scaling parallel efficiency
Zhao Liu +8 more
wiley +1 more source

