Results 71 to 80 of about 926 (201)

Lightweight GPU-Accelerated Parallel Processing of the SCHISM Model Using CUDA Fortran

open access: yesJournal of Marine Science and Engineering
The SCHISM model is widely used for ocean numerical simulations, but its computational efficiency is constrained by the substantial resources it requires.
Hongchun Zhang   +7 more
doaj   +1 more source

Performance Portability Strategies for Grid C++ Expression Templates

open access: yesEPJ Web of Conferences, 2018
One of the key requirements for the Lattice QCD Application Development as part of the US Exascale Computing Project is performance portability across multiple architectures.
Boyle Peter A.   +5 more
doaj   +1 more source

Portable multi- and many-core performance for finite-difference or finite-element codes – application to the free-surface component of NEMO (NEMOLite2D 1.0) [PDF]

open access: yesGeoscientific Model Development, 2018
We present an approach which we call PSyKAl that is designed to achieve portable performance for parallel finite-difference, finite-volume, and finite-element earth-system models.
A. R. Porter   +6 more
doaj   +1 more source

Toward exascale climate modelling: a python DSL approach to ICON's (icosahedral non-hydrostatic) dynamical core (icon-exclaim v0.2.0) [PDF]

open access: yesGeoscientific Model Development
A refactored atmospheric dynamical core of the ICON model implemented in GT4Py, a Python-based domain-specific language designed for performance portability across heterogeneous CPU-GPU architectures, is presented.
A. Dipankar   +28 more
doaj   +1 more source

Lattice Simulations using OpenACC compilers [PDF]

open access: yesProceedings of 31st International Symposium on Lattice Field Theory LATTICE 2013 — PoS(LATTICE 2013), 2014
OpenACC compilers allow one to use Graphics Processing Units without having to write explicit CUDA codes. Programs can be modified incrementally using OpenMP like directives which causes the compiler to generate CUDA kernels to be run on the GPUs. In this article we look at the performance gain in lattice simulations with dynamical fermions using ...
openaire   +2 more sources

XcalableACC: An Integration of XcalableMP and OpenACC [PDF]

open access: yes, 2020
AbstractXcalableACC (XACC) is an extension of XcalableMP for accelerated clusters. It is defined as a diagonal integration of XcalableMP and OpenACC, which is another directive-based language designed to program heterogeneous CPU/accelerator systems. XACC has features for handling distributed-memory parallelism, inherited from XMP, offloading tasks to ...
Akihiro Tabuchi   +4 more
openaire   +1 more source

Offline Atmospheric Transport on a Global Mesh of Hexagons

open access: yesJournal of Geophysical Research: Atmospheres, Volume 130, Issue 11, 16 June 2025.
Abstract We present a new version of the offline transport model from the LMDz atmospheric general circulation model. It paves the globe with hexagons and 12 pentagons of similar surface areas rather than with regular longitude‐latitude rectangles. It is available in a complete, nonlinear version and in a linearized configuration for use in variational
Frédéric Chevallier   +4 more
wiley   +1 more source

Multi-GPU programming with OpenACC and MPI [PDF]

open access: yes, 2016
International audienceIn this session you will learn how to program multi GPU systems or GPU clusters using the Message Passing Interface (MPI) and OpenACC.
Kraus, Jiri, Etancelin, Jean-Matthieu
core   +2 more sources

Operational numerical weather prediction with ICON on GPUs (version 2024.10) [PDF]

open access: yesGeoscientific Model Development
Numerical weather prediction and climate models require continuous adaptation to take advantage of advances in high-performance computing hardware. This paper presents the port of the ICON model to GPUs using OpenACC compiler directives for numerical ...
X. Lapillonne   +33 more
doaj   +1 more source

swPHoToNs: Toward trillion‐body‐scale cosmological N‐body simulations on Sunway TaihuLight supercomputer

open access: yesEngineering Reports, Volume 7, Issue 3, March 2025.
We present the details of implementing a highly efficient and scalable Cosmological N‐body simulation framework on the heterogeneous many‐core supercomputer Sunway TaihuLight. We manage to conduct cosmological simulations which contain up to 1.6 trillion particles, obtaining a sustained performance of 56.3 PFlops with a weak‐scaling parallel efficiency
Zhao Liu   +8 more
wiley   +1 more source

Home - About - Disclaimer - Privacy