Results 11 to 20 of about 926 (201)
An Early Performance Comparison of CUDA and OpenACC [PDF]
This paper presents a performance comparison between CUDA and OpenACC. The performance analysis focuses on programming models and underlying compilers. In addition, we proposed a Performance Ratio of Data Sensitivity (PRoDS) metric to objectively compare
Li Xuechao, Shih Po-Chou
doaj +4 more sources
ACC_TEST: Hybrid Testing Approach for OpenACC-Based Programs [PDF]
In recent years, OpenACC has been used in many supercomputers and attracted many non-computer science specialists for parallelizing their programs in different scientific fields, including weather forecasting and simulations.
Fathy Elbouraey Eassa +5 more
doaj +2 more sources
Portable LQCD Monte Carlo code using OpenACC [PDF]
Varying from multi-core CPU processors to many-core GPUs, the present scenario of HPC architectures is extremely heterogeneous. In this context, code portability is increasingly important for easy maintainability of applications; this is relevant in ...
Bonati Claudio +8 more
doaj +4 more sources
Towards OmpSs-2 and OpenACC interoperation [PDF]
The increasing demand in HPC to utilize accelerators has motivated the development of pragma-based directives to target these devices. OmpSs-2 and OpenACC are both directive-based solutions that allow application programmers to utilize accelerators. The two leverage distinct types of parallelism: task parallelism and data parallelism, respectively. Non-
Korakitis, Orestis +5 more
openaire +4 more sources
Gadget3 on GPUs with OpenACC [PDF]
We present preliminary results of a GPU porting of all main Gadget3 modules (gravity computation, SPH density computation, SPH hydrodynamic force, and thermal conduction) using OpenACC directives. Here we assign one GPU to each MPI rank and exploit both the host and accellerator capabilities by overlapping computations on the CPUs and GPUs: while GPUs ...
Antonio Ragagnin +7 more
openaire +4 more sources
Accelerating Lattice Boltzmann Applications with OpenACC [PDF]
An increasingly large number of HPC systems rely on heterogeneous architectures combining traditional multi-core CPUs with power efficient accelerators. Designing efficient applications for these systems has been troublesome in the past as accelerators could usually be programmed only using specific programming languages – such as CUDA – threatening ...
Calore, Enrico +3 more
openaire +3 more sources
This paper explores Machine Learning (ML) parameterizations for radiative transfer in the ICOsahedral Nonhydrostatic weather and climate model (ICON) and investigates the achieved ML model speed‐up with ICON running on graphics processing units (GPUs ...
Guillaume Bertoli +7 more
doaj +2 more sources
Numerical Integration of Slater Basis Functions Over Prolate Spheroidal Grids. [PDF]
Atom‐centered Slater functions are integrated over a prolate spheroidal grid, allowing the use of larger basis sets in correlated electronic structure computations. ABSTRACT Slater basis functions have desirable properties that can improve electronic structure simulations, but improved numerical integration methods are needed. This work builds upon the
Stark A +5 more
europepmc +2 more sources
Due to road curvature and sensors’ limited field of view, as‐built highway curves would pose an operational challenge to the adaptive cruise control (ACC) system and its shared control.
Shuyi Wang +5 more
doaj +2 more sources
NekBone with Optimized OpenACC directives [PDF]
Accelerators and, in particular, Graphics Processing Units (GPUs) have emerged as promising computing technologies which may be suitable for the future Exascale systems. Here, we present performance results of NekBone, a benchmark of the Nek5000 code, implemented with optimized OpenACC directives and GPUDirect communications. Nek5000 is a computational
Gong, Jing +7 more
openaire +3 more sources

