Results 91 to 100 of about 926 (201)
Efficient Task Flow Parallel System for New Generation Sunway Processor [PDF]
China’s independently developed next-generation Sunway supercomputer features a more powerful memory system and higher computational density compared to its predecessor,the Sunway TaihuLight.Its primary programming model remains the bulk synchronous ...
FU You, DU Leiming, GAO Xiran, CHEN Li
doaj +1 more source
Hardware for Deep Learning Acceleration
In this review, various platforms for deep learning accelerations, such as central processing units, graphics processing units, neural processing units, compute‐in‐memory units, and neuromorphic event processors, are overviewed and they are compared with regard to the key performance metrics, such as operational throughput, power efficiency ...
Choongseok Song +3 more
wiley +1 more source
CAPS OpenACC Compilers: Performance and Portability [PDF]
The announcement late 2011 of the new OpenACC directive-based programming standard supported by CAPS, CRAY and PGI compilers has open up the door to more scientific applications that can be ported on many-core systems.
Bihan, Stéphane
core
AliceVision: GPU Programming for Depth Estimation Feature descriptor matching with Locality Sensitive Hashing optimized using OpenACC [PDF]
Implements an approximate nearest neighbor algorithm for many-dimensional points, called Locality Sensitive Hashing in C, then optimizes it using OpenACC.
Widerberg, Hilmar Røtnes
core
Nowadays, the ocean numerical models are gradually developing towards multi-physical process and high resolution, with the increment of measured ocean data and more in-depth research in ocean field.
Tao Liu +6 more
doaj +1 more source
Application of performance portability solutions for GPUs and many-core CPUs to track reconstruction kernels [PDF]
Next generation High-Energy Physics (HEP) experiments are presented with significant computational challenges, both in terms of data volume and processing power.
Kwok Ka Hei Martin +13 more
doaj +1 more source
Multi-level Parallelism with MPI and OpenACC for CFD Applications [PDF]
High-level parallel programming approaches, such as OpenACC, have recently become popular in complex fluid dynamics research since they are cross-platform and easy to implement.
McCall, Andrew James
core
Quantum ESPRESSO: One Further Step toward the Exascale. [PDF]
Carnimeo I +10 more
europepmc +1 more source
Toward Automatic Translation: From OpenACC to OpenMP 4 [PDF]
For the past few years, OpenACC has been the primary directive-based API for programming accelerator devices like GPUs. OpenMP 4.0 is now a competitor in this space, with support from different vendors.
Sultana, Nawrin
core

