Results 11 to 20 of about 1,860 (195)
Efficient processing of complex XSD using Hive and Spark [PDF]
The eXtensible Markup Language (XML) files are widely used by the industry due to their flexibility in representing numerous kinds of data. Multiple applications such as financial records, social networks, and mobile networks use complex XML schemas with
Diana Martinez-Mosquera +2 more
doaj +2 more sources
SparkRA: Enabling Big Data Scalability for the GATK RNA-seq Pipeline with Apache Spark [PDF]
Zaid Al-Ars +2 more
exaly +2 more sources
Efficient Group K Nearest-Neighbor Spatial Query Processing in Apache Spark
Aiming at the problem of spatial query processing in distributed computing systems, the design and implementation of new distributed spatial query algorithms is a current challenge.
Panagiotis Moutafis +3 more
doaj +1 more source
Comparative Study of Record Linkage Approaches for Big Data
Record linkage is a challenging task for Big Data. This paper, hence, attempts to shed light on record linkage approaches for Big Data by comparing three dimensions involving record linkage phases, dataset properties, and parallel processing approach ...
Randa MOHAMED +3 more
doaj +3 more sources
Apache Spark usage and deployment models for scientific computing [PDF]
This talk is about sharing our recent experiences in providing data analytics platform based on Apache Spark for High Energy Physics, CERN accelerator logging system and infrastructure monitoring.
Castro Diogo +4 more
doaj +1 more source
Apache Spark ile Makine Öğrenmesi Destekli Diyabet Rahatsızlığı Tahmini
Diyabet rahatsızlığı, insan vücudunun organlarını etkileyen kritik sağlık sorunlarından biridir. Bu nedenle, diyabet, 21. yüzyılda küresel bir sağlık sorunu olarak kabul edilmektedir.
Emre Yıldırım, Ali Çalhan
doaj +1 more source
Alchemist: An Apache Spark ⇔ MPI interface [PDF]
SummaryThe Apache Spark framework for distributed computation is popular in the data analytics community due to its ease of use, but its MapReduce‐style programming model can incur significant overheads when performing computations that do not map directly onto this model. One way to mitigate these costs is to off‐load computations onto MPI codes.
Alex Gittens +8 more
openaire +2 more sources
CHANGE DETECTION OF MOBILE LIDAR DATA USING CLOUD COMPUTING [PDF]
Change detection has long been a challenging problem although a lot of research has been conducted in different fields such as remote sensing and photogrammetry, computer vision, and robotics. In this paper, we blend voxel grid and Apache Spark together
K. Liu, J. Boehm, C. Alis
doaj +1 more source
Performance Evaluation of Query Plan Recommendation with Apache Hadoop and Apache Spark
Access plan recommendation is a query optimization approach that executes new queries using prior created query execution plans (QEPs). The query optimizer divides the query space into clusters in the mentioned method.
Elham Azhir +3 more
doaj +1 more source
Дослідження продуктивності кластера Apache Spark на платформі Azure для методів машинного навчання
Розглянуто та досліджено питання підвищення продуктивності застосування моделей та методів задач машинного навчання з використанням Apache Spark Azure HDInsight.
С.В. Мінухін
doaj +1 more source

