Results 11 to 20 of about 142,887 (185)
A Parallel Multiobjective PSO Weighted Average Clustering Algorithm Based on Apache Spark [PDF]
Multiobjective clustering algorithm using particle swarm optimization has been applied successfully in some applications. However, existing algorithms are implemented on a single machine and cannot be directly parallelized on a cluster, which makes it ...
Mingxing nie
exaly +4 more sources
Large-scale digital forensic investigation for Windows registry on Apache Spark [PDF]
In this study, we investigate large-scale digital forensic investigation on Apache Spark using a Windows registry. Because the Windows registry depends on the system on which it operates, the existing forensic methods on the Windows registry have been ...
Jun-Ha Lee, Hyuk-Yoon Kwon
doaj +3 more sources
Framing Apache Spark in life sciences [PDF]
Advances in high-throughput and digital technologies have required the adoption of big data for handling complex tasks in life sciences. However, the drift to big data led researchers to face technical and infrastructural challenges for storing, sharing,
Andrea Manconi +4 more
doaj +2 more sources
Big Data in metagenomics: Apache Spark vs MPI. [PDF]
The progress of next-generation sequencing has lead to the availability of massive data sets used by a wide range of applications in biology and medicine.
José M Abuín +4 more
doaj +2 more sources
DECA: scalable XHMM exome copy-number variant calling with ADAM and Apache Spark [PDF]
Background XHMM is a widely used tool for copy-number variant (CNV) discovery from whole exome sequencing data but can require hours to days to run for large cohorts.
Michael D. Linderman +3 more
doaj +2 more sources
DNA short read alignment on apache spark [PDF]
The evolution of technologies has unleashed a wealth of challenges by generating massive amount of data. Recently, biological data has increased exponentially, which has introduced several computational challenges.
Maryam AlJame, Imtiaz Ahmad
doaj +1 more source
Efficient processing of complex XSD using Hive and Spark [PDF]
The eXtensible Markup Language (XML) files are widely used by the industry due to their flexibility in representing numerous kinds of data. Multiple applications such as financial records, social networks, and mobile networks use complex XML schemas with
Diana Martinez-Mosquera +2 more
doaj +2 more sources
Efficient Group K Nearest-Neighbor Spatial Query Processing in Apache Spark
Aiming at the problem of spatial query processing in distributed computing systems, the design and implementation of new distributed spatial query algorithms is a current challenge.
Panagiotis Moutafis +3 more
doaj +1 more source
Comparative Study of Record Linkage Approaches for Big Data
Record linkage is a challenging task for Big Data. This paper, hence, attempts to shed light on record linkage approaches for Big Data by comparing three dimensions involving record linkage phases, dataset properties, and parallel processing approach ...
Randa MOHAMED +3 more
doaj +3 more sources
Apache Spark usage and deployment models for scientific computing [PDF]
This talk is about sharing our recent experiences in providing data analytics platform based on Apache Spark for High Energy Physics, CERN accelerator logging system and infrastructure monitoring.
Castro Diogo +4 more
doaj +1 more source

