Results 1 to 10 of about 1,820 (155)

Framing Apache Spark in life sciences [PDF]

open access: yesHeliyon, 2023
Advances in high-throughput and digital technologies have required the adoption of big data for handling complex tasks in life sciences. However, the drift to big data led researchers to face technical and infrastructural challenges for storing, sharing,
Andrea Manconi   +4 more
doaj   +5 more sources

A Parallel Multiobjective PSO Weighted Average Clustering Algorithm Based on Apache Spark [PDF]

open access: yesEntropy, 2023
Multiobjective clustering algorithm using particle swarm optimization has been applied successfully in some applications. However, existing algorithms are implemented on a single machine and cannot be directly parallelized on a cluster, which makes it ...
Mingxing Nie
exaly   +4 more sources

Large-scale digital forensic investigation for Windows registry on Apache Spark [PDF]

open access: yesPLoS ONE, 2022
In this study, we investigate large-scale digital forensic investigation on Apache Spark using a Windows registry. Because the Windows registry depends on the system on which it operates, the existing forensic methods on the Windows registry have been ...
Jun-Ha Lee, Hyuk-Yoon Kwon
doaj   +3 more sources

Bioinformatics applications on Apache Spark. [PDF]

open access: yesGigascience, 2018
With the rapid development of next-generation sequencing technology, ever-increasing quantities of genomic data pose a tremendous challenge to data processing. Therefore, there is an urgent need for highly scalable and powerful computational systems. Among the state-of-the-art parallel computing platforms, Apache Spark is a fast, general-purpose, in ...
Guo R, Zhao Y, Zou Q, Fang X, Peng S.
europepmc   +4 more sources

Big Data in metagenomics: Apache Spark vs MPI. [PDF]

open access: yesPLoS ONE, 2020
The progress of next-generation sequencing has lead to the availability of massive data sets used by a wide range of applications in biology and medicine.
José M Abuín   +4 more
doaj   +2 more sources

Big data analytics on Apache Spark [PDF]

open access: yesInternational Journal of Data Science and Analytics, 2016
Apache Spark has emerged as the de facto framework for big data analytics with its advanced in-memory programming model and upper-level libraries for scalable machine learning, graph analysis, streaming and structured data processing. It is a general-purpose cluster computing framework with language-integrated APIs in Scala, Java, Python and R.
Xiaojun Chen   +2 more
exaly   +2 more sources

DECA: scalable XHMM exome copy-number variant calling with ADAM and Apache Spark [PDF]

open access: yesBMC Bioinformatics, 2019
Background XHMM is a widely used tool for copy-number variant (CNV) discovery from whole exome sequencing data but can require hours to days to run for large cohorts.
Michael D. Linderman   +3 more
doaj   +2 more sources

TRANSMUT‐Spark: Transformation mutation for Apache Spark [PDF]

open access: yesSoftware Testing, Verification and Reliability, 2022
SummaryThis paper proposesTRANSMUT‐Sparkfor automating mutation testing of big data processing code within Spark programs. Apache Spark is an engine for big data analytics/processing that hides the inherent complexity of parallel big data programming. Nonetheless, programmers must cleverly combine Spark built‐in functions within programs and guide the ...
João Batista de Souza Neto   +3 more
openaire   +3 more sources

DNA short read alignment on apache spark [PDF]

open access: yesApplied Computing and Informatics, 2023
The evolution of technologies has unleashed a wealth of challenges by generating massive amount of data. Recently, biological data has increased exponentially, which has introduced several computational challenges.
Maryam AlJame, Imtiaz Ahmad
doaj   +1 more source

Home - About - Disclaimer - Privacy