Results 21 to 30 of about 1,860 (195)
Implementing Apache Spark jobs execution and Apache Spark cluster creation for Openstack Sahara[1]
In this paper the problem of creating virtual clusters in clouds for big data analysis with Apache Hadoop and Apache Spark is discussed. Existing methods for Apache Spark clusters creation are described in this work.
A. . Aleksiyants +4 more
doaj +1 more source
Optimization Techniques for a Distributed In-Memory Computing Platform by Leveraging SSD
In this paper, we present several optimization strategies that can improve the overall performance of the distributed in-memory computing system, “Apache Spark”.
June Choi +3 more
doaj +1 more source
Apache Spark merupakan platform yang dapat digunakan untuk memproses data dengan ukuran data yang relatif besar (big data) dengan kemampuan untuk membagi data tersebut ke masing-masing cluster yang telah ditentukan konsep ini disebut dengan parallel ...
Aminudin Aminudin, Eko Budi Cahyono
doaj +1 more source
Deploying Apache Spark virtual clusters in cloud environments using orchestration technologies
Apache Spark is a framework providing fast computations on Big Data using MapReduce model. With cloud environments Big Data processing becomes more flexible since they allow to create virtual clusters on-demand. One of the most powerful open-source cloud
O. . Borisenko +2 more
doaj +1 more source
A distributed computing model for big data anonymization in the networks.
Recently big data and its applications had sharp growth in various fields such as IoT, bioinformatics, eCommerce, and social media. The huge volume of data incurred enormous challenges to the architecture, infrastructure, and computing capacity of IT ...
Farough Ashkouti, Keyhan Khamforoosh
doaj +1 more source
ReForeSt: Random Forests in Apache Spark [PDF]
Random Forests (RF) of tree classifiers are a popular ensemble method for classification. RF are usually preferred with respect to other classification techniques because of their limited hyperparameter sensitivity, high numerical robustness, native capacity of dealing with numerical and categorical features, and effectiveness in many real world ...
Lulli A., Oneto L., Anguita D.
openaire +1 more source
Laurelin: Java-native ROOT I/O for Apache Spark [PDF]
Apache Spark[1] is one of the predominant frameworks in the big data space, providing a fully-functional query processing engine, vendor support for hardware accelerators, and performant integrations with scientific computing libraries. One difficulty in
Melo Andrew, Shadura Oksana
doaj +1 more source
Privacy-Preserving Machine Learning on Apache Spark
The adoption of third-party machine learning (ML) cloud services is highly dependent on the security guarantees and the performance penalty they incur on workloads for model training and inference.
Claudia V. Brito +4 more
doaj +1 more source
Performance Analysis of the Distributed Support Vector Machine Algorithm Using Spark for Predicting Flight Delays [PDF]
In big data analysis requires powerful machine learning frameworks, strategies, and environments to analyze data at scale. Therefore, Apache Spark is used as a cluster computing framework to process big data in parallel and can run on multiple clusters ...
Khotimah Husnul +4 more
doaj +1 more source
Evaluasi Kinerja MLLIB APACHE SPARK pada Klasifikasi Berita Palsu dalam Bahasa Indonesia
Machine learning digunakan untuk menganalisis, mengklasifikasikan, atau memprediksi data. Untuk melakukan tugas dari machine learning diperlukan alat bantu dengan kinerja serta lingkungan yang kuat demi mendapatkan akurasi dan efisiensi waktu yang baik.
Antonius Angga Kurniawan +1 more
doaj +1 more source

