Results 21 to 30 of about 1,860 (195)

Implementing Apache Spark jobs execution and Apache Spark cluster creation for Openstack Sahara[1]

open access: yesТруды Института системного программирования РАН, 2018
In this paper the problem of creating virtual clusters in clouds for big data analysis with Apache Hadoop and Apache Spark is discussed. Existing methods for Apache Spark clusters creation are described in this work.
A. . Aleksiyants   +4 more
doaj   +1 more source

Optimization Techniques for a Distributed In-Memory Computing Platform by Leveraging SSD

open access: yesApplied Sciences, 2021
In this paper, we present several optimization strategies that can improve the overall performance of the distributed in-memory computing system, “Apache Spark”.
June Choi   +3 more
doaj   +1 more source

Pengukuran Performa Apache Spark dengan Library H2O Menggunakan Benchmark Hibench Berbasis Cloud Computing

open access: yesJurnal Teknologi Informasi dan Ilmu Komputer, 2019
Apache Spark merupakan platform yang dapat digunakan untuk memproses data dengan ukuran data yang relatif  besar (big data) dengan kemampuan untuk membagi data tersebut ke masing-masing cluster yang telah ditentukan konsep ini disebut dengan parallel ...
Aminudin Aminudin, Eko Budi Cahyono
doaj   +1 more source

Deploying Apache Spark virtual clusters in cloud environments using orchestration technologies

open access: yesТруды Института системного программирования РАН, 2018
Apache Spark is a framework providing fast computations on Big Data using MapReduce model. With cloud environments Big Data processing becomes more flexible since they allow to create virtual clusters on-demand. One of the most powerful open-source cloud
O. . Borisenko   +2 more
doaj   +1 more source

A distributed computing model for big data anonymization in the networks.

open access: yesPLoS ONE, 2023
Recently big data and its applications had sharp growth in various fields such as IoT, bioinformatics, eCommerce, and social media. The huge volume of data incurred enormous challenges to the architecture, infrastructure, and computing capacity of IT ...
Farough Ashkouti, Keyhan Khamforoosh
doaj   +1 more source

ReForeSt: Random Forests in Apache Spark [PDF]

open access: yes, 2017
Random Forests (RF) of tree classifiers are a popular ensemble method for classification. RF are usually preferred with respect to other classification techniques because of their limited hyperparameter sensitivity, high numerical robustness, native capacity of dealing with numerical and categorical features, and effectiveness in many real world ...
Lulli A., Oneto L., Anguita D.
openaire   +1 more source

Laurelin: Java-native ROOT I/O for Apache Spark [PDF]

open access: yesEPJ Web of Conferences, 2021
Apache Spark[1] is one of the predominant frameworks in the big data space, providing a fully-functional query processing engine, vendor support for hardware accelerators, and performant integrations with scientific computing libraries. One difficulty in
Melo Andrew, Shadura Oksana
doaj   +1 more source

Privacy-Preserving Machine Learning on Apache Spark

open access: yesIEEE Access, 2023
The adoption of third-party machine learning (ML) cloud services is highly dependent on the security guarantees and the performance penalty they incur on workloads for model training and inference.
Claudia V. Brito   +4 more
doaj   +1 more source

Performance Analysis of the Distributed Support Vector Machine Algorithm Using Spark for Predicting Flight Delays [PDF]

open access: yesE3S Web of Conferences, 2023
In big data analysis requires powerful machine learning frameworks, strategies, and environments to analyze data at scale. Therefore, Apache Spark is used as a cluster computing framework to process big data in parallel and can run on multiple clusters ...
Khotimah Husnul   +4 more
doaj   +1 more source

Evaluasi Kinerja MLLIB APACHE SPARK pada Klasifikasi Berita Palsu dalam Bahasa Indonesia

open access: yesJurnal Teknologi Informasi dan Ilmu Komputer, 2022
Machine learning digunakan untuk menganalisis, mengklasifikasikan, atau memprediksi data. Untuk melakukan tugas dari machine learning diperlukan alat bantu dengan kinerja serta lingkungan yang kuat demi mendapatkan akurasi dan efisiensi waktu yang baik.
Antonius Angga Kurniawan   +1 more
doaj   +1 more source

Home - About - Disclaimer - Privacy