icon
Quality Score

Content Quality
/
Video Quality
/
Qualified Instructor
/
Course Pace
/
Course Depth & Coverage
/

Overall Score : 0 / 100

icon
Course Description

Apache Spark is one of the most widely used and supported open-source tools for machine learning and big data. In this course, discover how to work with this powerful platform for machine learning. Instructor Dan Sullivan discusses MLlib - the Spark machine learning library - which provides tools for data scientists and analysts who would rather find solutions to business problems than code, test, and maintain their own machine learning libraries. He shows how to use DataFrames to organize data structure, and he covers data preparation and the most commonly used types of machine learning algorithms: clustering, classification, regression, and recommendations. By the end of the course, you will have experience loading data into Spark, preprocessing data as needed to apply MLlib algorithms, and applying those algorithms to a variety of machine learning problems.

icon
Instructor Details

Dan Sullivan

Dan Sullivan, PhD, is an enterprise architect and big data expert.

Dan specializes in data architecture, analytics, data mining, statistics, data modeling, big data, and cloud computing. In addition, he holds a PhD in genetics, bioinformatics, and computational biology. Dan works regularly with Spark, Oracle, NoSQL, MongoDB, Redis, R, and Python. He has extensive writing experience in topics including cloud computing, big data, Hadoop, and security.

icon
More courses by Dan Sullivan

Advanced SQL for Data Science: Time Series

Free

Introduction to Spark SQL and DataFrames

Free

Advanced SQL for Data Scientists

Free

icon
Reviews

0.0

0 total reviews

5 star 4 star 3 star 2 star 1 star
% Complete
% Complete
% Complete
% Complete
% Complete

No reviews yet. Be the first to review this course!