Distributed Machine Learning with Apache Spark

Learn the underlying principles required to develop scalable machine learning pipelines and gain hands-on experience using Apache Spark.

Created by: Jon Bates

icon
Quality Score

Content Quality
/
Video Quality
/
Qualified Instructor
/
Course Pace
/
Course Depth & Coverage
/

Overall Score : 0 / 100

icon
Course Description

Machine learning aims to extract knowledge from data, relying on fundamental concepts in computer science, statistics, probability and optimization. Learning algorithms enable a wide range of applications, from everyday tasks such as product recommendations and spam filtering to bleeding edge applications like self-driving cars and personalized medicine. In the age of big data', with datasets rapidly growing in size and complexity and cloud computing becoming more pervasive, machine learning techniques are fast becoming a core component of large-scale data processing pipelines.
This statistics and data analysis course introduces the underlying statistical and algorithmic principles required to develop scalable real-world machine learning pipelines. We present an integrated view of data processing by highlighting the various components of these pipelines, including exploratory data analysis, feature extraction, supervised learning, and model evaluation. You will gain hands-on experience applying these principles using Spark, a cluster computing system well-suited for large-scale machine learning tasks, and its packages spark.ml and spark.mllib. You will implement distributed algorithms for fundamental statistical models (linear regression, logistic regression, principal component analysis) while tackling key problems from domains such as online advertising and cognitive neuroscience.

icon
Instructor Details

placeholder

Jon Bates is the Databricks program manager for MOOCs, a Spark instructor, and data science consultant. He is passionate about data science, computer science, and management science. A pragmatist at heart, he enjoys using tools from these fields to build a competitive edge in business. He spent nine years as a proprietary bond trader, where he built portfolio infrastructure and data analysis tools to maximize his and his team's trading returns. He has a B.S. in Management Science from MIT and an M.S. in Predictive Analytics from Northwestern University. Jon lives in Boulder, CO where he runs a consulting business focused on providing data science solutions and training.

icon
More hadoop courses

Apache Kafka Series - Kafka Cluster Setup & Administration

$11.99

Apache Kafka Series - Confluent Schema Registry & REST Proxy

$11.99

Apache Spark for Java Developers

$11.99

Apache Kafka - Real-time Stream Processing (Master Class)

$11.99

Apache Kafka Series - Kafka Monitoring & Operations

$11.99

Apache Kafka for absolute beginners

$11.99

icon
Reviews

0.0

0 total reviews

5 star 4 star 3 star 2 star 1 star
% Complete
% Complete
% Complete
% Complete
% Complete

No reviews yet. Be the first to review this course!