Scalable Machine Learning on Big Data using Apache Spark

The rapid pace of innovation in Artificial Intelligence (AI) is creating enormous opportunity for transforming entire industries and our very existence. After competing this comprehensive 6 course Professional Certificate, you will get a practical understanding of Machine Learning and Deep Learning.You will master fundamental concepts of Machine Learning and Deep Learning, including supervised and unsupervised learning. You will utilize popular Machine Learning and Deep Learning libraries such as SciPy, ScikitLearn, Keras, PyTorch, and Tensorflow applied to industry problems involving object r

Created by: Romeo Kienzler

icon
Quality Score

Content Quality
/
Video Quality
/
Qualified Instructor
/
Course Pace
/
Course Depth & Coverage
/

Overall Score : 64 / 100

icon
Course Description

This course will empower you with the skills to scale data science and machine learning (ML) tasks on Big Data sets using Apache Spark. Most real world machine learning work involves very large data sets that go beyond the CPU, memory and storage limitations of a single computer. Apache Spark is an open source framework that leverages cluster computing and distributed storage to process extremely large data sets in an efficient and cost effective manner. Therefore an applied knowledge of working with Apache Spark is a great asset and potential differentiator for a Machine Learning engineer.After completing this course, you will be able to:- gain a practical understanding of Apache Spark, and apply it to solve machine learning problems involving both small and big data- understand how parallel code is written, capable of running on thousands of CPUs. - make use of large scale compute clusters to apply machine learning algorithms on Petabytes of data using Apache SparkML Pipelines. - eliminate out-of-memory errors generated by traditional machine learning frameworks when data doesn't fit in a computer's main memory- test thousands of different ML models in parallel to find the best performing one - a technique used by many successful Kagglers- (Optional) run SQL statements on very large data sets using Apache SparkSQL and the Apache Spark DataFrame API.Enrol now to learn the machine learning techniques for working with Big Data that have been successfully applied by companies like Alibaba, Apple, Amazon, Baidu, eBay, IBM, NASA, Samsung, SAP, TripAdvisor, Yahoo!, Zalando and many others.NOTE: You will practice running machine learning tasks hands-on on an Apache Spark cluster provided by IBM at no charge during the course which you can continue to use afterwards.Prerequisites:- basic python programming- basic machine learning (optional introduction videos are provided in this course as well)- basic SQL skills for optional contentThe following courses are recommended before taking this class (unless you already have the skills)https://www.coursera.org/learn/python-for-applied-data-science or similarhttps://www.coursera.org/learn/machine-learning-with-python or similarhttps://www.coursera.org/learn/sql-data-science for optional lectures

icon
Instructor Details

placeholder

Romeo Kienzler holds a M. Sc. (ETH) in Information Systems, informatics & Applied Statistics (Swiss Federal Institute of Technology). He has nearly two decades of experience in Software Enineering, Database Administration and Information Integration. Since 2012 he works as a Data Scientist for IBM. He published several works in the field with international publishers and on conferences. His current research focus is on massive parallel data processing architectures. Romeo also contributes to various open source projects.

icon
More courses by Romeo Kienzler

Fundamentals of Scalable Data Science

Free

Applied AI with DeepLearning

Free

Advanced Machine Learning and Signal Processing

Free

icon
More hadoop courses

Apache Kafka Series - Kafka Cluster Setup & Administration

$11.99

Apache Kafka Series - Confluent Schema Registry & REST Proxy

$11.99

Apache Spark for Java Developers

$11.99

Apache Kafka - Real-time Stream Processing (Master Class)

$11.99

Apache Kafka Series - Kafka Monitoring & Operations

$11.99

Apache Kafka for absolute beginners

$11.99

icon
Reviews

3.2

15 total reviews

5 star 4 star 3 star 2 star 1 star
% Complete
% Complete
% Complete
% Complete
% Complete

By Ruslan I M V on 9-Nov-19

Apache spark is great and powerful but the lectures are not clear and long.

By Lewis m on 12-Nov-19

So far the questions and quizes seem unrelated to machine learning. The videos are poorly set out, with breif explanations and the whole thing seems rushed.

By Benhur O J on 8-Oct-19

Too superficial. The python example codes are very cryptic and not very well commented. The programming videos are very difficult to follow because the instructor is literally reading the code instead of explaining it.

By Farrukh N A on 6-Nov-19

Course can be improved by focusing more on ML algorithms.... Explanation of GBT and Random Forest was not provided. But they were used.

By Jair J C C on 19-Oct-19

Very Good, but I think the course needs more challenging exams

By Suresh C on 2-Nov-19

There should be more details about Apache spark and some examples

By Yasser E H on 27-Oct-19

Really interesting content Unclear coding explanationsLimitations with the free access in IBM Watson Studio

By Gherbi H on 15-Nov-19

A very good course and will recommend it for anyone who has Apache Spark experience and wants to get an introduction to ML lib and machine learning in Apache Spark, the assignment submissions need some work but other than that a very good introductive course.

By Yuting K on 3-Oct-19

The quality of the videos could be better

By Ujjwal G on 11-Nov-19

For a intorductory course it is very good. Do not expect anything too advanced.

By Jay P on 29-Sep-19

Horrible

By Abdelrahman g e f on 16-Sep-19

the accent of the instructor was very hard to understand him during explanation but he was good instructor at all