Fundamentals of Scalable Data Science

As a coursera certified specialization completer you will have a proven deep understanding on massive parallel data processing, data exploration and visualization, and advanced machine learning & deep learning. You'll understand the mathematical foundations behind all machine learning & deep learning algorithms. You can apply knowledge in practical use cases, justify architectural decisions, understand the characteristics of different algorithms, frameworks & technologies & how they impact model performance & scalability.If you choose to take this specialization and earn th

Created by: Romeo Kienzler

icon
Quality Score

Content Quality
/
Video Quality
/
Qualified Instructor
/
Course Pace
/
Course Depth & Coverage
/

Overall Score : 78 / 100

icon
Course Description

Apache Spark is the de-facto standard for large scale data processing. This is the first course of a series of courses towards the IBM Advanced Data Science Specialization. We strongly believe that is is crucial for success to start learning a scalable data science platform since memory and CPU constraints are to most limiting factors when it comes to building advanced machine learning models.In this course we teach you the fundamentals of Apache Spark using python and pyspark. We'll introduce Apache Spark in the first two weeks and learn how to apply it to compute basic exploratory and data pre-processing tasks in the last two weeks. Through this exercise you'll also be introduced to the most fundamental statistical measures and data visualization technologies.This gives you enough knowledge to take over the role of a data engineer in any modern environment. But it gives you also the basis for advancing your career towards data science. Please have a look at the full specialization curriculum:https://www.coursera.org/specializations/advanced-data-science-ibmIf you choose to take this course and earn the Coursera course certificate, you will also earn an IBM digital badge. To find out more about IBM digital badges follow the link ibm.biz/badging.After completing this course, you will be able to:-Describe how basic statistical measures, are used to reveal patterns within the data -Recognize data characteristics, patterns, trends, deviations or inconsistencies, and potential outliers.-Identify useful techniques for working with big data such as dimension reduction and feature selection methods -Use advanced tools and charting libraries to:oimprove efficiency of analysis of big-data with partitioning and parallel analysis oVisualize the data in an number of 2D and 3D formats (Box Plot, Run Chart, Scatter Plot, Pareto Chart, and Multidimensional Scaling)For successful completion of the course, the following prerequisites are recommended: -Basic programming skills in python-Basic math-Basic SQL (you can get it easily from https://www.coursera.org/learn/sql-data-science if needed)In order to complete this course, the following technologies will be used:(These technologies are introduced in the course as necessary so no previous knowledge is required.)-Jupyter notebooks (brought to you by IBM Watson Studio for free)-ApacheSpark (brought to you by IBM Watson Studio for free)-PythonThis course takes four weeks, 4-6h per week

icon
Instructor Details

placeholder

Romeo Kienzler holds a M. Sc. (ETH) in Information Systems, informatics & Applied Statistics (Swiss Federal Institute of Technology). He has nearly two decades of experience in Software Enineering, Database Administration and Information Integration. Since 2012 he works as a Data Scientist for IBM. He published several works in the field with international publishers and on conferences. His current research focus is on massive parallel data processing architectures. Romeo also contributes to various open source projects.

icon
Reviews

3.9

161 total reviews

5 star 4 star 3 star 2 star 1 star
% Complete
% Complete
% Complete
% Complete
% Complete

By Piyapong B on 20-Nov-18

This course gives you nice experience with Apache Spark. There is lot of update going on interface which creates few problem but discussion forum helps you out. Good for beginners in Data Science who have basic knowledge of python and SQL.

By Jose A R N on 6-Feb-19

Sets you up well for working with Spark within the IBM Environment.

By Mukesh R on 6-Apr-18

Not so worth learning compare to the predecessor of this course. Should have included more assignment would have made the course very interesting.

By Neeraj on 14-Jul-19

Need more exercises related to wrangling data and manipulating SQL's with apache spark

By S V on 4-Jul-19

Nice but details are not discussed properly only taking names wasnt enough

By Felipe M M on 19-Sep-19

Videos are old. It feels like he had a bunch of material and put them together to create this course. For example: There are assignments that they give you the answer because the questions are not supposed to be there. He doesnt teach, instead, he reads a script. The assignments are not challenging and you dont feel like you learned. Horrible and painful.

By Markus W on 22-Sep-19

Romeo does a very good job of explaining things!However, the programming assignments are too easy to learn anything from.

By Eleni K on 10-Oct-19

I was really looking forward to this specialization but from the very first course I am really disappointed. The videos refer to various not updated information and then suddenly we are expected to do an assignment that was not at all explained in the course. I am not saying it is difficult, or not achievable but to be honest until now (week 2) it feels mostly like a waste of time.. Really sorry for this review.

By Csaba P O on 9-Sep-19

The content was OK, but I have expected more. Probably it was too basic for me. I would have been happy to see some more real life examples, like when to use the different statistics to solve real problems, not only the theoretical ones.

By Nikhilanj P on 6-Sep-19

Too many legacy issues. Would be better to start a new course altogether and maintain same syntax,etc.

By Tony H on 4-Nov-19

I felt that, for a course labelled as 'Advanced', there were too many trivial questions in the quizzes and too much hand-holding in the programming assignments. That being said I did enjoy the course and learned quite a lot and look forward to the next one in the specialisation.

By Cesar R on 6-Jul-19

Very basic lessons. Definitely what you would expect from an Advanced course.