Skip to content
UdemyHadoopPaid courseIntermediate levelCertificate

[LEGACY–SUPPORT END] Spark SQL & Hadoop (For Data Science) (Udemy.com)

Learn HDFS commands, Hadoop, Spark SQL, SQL Queries, ETL & Data Analysis| Spark Hadoop Cluster VM | Fully Solved Qs

Created by: Matthew Barr

Last updated September 2024

icon
What you will learn

  • Students will get hands-on experience working in a Spark Hadoop environment that’s free and downloadable as part of this course.
  • Students will have opportunities solve Data Engineering and Data Analysis Problems using Spark on a Hadoop cluster in the sandbox environment that comes as part
  • Issuing HDFS commands.
  • Converting a set of data values in a given format stored in HDFS into new data values or a new data format and writing them into HDFS.
  • Loading data from HDFS for use in Spark applications & writing the results back into HDFS using Spark.
  • Reading and writing files in a variety of file formats.
  • Performing standard extract, transform, load (ETL) processes on data using the Spark API.
  • Using metastore tables as an input source or an output sink for Spark applications.
  • Applying the understanding of the fundamentals of querying datasets in Spark.
  • Filtering data using Spark.

icon
Course Description

*Important Notice*

This course has been retired and is no longer receiving support. Originally designed to help students pass the now-retired Cloudera Certification exams, the material remains useful for those wanting to practice their skills on Spark and Hadoop clusters. However, its primary focus was certification preparation, which many students successfully completed.


Apache Spark is currently one of the most popular systems for processing big data.


Apache Hadoop continues to be used by many organizations that look to store data locally on premises. Hadoop allows these organisations to efficiently store big datasets ranging in size from gigabytes to petabytes.


As the number of vacancies for data science, big data analysis and data engineering roles continue to grow, so too will the demand for individuals that possess knowledge of Spark and Hadoop technologies to fill these vacancies.


This course has been designed specifically for data scientists, big data analysts and data engineers looking to leverage the power of Hadoop and Apache Spark to make sense of big data.


This course will help those individuals that are looking to interactively analyse big data or to begin writing production applications to prepare data for further analysis using Spark SQL in a Hadoop environment.


The course is also well suited for university students and recent graduates that are keen to gain exposure to Spark & Hadoop or anyone who simply wants to apply their SQL skills in a big data environment using Spark-SQL.


This course has been designed to be concise and to provide students with a necessary and sufficient amount of theory, enough for them to be able to use Hadoop & Spark without getting bogged down in too much theory about older low-level APIs such as RDDs.


On solving the questions contained in this course students will begin to develop those skills & the confidence needed to handle real world scenarios that come their way in a production environment.


(a) There are just under 30 problems in this course. These cover hdfs commands, basic data engineering tasks and data analysis.

(b) Fully worked out solutions to all the problems.

(c) Also included is the Verulam Blue virtual machine which is an environment that has a spark Hadoop cluster already installed so that you can practice working on the problems.


  • The VM contains a Spark Hadoop environment which allows students to read and write data to & from the Hadoop file system as well as to store metastore tables on the Hive metastore.

  • All the datasets students will need for the problems are already loaded onto HDFS, so there is no need for students to do any extra work.

  • The VM also has Apache Zeppelin installed. This is a notebook specific to Spark and is similar to Python’s Jupyter notebook.


This course will allow students to get hands-on experience working in a Spark Hadoop environment as they practice:


  • Converting a set of data values in a given format stored in HDFS into new data values or a new data format and writing them into HDFS.

  • Loading data from HDFS for use in Spark applications & writing the results back into HDFS using Spark.

  • Reading and writing files in a variety of file formats.

  • Performing standard extract, transform, load (ETL) processes on data using the Spark API.

  • Using metastore tables as an input source or an output sink for Spark applications.

  • Applying the understanding of the fundamentals of querying datasets in Spark.

  • Filtering data using Spark.

  • Writing queries that calculate aggregate statistics.

  • Joining disparate datasets using Spark.

  • Producing ranked or sorted data.

icon
Udemy Discount

The discount is applied through our link. Open the course from here and Udemy's current promotional price is applied at checkout on most courses, no code to type.

Some courses are excluded from Udemy's promotions. If the price does not drop, clear your browser cookies and use the button again.

icon
Instructor Details

Matthew Barr

I’m Matthew Barr, a data engineer and founder of Verulam Blue Mint, a UK-based SaaS platform that helps learners build real, portfolio-ready projects using SQL, dbt & PySpark for roles in data engineering, analytics, and data science. At Verulam Blue Mint, I focus on giving you the kind of realistic, messy-data experience you’d get in an actual data team—minus the office politics.

I started my career in data science and, over time, moved into roles with a greater focus on data engineering. I’m a certified data engineer across major cloud platforms: Azure, Fabric, AWS, and Databricks. I’ve spent years building the kind of end-to-end data-cleaning pipelines and KPI workflows that you’ll practise in my courses.

Before moving fully into data, I spent nearly a decade in financial services as an actuarial analyst in London. That background means I care a lot about trustworthy numbers, repeatable logic, and being able to explain results clearly to non-technical stakeholders.

Academically, I hold two master’s degrees from University College London: one in Data Science and one in Mathematics.

Through Verulam Blue Mint and my Udemy courses, my goal is simple: to help you bridge the gap between “I know some SQL/PySpark” and “I can run a serious data project that I’m proud to show on my CV, GitHub, and LinkedIn.”

icon
More courses by Matthew Barr

$54.99

$49.99

$64.99

icon
More Hadoop courses

$84.99

$15.00

$94.99

$84.99

$39.99

$89.99

icon
Reviews

4.4

96 ratings on Udemy

Select a bar to show only those reviews.Select the bar again to show every rating.

By Isabella Rossi on 2/3/2026

Really solid hands-on Spark SQL + Hadoop practice: the included VM gets you running fast, and the problems are great for practice. It’s retired and a bit certification-era in places, but it’s still a good way to build confidence with HDFS and Spark workflows.

By Luke Wilson on 12/31/2025

Great course even if it was meant for old Cloudera exams. But, had to get an old version of the Oracle virtual box to run the cluster on the Verulam Blue VM.

By Zofia Nowak on 12/27/2025

This course might be “legacy” but the content is still valuable. Got a great introduction to Spark and became familiar with hadoop. Also got a platform I could practice on a Spark cluster without the headache of setup. Great course!!!

By Georgios Varvaropoulos on 11/3/2023

Good flow (topics/Voice Over), plenty practical examples, significant coverage of topics, presentation could improved (long text phrases in most of the sections)

By Julián Esteban Márquez Quintero on 10/21/2023

A veces mostraba muy rápido los ejemplos. Quiero que haya más ejemplos para futuros cursos

By Veena Rathi on 3/6/2023

the course was good but no one responded to my questions

By Rajesh Kaushal on 3/10/2022

Got some knowledge about Big data. Looking forward to complete the course.

By Haindavi Kasireddy on 1/19/2022

Awesome course. It covers pretty much everything.

By Jeffrey Rhod Pestaño on 12/30/2021

i learn about more commands about HDFS and Spark SQL thank u for this lesson

By Cheikh Badiane on 12/22/2021

I love the way it explains even a complex topic becomes easy to understand I really recommend this course which more detail you on big data and its use immerses you in the heart of the matter through practice

Showing all 10 reviews on CourseDuck

Read all 96 reviews on Udemy

icon
Quality Score

No CourseDuck member has rated this course yet. Taken it? Give each part a thumbs up or down.

Content Quality
/
Video Quality
/
Qualified Instructor
/
Course Pace
/
Course Depth & Coverage
/

Overall Score : 88 / 100