Skip to content
UdemyHadoopPaid courseAll LevelsCertificate

PySpark & AWS: Master Big Data With PySpark and AWS (Udemy.com)

Mastering AWS & PySpark: Spark, PySpark, AWS, Spark Ecosystem, Hadoop, and Spark Applications [AWS, Hadoop, Pyspark]

Created by: AI Sciences

Last updated December 2025

icon
What you will learn

  • ● The introduction and importance of Big Data.
  • ● Practical explanation and live coding with PySpark.
  • ● Spark applications
  • ● Spark EcoSystem
  • ● Spark Architecture
  • ● Hadoop EcoSystem
  • ● Hadoop Architecture
  • ● PySpark RDDs
  • ● PySpark RDD transformations
  • ● PySpark RDD actions
  • ● PySpark DataFrames
  • ● PySpark DataFrames transformations
  • ● PySpark DataFrames actions
  • ● Collaborative filtering in PySpark
  • ● Spark Streaming
  • ● ETL Pipeline
  • ● CDC and Replication on Going

icon
Course Description

Comprehensive Course Description:

The hottest buzzwords in the Big Data analytics industry are Python and Apache Spark. PySpark supports the collaboration of Python and Apache Spark. In this course, you’ll start right from the basics and proceed to the advanced levels of data analysis. From cleaning data to building features and implementing machine learning (ML) models, you’ll learn how to execute end-to-end workflows using PySpark.


Right through the course, you’ll be using PySpark for performing data analysis. You’ll explore Spark RDDs, Dataframes, and a bit of Spark SQL queries. Also, you’ll explore the transformations and actions that can be performed on the data using Spark RDDs and dataframes. You’ll also explore the ecosystem of Spark and Hadoop and their underlying architecture. You’ll use the Databricks environment for running the Spark scripts and explore it as well.


Finally, you’ll have a taste of Spark with AWS cloud. You’ll see how we can leverage AWS storages, databases, computations, and how Spark can communicate with different AWS services and get its required data.   


How Is This Course Different? 

In this Learning by Doing course, every theoretical explanation is followed by practical implementation.   


The course ‘PySpark & AWS: Master Big Data With PySpark and AWS’ is crafted to reflect the most in-demand workplace skills. This course will help you understand all the essential concepts and methodologies with regards to PySpark. The course is:

• Easy to understand. 

• Expressive. 

• Exhaustive. 

• Practical with live coding. 

• Rich with the state of the art and latest knowledge of this field. 


As this course is a detailed compilation of all the basics, it will motivate you to make quick progress and experience much more than what you have learned. At the end of each concept, you will be assigned Homework/tasks/activities/quizzes along with solutions. This is to evaluate and promote your learning based on the previous concepts and methods you have learned. Most of these activities will be coding-based, as the aim is to get you up and running with implementations.   

High-quality video content, in-depth course material, evaluating questions, detailed course notes, and informative handouts are some of the perks of this course. You can approach our friendly team in case of any course-related queries, and we assure you of a fast response.   


The course tutorials are divided into 140+ brief videos. You’ll learn the concepts and methodologies of PySpark and AWS along with a lot of practical implementation. The total runtime of the HD videos is around 16 hours.


Why Should You Learn PySpark and AWS? 

PySpark is the Python library that makes the magic happen.   

PySpark is worth learning because of the huge demand for Spark professionals and the high salaries they command. The usage of PySpark in Big Data processing is increasing at a rapid pace compared to other Big Data tools.   

AWS, launched in 2006, is the fastest-growing public cloud. The right time to cash in on cloud computing skills—AWS skills, to be precise—is now.


Course Content:

The all-inclusive course consists of the following topics:

1. Introduction:

a. Why Big Data?

b. Applications of PySpark

c. Introduction to the Instructor

d. Introduction to the Course

e. Projects Overview

2. Introduction to Hadoop, Spark EcoSystems, and Architectures:

a. Hadoop EcoSystem

b. Spark EcoSystem

c. Hadoop Architecture

d. Spark Architecture

e. PySpark Databricks setup

f. PySpark local setup


3. Spark RDDs:

a. Introduction to PySpark RDDs

b. Understanding underlying Partitions

c. RDD transformations

d. RDD actions

e. Creating Spark RDD

f. Running Spark Code Locally

g. RDD Map (Lambda)

h. RDD Map (Simple Function)

i. RDD FlatMap

j. RDD Filter

k. RDD Distinct

l. RDD GroupByKey

m. RDD ReduceByKey

n. RDD (Count and CountByValue)

o. RDD (saveAsTextFile)

p. RDD (Partition)

q. Finding Average

r. Finding Min and Max

s. Mini project on student data set analysis

t. Total Marks by Male and Female Student

u. Total Passed and Failed Students

v. Total Enrollments per Course

w. Total Marks per Course

x. Average marks per Course

y. Finding Minimum and Maximum marks

z. Average Age of Male and Female Students

4. Spark DFs:

a. Introduction to PySpark DFs

b. Understanding underlying RDDs

c. DFs transformations

d. DFs actions

e. Creating Spark DFs

f. Spark Infer Schema

g. Spark Provide Schema

h. Create DF from RDD

i. Select DF Columns

j. Spark DF with Column

k. Spark DF with Column Renamed and Alias

l. Spark DF Filter rows

m. Spark DF (Count, Distinct, Duplicate)

n. Spark DF (sort, order By)

o. Spark DF (Group By)

p. Spark DF (UDFs)

q. Spark DF (DF to RDD)

r. Spark DF (Spark SQL)

s. Spark DF (Write DF)

t. Mini project on Employees data set analysis

u. Project Overview

v. Project (Count and Select)

w. Project (Group By)

x. Project (Group By, Aggregations, and Order By)

y. Project (Filtering)

z. Project (UDF and With Column)

aa. Project (Write)

5. Collaborative filtering:

a. Understanding collaborative filtering

b. Developing recommendation system using ALS model

c. Utility Matrix

d. Explicit and Implicit Ratings

e. Expected Results

f. Dataset

g. Joining Dataframes

h. Train and Test Data

i. ALS model

j. Hyperparameter tuning and cross-validation

k. Best model and evaluate predictions

l. Recommendations


6. Spark Streaming:

a. Understanding the difference between batch and streaming analysis.

b. Hands-on with spark streaming through word count example

c. Spark Streaming with RDD

d. Spark Streaming Context

e. Spark Streaming Reading Data

f. Spark Streaming Cluster Restart

g. Spark Streaming RDD Transformations

h. Spark Streaming DF

i. Spark Streaming Display

j. Spark Streaming DF Aggregations

7. ETL Pipeline

a. Understanding the ETL

b. ETL pipeline Flow

c. Data set

d. Extracting Data

e. Transforming Data

f. Loading data (Creating RDS)

g. Load data (Creating RDS)

h. RDS Networking

i. Downloading Postgres

icon
Udemy Discount

The discount is applied through our link. Open the course from here and Udemy's current promotional price is applied at checkout on most courses, no code to type.

Some courses are excluded from Udemy's promotions. If the price does not drop, clear your browser cookies and use the button again.

icon
Instructor Details

AI Sciences

Welcome to the epicenter of innovation, where a collective of visionaries, PhDs, and leading practitioners in Artificial Intelligence, Computer Science, Machine Learning, and Statistics unite. Our team hails from the tech titans - Amazon, Google, Facebook, Microsoft, KPMG, BCG, and IBM.

In our commitment to demystify the complex world of tech, we've crafted an extensive series of courses. Tailored primarily for beginners and newcomers, these courses are your gateway into the realms of Machine Learning, Statistics, Artificial Intelligence, and Data Science. We embarked on this journey with a simple goal: to make these advanced concepts accessible, minimizing theory and lengthy texts, allowing eager minds to dive straight into practice.

As our mission evolved, so did our offerings. We now present comprehensive courses that cater to a broader audience, ensuring everyone can navigate and master these fields with ease.

The impact of our courses has been nothing short of remarkable. We've empowered over 100,000 students, transforming them into masters of AI and Data Science. Join us, and be part of this journey of learning and empowerment, where your mastery of the future begins today.

icon
More courses by AI Sciences

$74.99

$64.99

$79.99

$10.99

icon
More Hadoop courses

$11.99

$12.99

$14.99

$94.99

$10.99

$39.99

icon
Reviews

4.4

3,347 ratings on Udemy

Select a bar to show only those reviews.Select the bar again to show every rating.

By Heloisa Colombo on 7/10/2024

Sections 2, 3, and 4 (Introduction to Ecosystem and Architecture, RDDs, and Dataframes) have good and interesting content. The Collaborative Filtering and Spark Streaming modules seem rushed and lack context. I was expecting more focus on integration with AWS, but that part was disappointing. Section 9, about Chatbot with Amazon Lex, doesn't seem to align well with the course content.

By JG Coding on 5/16/2024

I actually prefer this instructor including the errors he encounters rather than editing them out of his presentation. Tech services are ever-changing making producing tutorials a challenge. The screens are updated or improved between releases. Capabilities are enriched or made more streamline. The instructor does a good job of live problem-solving and clearing explaining and illustrating how to resolve these issues we encounter.

By Glorian Yapinus on 3/22/2024

In this course, you learn the basic knowledge of the Spark ecosystem, such as Spark DataFrames, RDD, Spark Streaming, and a short introduction to Spark MLlib when applied to a recommendation system problem (collaborative filtering). The instructor gave a detailed explanation geared towards beginners, which can be felt as a waste of time for a more experienced audience. In my opinion, the Amazon Lex part is less useful given the rise of OpenAI chatbots.

By Mohamed Ali SIDI YAHYA on 9/7/2023

I wouldn't recommend this course for beginners without foundational knowledge of Spark and Hadoop. The instructor occasionally glosses over certain concepts without in-depth explanations. Moreover, there are 5-6 hour segments solely on AWS. Some vital concepts are not elucidated well, while on the other hand, some obvious topics are stretched out for as long as 10 minutes. I believe this course would suit those with prior knowledge of Spark, Hadoop, and a basic understanding of AWS. For context, I took a precursor course titled "Spark Starter Kit," which provided an excellent introduction to Spark.

By Anonymized User on 7/1/2023

This is a great course on Pyspark and AWS. The first sections of the course even helped me to prepare and pass the Databricks Apache Spark 3.0 - Python certification. Every question I asked was answered in detail in a short time. Note: The Databrics section is a bit old, but the AWS sections are recorded not long ago!

By Alejandro Juárez Toribio on 10/2/2022

Es un curso bastante completo, respecto a las bases para poder implementar Pyspark con Databricks y AWS en proyectos propios y profesionales. Lo más relevante a mi parecer son los proyectos que se presentan al final de cada módulo, lo que permite reforzar todo lo aprendido durante las clases. Además después de solucionarlos por tu cuenta, al comparar con la solución del profesor, puedes ver algunas alternativas de solución a los problemas que plantea lo cual es muy formativo. Sin duda es un curso obligatorio para todos aquellos que quieran aprender estas herramientas tan demandadas hoy en día en todo el ambiente de Data Science.

By Naveen Thumu on 6/9/2022

very good course for beginner who has some knowledge of python. This Course will help budding developers. i recommend installing Hadoop File System in your local machine, and execute the scripts in your local machine instead of running on DataBricks. This will give good information about how jobs are executed as Pyspark has UI which provide this information http:// :4040 Happy learning

By Bilal Ali on 4/26/2022

This course help me form a basic understanding of Spark and how to use it to analyze large scale dataset. Besides fundamental knowledge of how to use, the lecturer also provide students with some deeper concept of how to optimize the performance of spark programming, which can be very useful in running code on large dataset.

By Anne Nguyen on 12/1/2021

The course gave a good introduction to using PySpark and AWS with detailed steps for tutorials. Installation steps were accurate together with clarifications on all dependencies. Some AWS steps were not exactly reproducible (understandably because AWS changes their interface all the time) and could use some update. The intro/ending music as well as the round-up for each lecture were rather repetitive and could be done away with. Overall it is a good course to quickly learn PySpark syntax and how it can be incorporated into AWS.

By Robert Vallance on 8/9/2021

Overall very good course. I liked the overall structure that started with the fundamental RDDs, then onto dataframes and streaming, etc. Explanations are excellent, breaking down some of the more complex ideas. Quick at responding to questions in the Q&A session too. I would recommend additional videos showing setup of packages such as MySQL Workbench in Linux and Mac as well as Windows.

Showing all 10 reviews on CourseDuck

Read all 3,347 reviews on Udemy

icon
Quality Score

No CourseDuck member has rated this course yet. Taken it? Give each part a thumbs up or down.

Content Quality
/
Video Quality
/
Qualified Instructor
/
Course Pace
/
Course Depth & Coverage
/

Overall Score : 88 / 100