New in Big Data: Hive, HiveMall, AWS Lambda, Solr, Kibana (Udemy.com)
Big Data ETL, Machine Learning and Data Visualization. Concise hands-on course with full code examples. Learn to excel!
Created by: Elena Akhmatova
Produced in 2017
What you will learn
- craft a solution to your BigData tasks using building blocks shown in the course
- create deliverables for your work in a form of an Amazon microservice, Search Engine web service, or a Dashboard
- learn new technologies: AWS Lambda, Hivemall Machine Learning Library, HyperLogLog cardinality estimation technique, connecting Apache Hive to Solr and Hue, writing custom Hive UDFs, and a few other things
Quality Score
Overall Score : 96 / 100
Course Description
This course works you through the full Big Data process:
- Data Input
- ETL
- Predictive Modelling using Machine Learning
- Data Visualization
- Deployment to AWS using AWS Lambda and Amazon EMR bundle
Apache Hive is an easy SQL based tool that allows to process large amounts of data on Hadoop fast. Hive gained popularity immediately after Hadoop MapReduce became widely used as it allows to work with data by means of SQL queries. It is used by many organisations to process their data. This course shows a number of interesting Hive queries and explains what Hive UDFs are.
Apache HiveMall is a Machine Learning library of tomorrow. Like Hive it allows to use complex machine learning algorithms knowing SQL only. No need to code, compile and debug! It is really easy to use for programmers and non-programmers. Apache HiveMall Machine Learning library implements many useful Machine Learning algorithms (Supervised classification, LDA, RandomForest, etc.) using Hive UDFs. This course focuses on Text Classification when presenting HiveMall.
Hive + HiveMall is no less (or maybe even more) attractive and efficient than Spark + Spark MLib. Also, as HiveQL is more or less SQL. Knowing SQL and knowing only SQL will allow many non-developers to enter BigData world.
AWS Lambda is a must to know now. I show how to use it with Java to make it suitable to be a part of a BigData pipeline. AWS Lambda + Amazon EMR + Hive combination is also explained.
Solr and Hue is a search engine and visualisation dashboard combination. ElasticSearch and Kibana is another such combination. Both technologies use the same idea: use connectors to push data from Hive or Spark directly to Solr or ElasticSearch. Hue and Kibana use properties and inner data representations of their corresponding search engines to display data on a dashboard. This course shows how to integrate Hive with both technologies.
Instead of being comprehensive this course assumes a bit of prior knowledge of the topic. It teaches by presenting solutions for the problems that occurred repeatedly during the time i worked on different BigData projects. It shows how mastering small things gives you an ability to create a simple solution to almost every problem from concept to delivery.
We start with importing data to Apache Hive correctly, and slowly progress to an ability to quickly deliver results of your work as an AWS service, a Search Engine service, or a Hue dashboard.
The course shows data processing with Hive (also teaching how to write User Defined Functions for Hive of different levels of complexity: UDF, GenericUDF, UDAF and UDTF), it shows an application of Machine Learning to Text Classification using HiveMall, and then exporting data from Hive to Solr & Hue or ElasticSearch & Kibana. You will also learn how to write an AWS Lambda that runs Hive.
All together that gives you an ability to build a simple data processing pipeline. A data pipeline that is simple, robust and ready to be delivered and used in no time. Who this course is for:
- Anyone who has done at least one introductory BigData course.
- People who think they know it all will most probably learn something new too.
Instructor Details
- 4.8 Rating
4 Reviews
Elena Akhmatova
Elena works in the field of Natural Language Processing. She graduated with a degree from Saint-Petersburg State University in Russia first and then acquired PhD from Macquarie University in Sydney, Australia, where she works currently. Now she applies theoretical concepts developed in the field of Natural Language Processing to solve business problems of different big and small enterprises.
As an early adopter of BigData tools and concepts she finds existing BigData frameworks to be attractive means of working with data. She started using such tools and advising other people to adopt BigData concepts way before Hadoop, Spark and other related technologies became "must to know" tools for many IT professionals.
Sharing knowledge is something Elena enjoys doing. She believes that sharing knowledge enriches her as much as other people.
More elastic search courses
Apache Kafka Series - Learn Apache Kafka for Beginners v2 (2021)
4.8 (388 Reviews)
Provider: Udemy
Time: 7.5h
$11.99
Spring Boot Microservices and Spring Cloud (2021)
4.7 (407 Reviews)
Provider: Udemy
Time: 16h
$11.99
The Flask Mega-Tutorial (Python Web Development) (2018)
4.7 (114 Reviews)
Provider: Udemy
Time: 11.5h
$11.99
Part 2: AWS Certified Solutions Architect (and CD,SO) - 2021 (2021)
4.6 (61 Reviews)
Provider: Udemy
Time: 11h
$11.99
AWS Certified Data Analytics Specialty - Hands On! (2021)
4.5 (449 Reviews)
Provider: Udemy
Time: 12.5h
$11.99








