This course provides an introduction to big data concepts and hands-on experience with Apache Spark, a powerful distributed data processing framework. Participants will learn how to process large-scale datasets efficiently using Spark’s core components, including Spark SQL, DataFrames, and basic data transformations.
Through practical labs and examples, learners will understand how Spark enables scalable data processing and analytics for real-world applications such as ETL pipelines, data analysis, and machine learning workflows.
Duration:
2 Days
Course Code: BDT152
Learning Objectives:
After this course, you will be able to:
Data engineers, data analysts, developers, and IT professionals interested in big data processing and analytics
Basic knowledge of Python or Scala, SQL, and fundamental data processing concepts; familiarity with distributed systems is a plus
Module 1: Introduction to Big Data
Module 2: Introduction to Apache Spark
Module 3: Working with Spark Core and RDDs
Module 4: Spark DataFrames and Spark SQL
Module 5: Data Processing and ETL Pipelines
Module 6: Introduction to Spark MLlib
Module 7: Performance Optimization and Best Practices