Introduction to Apache Spark and PySpark Topics
- What is Distributed Data Processing?
- Why Single-Machine Tools Run Out
- Apache Spark in the Data Stack
- Spark vs Hadoop MapReduce
- PySpark and the JVM Bridge
- Installing PySpark Locally
- Spark Version Check
- The Driver and the Executors
- Cluster Managers Overview
- Local Mode for Learning
- Creating a SparkSession
- SparkContext vs SparkSession
- Reading the Spark UI
- Jobs, Stages and Tasks
- Notebook Setup for PySpark
- Your First DataFrame
- Stopping a Session Cleanly



















