Skip to main content

Introduction to Apache Spark™

This beginner-friendly course covers the fundamentals of Apache Spark for large-scale data processing. You will explore Spark’s distributed architecture, master the DataFrame API, and learn to read, write, and process data using Python. Through hands-on exercises, you will build the skills needed to execute Spark transformations and actions efficiently.


Languages Available: English | 日本語 | 한국어

Skill Level
Associate
Duration
4h
Prerequisites
The content was developed for participants with these skills/knowledge/abilities:

• Basic programming knowledge

• Familiarity with Python

• Basic understanding of SQL queries (SELECT, JOIN, GROUP BY)

• Familiarity with data processing concepts

• No prior Spark or Databricks experience required

Outline

1. Apache Spark Runtime Architecture

• What is Apache Spark

• Spark Runtime Architecture

• Demo: Exploring Spark Architecture in Databricks


2. Spark DataFrames and SQL

• Introduction to DataFrames

• Reading and Writing Data

• Demo: Reading and Writing Data with DataFrames


3. Distributed Systems Programming Fundamentals

• Distributed Systems Programming Fundamentals


4. ETL with the DataFrame API

• Basic ETL Operations with the DataFrame API

• Demo: Flight Data ETL with the DataFrame API

• Lab: Analyzing Transaction Data with DataFrames

Upcoming Public Classes

Date
Time
Your Local Time
Language
Price
Nov 24
01 PM - 05 PM (Australia/Sydney)
-
English
$750.00
Nov 24
09 AM - 01 PM (Europe/Paris)
-
English
$750.00
Nov 24
09 AM - 01 PM (America/Los_Angeles)
-
English
$750.00
Dec 15
09 AM - 01 PM (Asia/Kolkata)
-
English
$750.00
Dec 15
01 PM - 05 PM (Europe/Paris)
-
English
$750.00
Dec 15
09 AM - 01 PM (America/New_York)
-
English
$750.00
Jan 26
09 AM - 01 PM (Asia/Singapore)
-
English
$750.00
Jan 26
09 AM - 01 PM (Europe/Paris)
-
English
$750.00
Jan 26
01 PM - 05 PM (America/New_York)
-
English
$750.00

Public Class Registration

If your company has purchased success credits or has a learning subscription, please fill out the Training Request form. Otherwise, you can register below.

Private Class Request

If your company is interested in private training, please submit a request.

See all our registration options

Registration options

Databricks has a delivery method for wherever you are on your learning journey

Runtime

Self-Paced

Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos

Register now

Instructors

Instructor-Led

Public and private courses taught by expert instructors across half-day to two-day courses

Register now

Learning

Blended Learning

Self-paced and weekly instructor-led sessions for every style of learner to optimize course completion and knowledge retention. Go to Subscriptions Catalog tab to purchase

Purchase now

Scale

Skills@Scale

Comprehensive training offering for large scale customers that includes learning elements for every style of learning. Inquire with your account executive for details

Upcoming Public Classes

Databricks Performance Optimization - Mandarin Chinese

Databricks Performance Optimization 课程向数据工程师和分析师讲授如何在 Databricks Data Intelligence Platform 上诊断、衡量和修复性能瓶颈,以及如何将性能改进与成本关联起来。本课程遵循"先衡量、先利用平台"的工作流:让平台的自动优化功能承担繁重工作,通过 Query Profile 和系统表验证已应用的优化,再在必要时进行手动调优。

学员将以 Query Profile、Performance Insights 和查询历史记录系统表为基础,建立衡量性能的能力。在此基础上,他们将探索关键优化技术,包括数据布局与自调优托管表、liquid clustering、缓存与中间结果、shuffle、数据倾斜、溢出、行爆炸、驱动程序性能、Python UDF、serverless compute、Photon 以及成本归因。

在整个课程中,学员将针对合成零售数据中刻意设计的慢查询进行探索,并观察不同优化技术对性能的影响。他们将利用文件裁剪、任务执行时间、Photon 覆盖率和成本等依据来评估改进效果。两个基于场景的实验将强化从诊断到验证的完整优化工作流。

注意:Databricks Academy 正在将 Databricks 环境中的课堂教学转为基于 notebook 的形式,不再使用幻灯片进行授课。您可以在 Vocareum 实验环境中访问课程 notebook。

Languages Available: English | 日本語 | Português BR | 한국어

Paid
4h
Lab
instructor-led
Professional

Questions?

If you have any questions, please refer to our Frequently Asked Questions page.