Build Data Pipelines with Lakeflow Spark Declarative Pipelines - Mandarin Chinese
本课程向用户介绍使用 Databricks 中的 Apache Spark™ Declarative Pipelines (SDP) 构建数据管道所需的基本概念和技能,涵盖通过多个流式处理表和物化视图进行增量批处理或流式处理摄取与处理。本课程专为初次接触 Spark Declarative Pipelines 的数据工程师设计,全面介绍核心组件,包括增量数据处理、流式处理表、物化视图和临时视图,并重点说明其各自的用途与区别。
课程涵盖以下主题:
• 使用 Spark Declarative Pipelines 中的多文件编辑器,通过 SQL 开发和调试 ETL 管道(并提供 Python 代码示例)
• Spark Declarative Pipelines 如何通过管道图形跟踪管道中的数据依赖关系
• 配置管道 compute 资源、数据资产、触发器模式及其他高级选项
随后,课程介绍 Spark Declarative Pipelines 中的数据质量期望,引导用户将期望集成到管道中,以验证和强制执行数据完整性。学员还将浏览如何将管道投入生产,包括调度选项,以及启用管道事件日志记录以监控管道性能和健康状况。
最后,课程介绍如何在 Spark Declarative Pipelines 中使用 AUTO CDC INTO 语法实现变更数据捕获 (CDC),以管理缓慢变化维度(SCD Type 1 和 Type 2),帮助用户将 CDC 集成到自己的管道中。
注意:对于 SCORM 课程文件,请确保完成内容后关闭 SCORM 窗口。请勿点击‘Next Lesson’按钮,否则可能会导致 SCORM 模块无法标记为已完成。
本课程内容面向具备以下技能、知识和能力的学员:
• 对 Databricks Data Intelligence Platform 的基本了解,包括 Databricks Workspaces、Apache Spark、Delta Lake、金银铜架构、Lakeflow Jobs 和 Unity Catalog。
• 具备将原始数据摄取到 Delta 表的经验,包括使用 `read_files` SQL 函数加载 CSV、JSON、TXT 和 Parquet 等格式。
• 熟练使用 SQL 转换数据,包括编写中级查询以及对 SQL 连接的基本理解。
• 了解 ETL 概念以及批处理/流式处理工作流。
Self-Paced
Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos
Registration options
Databricks has a delivery method for wherever you are on your learning journey
Self-Paced
Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos
Register nowInstructor-Led
Public and private courses taught by expert instructors across half-day to two-day courses
Register nowBlended Learning
Self-paced and weekly instructor-led sessions for every style of learner to optimize course completion and knowledge retention. Go to Subscriptions Catalog tab to purchase
Purchase nowSkills@Scale
Comprehensive training offering for large scale customers that includes learning elements for every style of learning. Inquire with your account executive for details

