Feature Engineering at Scale - Mandarin Chinese
在本课程中,您将全面了解如何在 Databricks 平台上设计、扩展和运营化端到端特征工程管道。课程内容分为三个递进式模块:掌握 Spark 分布式执行与优化的基础知识、使用 Auto Loader 和声明式 LakeFlow 管道实现可扩展的数据摄取,以及借助 Databricks 特征存储进阶至生产级 MLOps。
您将参与实践学习,内容包括:使用 Catalyst optimizer 和 Spark UI 调试 Spark 性能、构建具备自动化质量检查的稳健铜-白银-黄金奖章架构、以及使用 SparkML 实现可扩展的特征转换。课程最终将涵盖通过在线特征存储部署实时特征服务、使用按需转换定义 FeatureSpec,以及借助 Unity Catalog 应用治理与血缘追踪。
完成本课程后,您将具备在企业级规模下管理机器学习完整特征生命周期的能力——设计出高性能、可靠、受治理且面向生产的管道。
注意:对于 SCORM 课程文件,请确保完成内容后关闭 SCORM 窗口。请勿点击“Next Lesson”按钮,否则可能导致 SCORM 模块无法标记为已完成。
本课程内容面向具备以下技能、知识和能力的学员:
1. 已完成"Apache Spark 简介"课程,或具备同等的 Spark 基础知识,包括基本数据转换和 Spark SQL。
* 学习者应熟悉 Spark 在分布式数据处理中的作用。本课程将在此基础上进一步讲解 Spark 如何支持可扩展的机器学习工作流。
2. 具备中级 Python 编程能力,尤其是使用 `pandas`、`numpy` 或 `scikit-learn` 等库进行数据处理的能力。
3. 对传统机器学习工作流有中级程度的理解,包括模型训练、评估和超参数调优。
4. 熟悉 Databricks 平台及 Workflows。
* 强烈建议学习者在学习本课程之前先完成 Databricks 机器学习 Associate 课程。本课程默认学习者已具备在 Databricks 环境中进行 ML 开发的相关知识。
Self-Paced
Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos
Registration options
Databricks has a delivery method for wherever you are on your learning journey
Self-Paced
Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos
Register nowInstructor-Led
Public and private courses taught by expert instructors across half-day to two-day courses
Register nowBlended Learning
Self-paced and weekly instructor-led sessions for every style of learner to optimize course completion and knowledge retention. Go to Subscriptions Catalog tab to purchase
Purchase nowSkills@Scale
Comprehensive training offering for large scale customers that includes learning elements for every style of learning. Inquire with your account executive for details

