Machine Learning at Scale - Mandarin Chinese
在本课程中,您将深入了解 Apache Spark 的架构及其在 Databricks 中应用于机器学习工作负载的理论与实践知识。您将学习何时使用 Spark 进行数据准备、模型训练和部署,同时获得使用 Spark ML 以及 Pandas APIs on Spark 的实践经验。本课程将向您介绍高级概念,例如超参数调优以及使用 Spark 扩展 Optuna。本课程将使用助理课程中介绍的功能和概念,例如 MLflow 和 Unity Catalog,以实现全面的模型打包与治理。
注意: 对于 SCORM 讲授文件,请确保在完成内容后关闭 SCORM 窗口。请勿点击“Next Lesson”按钮,否则可能导致 SCORM 模块无法被标记为已完成。
本课程内容面向具备以下技能、知识和能力的学员:
• 熟悉 Databricks Data Intelligence Platform 及基本工作区运营(创建集群、在笔记本中运行代码、使用基本笔记本操作、从 Git 导入 Repos)
• 具备中级 Python 编程经验,包括数据处理库(Pandas、NumPy)和机器学习框架(Scikit-Learn)
• 掌握 Apache Spark 和 PySpark 基础知识,包括 DataFrames、转换及分布式数据处理操作
• 理解机器学习概念,包括模型训练、评估、超参数调优及部署工作流
• 具备中级 Delta Lake 运营经验(创建表、执行更新、优化文件、时间旅行功能)
• 基本熟悉 MLflow 的实验追踪、模型记录及 Model Registry 运营
• 理解分布式计算概念(集群架构、并行化、可扩展性考量)
• 掌握基本 SQL 知识,用于在 Spark 环境中进行数据查询与处理
Self-Paced
Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos
Registration options
Databricks has a delivery method for wherever you are on your learning journey
Self-Paced
Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos
Register nowInstructor-Led
Public and private courses taught by expert instructors across half-day to two-day courses
Register nowBlended Learning
Self-paced and weekly instructor-led sessions for every style of learner to optimize course completion and knowledge retention. Go to Subscriptions Catalog tab to purchase
Purchase nowSkills@Scale
Comprehensive training offering for large scale customers that includes learning elements for every style of learning. Inquire with your account executive for details

