Skip to main content

Get Started with Databricks for Machine Learning

In this course, you will learn basic skills that will allow you to use the Databricks Data Intelligence Platform to perform a simple data science and machine learning workflow. You will be given a tour of the workspace, and you will be shown how to work with notebooks. You will explore exploratory data analysis and feature engineering, track and manage models with MLflow, experiment using Databricks Genie Code, and deploy a model with Databricks Model Serving. Finally, the course brings these skills together in a comprehensive lab covering the end-to-end machine learning.


Languages Available: English | 日本語 | Português BR | 한국어

Skill Level
Onboarding
Duration
3h
Prerequisites

In this course, the content was developed for participants with these skills/knowledge/abilities: 
• A beginner-level understanding of Python.

• Basic understanding of DS/ML concepts (e.g. classification and regression models), common model metrics (e.g. F1-score), and Python libraries (e.g. scikit-learn and XGBoost)

Self-Paced

Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos

See all our registration options

Registration options

Databricks has a delivery method for wherever you are on your learning journey

Runtime

Self-Paced

Custom-fit learning paths for data, analytics, and AI roles and career paths through on-demand videos

Register now

Instructors

Instructor-Led

Public and private courses taught by expert instructors across half-day to two-day courses

Register now

Learning

Blended Learning

Self-paced and weekly instructor-led sessions for every style of learner to optimize course completion and knowledge retention. Go to Subscriptions Catalog tab to purchase

Purchase now

Scale

Skills@Scale

Comprehensive training offering for large scale customers that includes learning elements for every style of learning. Inquire with your account executive for details

Upcoming Public Classes

Data Engineer

Build Data Pipelines with Apache Spark Declarative Pipelines - Mandarin Chinese

本课程向用户介绍使用 Databricks 中的 Apache Spark™ Declarative Pipelines (SDP) 构建数据管道所需的基本概念和技能,涵盖通过多个流式处理表和物化视图进行增量批处理或流式处理摄取与处理。本课程专为初次接触 Spark Declarative Pipelines 的数据工程师设计,全面介绍核心组件,包括增量数据处理、流式处理表、物化视图和临时视图,并重点说明其各自的用途与区别。

课程涵盖以下主题:

• 使用 Spark Declarative Pipelines 中的多文件编辑器,通过 SQL 开发和调试 ETL 管道(并提供 Python 代码示例)

• Spark Declarative Pipelines 如何通过管道图形跟踪管道中的数据依赖关系

• 配置管道 compute 资源、数据资产、触发器模式及其他高级选项

随后,课程介绍 Spark Declarative Pipelines 中的数据质量期望,引导用户将期望集成到管道中,以验证和强制执行数据完整性。学员还将浏览如何将管道投入生产,包括调度选项,以及启用管道事件日志记录以监控管道性能和健康状况。

最后,课程介绍如何在 Spark Declarative Pipelines 中使用 AUTO CDC INTO 语法实现变更数据捕获 (CDC),以管理缓慢变化维度(SCD Type 1 和 Type 2),帮助用户将 CDC 集成到自己的管道中。

注意:对于 SCORM 课程文件,请确保完成内容后关闭 SCORM 窗口。请勿点击‘Next Lesson’按钮,否则可能会导致 SCORM 模块无法标记为已完成。

Paid & Subscription
3h
Lab
Associate

Questions?

If you have any questions, please refer to our Frequently Asked Questions page.