Skip to main content
DATA ENGINEERING

Lakeflow: Unified agentic data engineering

Ingest, transform and orchestrate ETL and data pipelines with agentic authoring and operations
Databricks pipeline UI with code, sources, and DAG
TOP COMPANIES USE LAKEFLOW
BENEFITS

Lakeflow delivers high-quality data at agent scale

A unified platform with purpose-built AI to safely accelerate pipeline creation, operations and governance.

Unified ingestion, transformation and orchestration

Connect ingestion, transformation and orchestration in one execution engine governed by Unity Catalog. Eliminate tool sprawl and fragmented security across batch and streaming pipelines on a foundation built to scale.

Agentic pipeline development with Genie Code

Accelerate reliable pipeline delivery with AI built for data engineering context. Author, test and deploy declarative pipelines using natural language with Genie Code, coding agents or visual experiences.

Autonomous pipeline operations with Genie ZeroOps

Continuously monitor pipeline health, catch schema drift and remediate runtime failures. Genie ZeroOps uses telemetry and lineage to surface root causes and deploy validated fixes, keeping production pipelines running smoothly.

85% faster development

Porsche uses the Lakeflow Salesforce connector to ingest CRM data, improving customer experience and strengthening the bond to their brand throughout the customer journey.

Read the Porsche story

50% cost reduction

Hinge Health improves patient outcomes with personalized care plans and managed a 10x data growth while keeping total cost of ownership in check.

Read the Hinge Health story

99% reduction in pipeline latency

Volvo uses Lakeflow to efficiently process and orchestrate real-time data, fueling their global inventory management system for hundreds of thousands of spare parts.

Read the Volvo story

4,500+ weekly jobs orchestrated

AccuWeather uses Lakeflow to orchestrate the consolidation of high-volume weather data. Moving to Databricks serverless infrastructure helped the team reduce maintenance burden and cut costs.

Read the AccuWeather story
DATA ENGINEERING PRODUCTS

Data engineering products for production pipelines

Genie Code

Build and maintain data pipelines with agentic AI that understands your data.

Lakeflow Connect

Efficient data ingestion connectors unlock easy access to analytics and AI with unified governance.

Apache Spark Declarative Pipelines

An open, declarative framework for data teams and AI agents to easily build and operate reliable batch and streaming pipelines.

Lakeflow Jobs

Equip teams to better automate and orchestrate any ETL, analytics and AI workflow with deep observability, high reliability and broad task type support.

Unity Catalog

Seamlessly govern your data and AI assets with a unified, open governance solution — providing automated, column-level data lineage, fine-grained access controls enforced at scale and audit logs to help support audit readiness and regulatory and data-privacy requirements.

Lakeflow Designer

Prepare and transform data with AI-first authoring, directly on Databricks.

use cases

Agentic data engineering use cases with Lakeflow

Logos of various software and tech companies.

Ingest data from any source with Lakeflow Connect

Ingest data from databases, SaaS applications, event streams and cloud storage into the Databricks Platform without building custom ingestion pipelines. Automated change data capture (CDC) and schema evolution continuously deliver high-quality data, giving AI agents the rich enterprise context they need to operate reliably.

Leading organizations that run on Lakeflow

Get started with Lakeflow training and documentation

Agentic data engineering FAQ

Data engineering is the practice of taking raw data from a data source and processing it so it’s stored and organized for a downstream use case such as data analytics, business intelligence (BI) or machine learning (ML) model training. In other words, it’s the process of preparing data so value can be extracted from it. An example of a common data engineering pattern is ETL (extract, transform, load), which defines a data pipeline that extracts data from a data source, transforms it and loads (or stores) it into a target system like a data warehouse.

Ready to become a data + AI company?

Take the first steps in your data transformation