Skip to main content

Feature Store

Feature Store: build, govern and serve ML features

feature-store-img-1
Databricks ML pipeline: data processing, training, inference.

Databricks Feature Store connects your data to your models and lets your team focus on what matters. It provides a central, governed layer for building, sharing, and serving ML features. 

Author features with Feature Views, serve them with the Online Feature Store, and govern and discover with Unity Catalog. It’s integrated with Model Serving, MLflow, and Genie Code for a managed, end-to-end ML platform, so you can build, train, and serve a model in hours, not weeks.

Build features with Feature Views

Feature Views define the 'What', managed pipelines handle the 'How'. A single Feature abstraction supports batch and stream features. Use the Feature Engineering SDK for rapid notebook experimentation and point-in-time accurate training data computation. 

Feature Views ensure that the values used during training are consistent with those served online, so you can minimize training-serving skew. When you’re ready to ship, a single API call materializes managed, production-ready feature pipelines, writing data to both online and offline stores.

Better recommendations for hundreds of millions of travelers start with better features. Feature Views cut our feature code dramatically — our data scientists go faster and focus on what drives traveler value, not how to compute it.

 

-- Jules Marshall, Senior Director, Product Management, Data

Serve features at scale with Online Feature Store

Real-time models are only as good as the feature data they consume. Online Feature Store serves features to production applications and model endpoints with low latency at high scale. Feature Store manages sync from Materialized Feature Views or Feature Tables to keep data up to date, while Feature Engineering SDKs ensure models train on the same point-in-time correct feature data served online.

Backed by Lakebase Autoscaling, Databricks’ managed Postgres, Online Feature Store is serverless and fully managed, scaling with your traffic without the complexity of a third-party key-value store. Model Serving endpoints integrate with Lakebase to fetch features or serve predictions without requiring a separate feature fetch.

Databricks UI showing a table of features.

Govern and share features with Unity Catalog

Treat features as governed, reusable assets — not logic buried in individual notebooks and pipelines.

Databricks Feature Store is built into the wider platform, supporting the full ML lifecycle: train models with MLflow, serve them on Databricks, and govern everything through Unity Catalog. Tag, share, and reuse features across use cases.

Lineage is tracked from source data through the features and models that depend on it. When you log a model with Feature Store, its training features are recorded alongside it, so teams can understand dependencies, reuse trusted features, and reproduce how production models were built.

Data processing latency across timeframes.

Real-time feature freshness with Spark Real-Time Mode

Streaming data adds critical context for real-time ML, but supporting it consistently across training and inference is difficult. Feature Views orchestrate Spark Real-Time Mode (RTM) pipelines to materialize streaming features, writing model-ready aggregates directly to Lakebase with <200ms freshness.

Sawtooth Windows combines leading-edge stream data with batch history to support long- or lifetime-window features. For training, managed Ingestion Tables snapshot the stream, while Feature Engineering SDKs use the offline copy for historical feature computation and experimentation.

Resources

Blog

Python code configuring Databricks features

Demos

'How to Build and Serve Production ML Features'

Frequently Asked Questions

Databricks Feature Store is a managed platform for building, governing, and serving machine learning features across the full lifecycle, using either stream or batch data for real-time or batch inference. Features are authored with Feature Views, which define 'what' a feature is while managed pipelines handle 'how' it is computed, and a single API call materializes production-ready pipelines that write to both online and offline stores. It is integrated with Model Serving, MLflow, Genie Code, and Unity Catalog so features, training, and serving stay connected in one platform. Databricks describes this as enabling teams to build, train, and serve a model in hours rather than weeks.

Ready to get started?