bp is modernizing its global data platform to better support thousands of pipelines, decentralized engineering teams, and a growing need for governed, cost-efficient data operations. Across AWS and Azure environments, teams historically relied on a mix of homegrown frameworks, Spark workloads, orchestration tools, and manually managed infrastructure. While this gave teams flexibility, it also introduced fragmentation, inconsistent development practices, and rising operational overhead.
As bp expanded its Unified Data Platform (UDP) initiative on Databricks under the leadership of Srinivas Chandolu, the company needed a way to standardize data engineering without slowing teams down.
“We had this autonomy versus centralization tradeoff,” explained Hitesh Chouhan, bp’s Tech Lead for the UDP. “Centralization can become a bottleneck, but with autonomy, people start doing things differently, and it becomes difficult to govern and maintain at scale.”
To create a more consistent foundation for data engineering, bp standardized on Databricks Lakeflow for ingestion, transformation, and orchestration workloads, with Apache Spark™ Declarative Pipelines (SDP) becoming the default framework for governed curation and transformation.
Standardizing data pipelines across decentralized teams
Before adopting SDP, teams across bp often implemented pipelines differently depending on the environment, business unit, or individual developer preferences. Some workloads ran on AWS Glue, others relied on Azure-native tooling, and many pipelines required teams to manually manage Spark versions, dependencies, orchestration patterns, and operational safeguards.
The technical challenge wasn’t simply building pipelines. It was maintaining them over time.
“It’s not writing the pattern that’s the problem, it’s maintaining it,” Chouhan said. “As the world moved ahead, we didn’t have the bandwidth to constantly go back and modernize everything.”
bp wanted a standardized framework that could simplify development while enforcing governance and operational consistency across thousands of pipelines. Spark Declarative Pipelines became the default approach.
Today, transformation workloads under the UDP are almost entirely standardized on SDP, representing more than 7,000 transformation pipelines and over 11,000 total pipelines (including the ingestion layer) under the broader platform’s purview. By adopting a declarative framework, bp reduced the need for teams to hand-code operational logic and infrastructure management while gaining tighter governance controls across the organization.
“We do have hand-coded Spark workloads,” Chouhan said. “The problem is that you don’t always know what people are doing. SDP makes sure no one goes out of bounds.” SDP also makes it easy for the team to simplify development across both batch and streaming workloads. “With SDP I can do batch ingestion and streaming — I just flip the switch,” Chouhan said. “I can use the same pipeline for both.”
Because the platform is metadata-driven and standardized around SDP, bp can more consistently govern costs, enforce standards, and improve observability across teams. “In the past, segregating costs was a big pain,” Chouhan said. “With Lakeflow, I can scan system tables and get everything I need with a cost-based model.”
Unifying data engineering on Lakeflow
Standardizing on SDP helped bp simplify the developer experience while reducing the operational burden associated with maintaining large-scale data infrastructure.
Instead of relying on teams to independently manage orchestration patterns, incremental processing logic, streaming reliability, and dependency management, much of that operational complexity is now handled directly within the Lakeflow platform. “We’ve taken the position that if you’re building curation pipelines, you use Lakeflow and you use SDP,” Chouhan said. “Lakeflow is doing a lot of the heavy lifting.”
bp also standardized orchestration on Lakeflow, creating a more unified operational model across ingestion, transformation, and scheduling workflows. Instead of stitching together separate systems for orchestration, governance, and pipeline execution, teams can operate within a single platform while maintaining centralized visibility into costs, observability, and operational health. “Having everything together is the single greatest factor for control and cost management,” said Chouhan.
The shift has also improved cost management and infrastructure efficiency. For one aviation-related workload, bp reduced overall pipeline volume by 24% by consolidating pipelines and standardizing around shared libraries and reusable patterns.
Beyond reducing overhead, the standardized framework allows teams to onboard faster, develop more consistently, and spend less time troubleshooting infrastructure differences across environments.
Accelerating operational decision-making with Databricks
The impact of bp’s modernization effort became especially visible in one of the company’s early operational workloads migrated onto Databricks.
Previously, operational data lived across many disconnected systems, each with its own dashboards and reports. Over time, teams accumulated increasingly large and complex BI environments that made analysis slower and more difficult to operationalize. “One report had over 60 pages and more than 150 visuals,” said Pradeep Ganguru, Airfield Digital Manager at bp.
To simplify the experience, bp consolidated operational data onto Databricks, using Lakeflow as the underlying data engineering framework and Genie for ad hoc analysis. Rather than navigating sprawling reports, operations teams could interact directly with governed operational data through a simpler analytics experience.
The results were immediate. During pre-launch testing, a task that previously required more than five hours for an operations lead to complete was reduced to roughly 20 minutes — an 18x improvement.
Building a governed foundation for long-term scale
By standardizing on Spark Declarative Pipelines as the foundation for transformation and curation, bp established a scalable framework for governed data engineering across decentralized teams. Instead of managing fragmented tooling and inconsistent operational patterns, the company now has a consistent foundation that balances developer flexibility with centralized governance and operational control.
By standardizing on Databricks and unifying data engineering on Lakeflow, bp established a more consistent operational foundation across ingestion, transformation, orchestration, governance, and analytics. Within that broader strategy, Spark Declarative Pipelines became the standard framework for building and operating governed pipelines at scale.
Instead of managing fragmented tooling and inconsistent operational patterns across teams, bp now has a unified platform for data and AI, a unified operational model for data engineering, and a standardized approach to pipeline development and governance.



