Skip to main content
Industries

Manufacturing data and AI: Connecting the product value chain

Give every stage of the manufacturing value chain access to trusted data and AI

by Dr. Philip Laserstein and Dr. Max Köhler

  • Manufacturing data spans a connected product value chain, but the systems that capture it remain separated by function and plant
  • Databricks brings data from disparate systems together – or queries it in place – so teams can answer cross-stage questions across the value chain
  • Governed semantics, natural-language analytics and agentic applications help more people move from finding data to taking action without becoming data engineers

A manufacturing defect rarely belongs to one system. A scrap spike may relate to a machine setting, a supplier batch, a logistics event, or a recurring issue recorded in a quality system. Yet the data needed to investigate it is usually split across plant, functional, and system boundaries.

More than 50 years ago, Dr. Joseph Harrington’s vision of Computer Integrated Manufacturing (CIM) recognized that manufacturing depends on a connected flow of information across functions. Today, that vision is becoming practical as data and AI connect the stages of the product value chain.

The hardest manufacturing questions are cross-stage questions:

  • Which supplier lot reached the affected products?
  • Has this defect appeared before? Did the corrective action hold?
  • Which customers or service cases could be affected?

Answering any of these requires joining data from systems that were never designed to talk to each other. Manufacturers do not need more isolated reports. They need a connected flow of information across the product value chain, with the governance and business context to make that information usable. This is the role a modern Data and AI Platform can play.

What is the manufacturing product value chain?

The product value chain is the end-to-end sequence of functions that transforms raw materials and ideas into products delivered to customers and supported in the field. It connects research and development (R&D) and engineering, purchasing, production and quality, sales and marketing, and aftermarket and field service. Each stage has its own goals, teams, and operational systems:

  • Research and development (R&D) and engineering work with Product Lifecycle Management (PLM), Computer-Aided Design (CAD), Computer-Aided Engineering (CAE) and simulation, test data, requirements, and engineering Bills of Materials (BOMs).
  • Purchasing works with Enterprise Resource Planning (ERP) purchasing, source-to-pay, contracts, supplier risk and supplier portals.
  • Production and quality work with Manufacturing Execution Systems (MES), Supervisory Control and Data Acquisition (SCADA)/Programmable Logic Controller (PLC) data, process historians, Quality Management Systems (QMS)/Laboratory Information Management Systems (LIMS), and maintenance systems.
  • Logistics and supply chain work with Enterprise Resource Planning (ERP), Warehouse Management Systems (WMS), Transportation Management Systems (TMS), planning systems, Electronic Data Interchange (EDI), and telematics.
  • Sales and marketing work with Customer Relationship Management (CRM), Configure-Price-Quote (CPQ), pricing, dealer management, marketing automation, and e-commerce.
  • Aftermarket and field service work with service management, warranty, service-parts planning, connected-product data, diagnostics, and tickets.
image2.png
Figure 1. The product value chain as one connected flow of information, not a set of isolated silos.

Each stage produces valuable operational data. The larger opportunity comes from connecting a finding in one part of the chain to an action in another: a quality issue found in production investigated through logistics and supplier records, or a supplier alert traced forward to every product that received the affected material.

Why system boundaries are the real problem

When every stage of the value chain is isolated, a cross-stage question becomes a manual project of tickets, exports and reconciliation. When the data is accessible as one governed system, the same question becomes a query.

Example 1: A plant quality engineer needs to answer three questions:

  1. Is a scrap spike caused by the batch, the machine, or the setup, and is it still happening now?
  2. Have we seen this defect before, and did the fix hold?
  3. Why does one plant scrap far more of the same part than another?

Each question spans multiple systems: Manufacturing Execution System (MES) records, process historian data, supplier and Supplier Quality Management (SQM) data, Quality Management System (QMS) history, Eight Disciplines (8D) records, and often, multiple plant instances. Today, answering each one means raising tickets, pulling manual exports and relying on a few specialists.

Example 2: A purchasing analyst needs to answer three questions:

  • Which critical parts depend on a single supplier that is now flagged as delivery-risk?
  • Has a supplier's on-time and quality performance been slipping across recent orders?
  • If one supplier fails, which products, plants and open orders are exposed?

Each question spans multiple systems: Enterprise Resource Planning (ERP) purchasing and source-to-pay records, contracts, supplier risk feeds and supplier portals. Today, answering each one can become a small project of tickets, specialist knowledge, and manual exports.

Pooling those operational data sources removes that friction. Bring the value chain together once, govern it once, and use a shared identifier—a serial, lot, or part number—as the join key, and traceability becomes a query. Six questions, one underlying need: join data across systems that were never designed to connect, and trust the result. Four platform capabilities make that possible.

How Databricks connects manufacturing data

No disruptive migration required

A connected value chain does not require one disruptive migration of every source system. Data can be copied when that is the right choice, or it can remain in place and still be queried.

With zero-copy Open Sharing and Lakehouse Federation, organizations can access data in its source systems without creating another extract, transform and load (ETL) pipeline or copy for every use case. When mirroring is appropriate, connectors and cloud object storage provide a scalable path for bringing data into the lakehouse.

The Databricks Data and AI Platform brings this flexibility together with the capabilities manufacturing teams need: historical analysis across production, quality and supply data; low-latency processing for machine and vehicle telemetry; and applications that can read and write individual records quickly.

The four capabilities the examples rely on

Pool or federate the source data, without a disruptive migration: The quality engineer's question spans MES, process historian, supplier and SQM, QMS and 8D records across plant instances; the purchasing analyst's spans ERP purchasing, contracts and supplier-risk feeds. With zero-copy Open Sharing and Lakehouse Federation, that data can be queried in place, and mirrored through connectors when copying is the better choice. Either way, no new ETL copy is required per question.

Orchestration and refinement: Both examples join against governed gold tables, not raw extracts. Lakeflow helps teams build, schedule, and monitor the pipelines that turn raw inputs into trusted, analysis-ready data, commonly through Bronze, Silver, and Gold layers.

Governance for data and AI: Unity Catalog is the single control plane across mirrored and federated data: one permission model, full lineage, and discovery spanning data, models, and AI agents, so the same governed surface answers both examples. Unity Gateway enables you to control AI access, spend and observability across agents, tools, models and MCPs.

Agentic capabilities: Built on this governed foundation, Genie One is an AI coworker that connects to your data, Agent Bricks helps build AI agents grounded in enterprise data, and Genie App Builder lets anyone create agents and applications in natural language. The engineer and the analyst can ask their questions in plain language and get answers grounded in definitions the business recognizes, the subject of the next two sections.

The practical design principle is simple: pool the data where it adds value, federate it where copying does not make sense, and govern both through the same control plane. The result is a way to work across manufacturing data types and stages without creating a new silo for every analytical or AI use case.

Enhancing data literacy – making data useful to more than specialists

A platform is only as valuable as the people who can use it. In many manufacturing organizations, a small group of experts understands the plant-specific systems, while business users wait for reports or exports.

Data literacy grows when people can move through a practical progression: finding relevant data, understanding trusted definitions, analyzing it, asking questions in natural language, building governed agents or applications, and sharing those assets with others. Not everyone needs to become a data engineer to participate. That growth follows six steps, shown below.

image1.png
Figure 2. The data literacy ladder. The same person grows stage by stage, from finding data to building governed apps, without ever having to become a data engineer.

Technology is only part of the change. Training, communities of practice, and a champions network help each function develop confidence and share reusable patterns. A common, governed data surface gives those communities something concrete to work from: shared definitions, common vocabulary, and answers that can be reused across teams.

How "Talk to Data" works in manufacturing (and why it requires governed semantics)

The fastest way to enhance data literacy is to let people talk to their data: no query language to learn, no ticket to raise, and no report to wait for. Natural-language analytics can lower the barrier to data, but a conversational interface is not enough. The answer must be grounded in definitions that the business recognizes.

Example: Purchasing

A question such as “Which critical parts depend on a single supplier that is now flagged for delivery risk?” may require knowledge of SAP specific table headers, joins, and business rules. A purchasing analyst should not need to become a data specialist to investigate it.

The Purchasing Genie Demo makes this pattern concrete. The runnable project follows three steps: prepare governed data, build expert agents, and then compose them under a supervisor and share as one governed app.

image3.png
Figure 3. The end-to-end example in three steps: prepare governed data, build expert agents, then compose them under a supervisor and share as one governed app.

This pattern separates the work that requires technical expertise – preparing and governing the data – from the work that business users should be able to do themselves: asking questions, reviewing answers and taking action.

Example: KPI reporting

The same pattern applies to everyday reporting. When KPI definitions live in a governed semantic layer, both BI tools and AI work from the same business logic. Genie Agents answer questions within a function using those trusted definitions, and Agent Bricks composes them into persona-based agents that work across functions. The result: the same question gets the same answer every time, grounded in definitions the business recognizes rather than inferred from complex schemas or disconnected reports.

Mercedes-Benz Korea applies this pattern in practice, you can learn more here.

A practical path forward

More than 50 years ago, Dr. Joseph Harrington’s vision of Computer Integrated Manufacturing described manufacturing as one cohesive system, unified by the flow of information. That vision is now achievable at scale.

The implementation recipe has three steps:

  1. Pool the data once and govern it once with Unity Catalog as the single control plane, covering both mirrored and federated data sources.
  2. Make a shared identifier the join key, a serial, lot, or part number, so findings in one part of the value chain can drive action in another, enabling end-to-end traceability as a query. Which identifier fits depend on the domain: serialized units and VINs in some, lot or batch numbers in process manufacturing.
  3. Let people talk to their data by grounding natural-language access in governed business semantics.

When business users can ask questions in plain language and receive trustworthy answers, data literacy stops being the privilege of a few specialists and becomes an organizational capability. The outcome is not another dashboard. It is thousands of people across the value chain who can find, understand, and act on trusted data.

Frequently asked questions (FAQ)

What is the biggest benefit of connecting manufacturing data across the value chain?

The biggest benefit of connecting manufacturing data across the value chain is the ability to answer questions that begin in one stage and require action in another. A production defect can be traced through logistics to the supplier batch that caused it, or a suspect supplier lot can be traced forward to every product it reached. Without connected data, each of these investigations requires days of manual work. With a connected lakehouse, the same question becomes a query.

Do manufacturers need to migrate all source data to Databricks?

Manufacturers do not need to migrate all source data to Databricks. Data can be copied when that is useful, but zero-copy Open Sharing and Lakehouse Federation allow data to remain in source systems and be queried in place. Unity Catalog can govern both federated and mirrored data.

What is traceability?

Traceability links manufacturing process steps, products, materials and operational records so teams can trace backward from an affected product or forward from a suspect material. It supports faster problem solving and more precise recall analysis. One example is to trace the serial or lot numbers of the products produced.

What is LTAP, and why do manufacturers need it?

LTAP means Lake Transactional/Analytical Processing. It describes running analytical and transactional workloads on one governed platform. Manufacturing needs both large-scale analysis of historical data and responsive applications that read and write operational records. Historically these lived in two separate stacks with data copied between them; LTAP runs both on one copy of governed data.

What makes data literacy grow beyond platform features?

Features alone are not enough. Training, communities of practice, and champions in each function help people build confidence, share patterns, and use a common governed data surface. The platform provides the foundation; an active community turns it into broad capability.

How does “Talk to Data” become trustworthy?

Natural-language access should be grounded in governed business semantics: documented measures, dimensions, joins and definitions. Genie can then answer questions using the same trusted context rather than inferring meaning from complex schemas or disconnected reports.

What is Lakehouse Federation?

Lakehouse Federation is a Databricks capability that allows data to remain in its source system while still being queried through Databricks. It avoids creating additional ETL pipelines or data copies. Source data is governed through Unity Catalog alongside any mirrored data, giving a unified governance layer regardless of where data physically lives.

What is Unity Catalog in the context of manufacturing data?

Unity Catalog is Databricks's unified governance layer for data, models, and AI agents. It provides one permission model, full data lineage, and discovery across the entire data landscape. For manufacturing, it means a single control plane that governs production data, quality records, supplier data, and AI agents, whether that data lives in the lakehouse or is federated from source systems.

Does the platform handle high-volume machine and vehicle telemetry?

Yes. The platform combines batch and streaming. High-volume machine and vehicle telemetry can be ingested as a continuous, very-low-latency stream and processed reliably, so no events are dropped along the way.

How does data get into Databricks?

Ingestion is designed to be easy. Zerobus Ingest supports push-based streaming directly into governed tables, Lakeflow Connect provides managed connectors and change data capture for business systems, and Auto Loader and Structured Streaming handle files and events.

Is the platform open and multi-cloud?

Yes. Delta tables are stored on cloud object storage, with storage and compute scaling independently. Databricks runs on AWS, Azure and Google Cloud, and data lands in open formats such as Delta Lake and Iceberg.

Get the latest posts in your inbox

Subscribe to our blog and get the latest posts delivered to your inbox.