Securely store and share data across public sector agencies and maintain regulatory compliance with Deloitte and Databricks
In today's environment, proactive cybersecurity is crucial to any public sector agency. For many organizations, log data that security professionals need for effective threat monitoring and incident response is not readily accessible in one place, or it lives in siloed departments. In some instances, the data may also be stored only for short-term operational purposes. This severely limits the ability to effectively manage security, and underscores the need for effective log retention as well as secure access to critical cyber information.
Federal mandates are requiring agencies to retain information systems logs over a multi-year period to support the detection, investigation, and remediation of cyber incidents. This creates multiple challenges for agencies to navigate. First, storing massive volumes can be costly, particularly if done in relatively high-cost on-premises or proprietary storage. Furthermore, transferring large volumes of data to a single monolithic repository to provide centralized access can also be expensive and result in data duplication across multiple environments. In short, the memorandum significantly increases data management and cybersecurity demands on federal organizations.

Deloitte's Cyber Data Optimization solution looks to address these challenges by employing a hub-and-spoke model on the Databricks Data + AI Platform. A central analytics "Lakehouse Hub" coordinates with enterprise clouds and source systems, the "Nodes", to establish a centralized analytics layer for log data. Data is retained in low-cost cloud storage at the nodes and accessible by centralized queries from the hub, avoiding transfer of raw data across cloud boundaries. This multi-node, federated model allows data to be securely shared from individual nodes to the central hub, enabling comprehensive log access to address potential cyber threats more efficiently. This approach allows organizations to navigate the changing cyber landscape more effectively while avoiding costly data storage and egress.
Federal compliance requires that organizations not only collect an extensive list of system logs for an extended retention period, but also ensure comprehensive data visibility in order to support cybersecurity operations. The scale of log data volumes can make it technically and financially unsupportable for many organizations within their current toolbox.
Deloitte’s Cyber Data Optimization solution addresses these cost and scale challenges by leveraging low-cost cloud storage, reducing the need for expensive data indexing in proprietary systems. This is particularly impactful for high-volume telemetry data that is growing to petabyte scale.
The federated model provides centralized access and visibility to remote data distributed across the organization. Security operations center (SOC) analysts then have the opportunity to compile, search and perform advanced analytics on log data, enabling rapid response to cyber investigations that require significant historical data.
The hub-and-spoke architecture manages large volume log data across multi-cloud environments by eliminating data duplication and reducing data egress transfer. The framework is a federation of Databricks workspaces that take advantage of a distributed medallion data pattern, incrementally increasing data quality at each node as data flows from raw to consumption-ready. Nodes are deployed at or near source systems as much as possible. Raw log data is ingested at the node, processed, and made available to be queried by the central hub. This eliminates costly data egress across clouds and regions by keeping the source log data at a single node. Only curated responses to federated queries by the hub are transferred from node to hub.

Ensuring the right users have the right access to log data is vital. By leveraging the Databricks governance framework, the hub defines and enforces access control rules that associate role-based user pools with collections of log datasets. In cases where more granular access management is needed, dynamic view functions can be constructed for row/column-level permissions or data masking.
The Cyber Lakehouse integrates with common systems familiar to the organization’s workforce, augmenting the existing toolset while maintaining continuity and accelerating adoption. This eliminates the need for additional training while leveraging the benefits of the Databricks Data + AI Platform. With the Cyber Data Optimization solution, several use cases have been exercised such as:
The Cyber Data Optimization Solutions pairs the deep industry experience of Deloitte with the Databricks Data + AI Platform. With Brickbuilder Solutions, you are guaranteed to get:
Deloitte will be at the Databricks Government Forum on December 11. Come meet the team in person and see the Cyber Data Optimization solution in action by registering here.
It solves the challenge of storing and accessing years of log data that federal mandates require for cyber threat detection, investigation, and remediation, when that data is otherwise siloed or kept only for short-term operational purposes. Traditional approaches force agencies to choose between high-cost on-premises or proprietary storage and the expense of moving massive log volumes into one monolithic repository, which also creates data duplication. Deloitte's Cyber Data Optimization solution addresses this by using a hub-and-spoke model on the Databricks Data + AI Platform instead.
A central analytics "Lakehouse Hub" coordinates with enterprise clouds and source systems, called "Nodes," to create a centralized analytics layer for log data without moving the raw data itself. Log data stays in low-cost cloud storage at each node and is accessible through centralized queries from the hub, and the framework uses a distributed medallion data pattern so data quality increases as it flows from raw to consumption-ready at each node. Only curated responses to federated queries from the hub are transferred back from a node, keeping source log data at a single location.
It reduces costs by keeping log data in low-cost cloud storage at its source node rather than transferring it into expensive, high-cost on-premises or proprietary systems. Because raw data is never duplicated across environments or moved across cloud boundaries, agencies avoid both data duplication and costly data egress. This also cuts down on the expensive data indexing typically required by proprietary systems, which matters most for high-volume telemetry data growing to petabyte scale.
Access is governed centrally through the Databricks governance framework, which lets the hub define and enforce access control rules tied to role-based user pools for specific collections of log datasets. When more granular control is needed, dynamic view functions can be built to enforce row- or column-level permissions or apply data masking. This ensures the right users have the right access to log data even though it remains distributed across multiple nodes.
No, the Cyber Lakehouse is designed to integrate with the systems an organization's workforce already uses rather than replace them, which avoids added training and speeds adoption. Demonstrated use cases include BI tool dashboards populated with log data aggregated from across the enterprise, SIEM tool queries pushed down to the lakehouse and returned as results without requiring SIEM data ingestion and indexing, and alerts detected at the nodes being pushed up to the BI or SIEM interface.
Subscribe to our blog and get the latest posts delivered to your inbox.