DB Databricks Financial Services Australia

National Australia Bank (NAB)

Data Lake / Data Warehouse Modernization · Real-time Analytics

NAB had standardized on Databricks but its 300–400 data engineers were building 1,600+ Spark pipelines in inconsistent ways, with hand-written Spark and legacy SQL DMLs of 5,000–8,000 lines and some pipelines containing 64 union-alls. The bank needed a single, governed way to run, standardize, and productionize Spark workloads from ingestion to transformation without slowing teams down.

100% of Bronze pipelines on Spark Declarative Pipelines
120 data sources onboarded in the first year
Data quality improved by 38%
End-to-end latency under 15 minutes

Solution

National Australia Bank (NAB) is consolidating 1,600+ hand-written Spark and legacy SQL pipelines onto Lakeflow Spark Declarative Pipelines on Databricks — 1,800 declarative pipelines run today — to give 300–400 engineers a single, consistent way to build, run, and operationalize Spark workloads. Engineers define what the data should look like (target tables, merge keys, business logic); Lakeflow handles orchestration, scaling, incremental processing, and state management, while AutoCDC eliminates the thousands of lines of hand-coded SCD logic that previously spanned Spark and SQL. Bronze ingestion is 100% declarative, ~50% of the Silver layer is migrated, and Gold is the next target — moving the bank toward a fully declarative, streaming-first architecture with under-15-minute end-to-end latency across AWS, Azure, and Google Cloud.

Data flow

More than 1,600 data pipelines (1,800 declarative pipelines running today) flow from upstream banking source systems into a Bronze→Silver→Gold medallion on Databricks. Spark Declarative Pipelines define the target tables, merge keys, and business logic; Lakeflow handles orchestration, scaling, incremental processing, and state management; AutoCDC eliminates hand-coded SCD logic. Bronze ingestion is 100% declarative, ~50% of Silver is migrated, and Gold is next — enabling end-to-end streaming with under-15-minute latency.

Solution architecture

3 components · 2 layers
  1. Compute
    • Lakeflow Spark Declarative Pipelines Declarative pipeline framework replacing 5,000–8,000-line hand-written Spark DMLs with target-table/merge-key definitions; Bronze 100% migrated, Silver ~50%, Gold next
    • Databricks AutoCDC Built-in change-data-capture that eliminates hand-coded SCD logic previously spanning thousands of lines
  2. Orchestration
    • Lakeflow Runs and operationalizes the declarative pipelines with built-in orchestration, scaling, and reliability

Architecture clues

  • AutoCDC for change data capture across SCDs
  • Lakeflow for orchestration and operationalization
  • Runs on AWS, Azure, and Google Cloud
  • Spark Declarative Pipelines as the standardization layer

Evidence from the source

I call Spark Declarative Pipelines hyper-standardization. There’s only one way to do things - and that consistency is what we were missing.
Latency has dropped dramatically, now supporting under 15 minutes for end to end use cases
Pipeline success rates have increased from 86% to 99.6%