DB Databricks Financial Services United States

Capital One

Real-time Analytics · Risk Management

Capital One's credit line increase program relies on intelligent, trusted data to deliver financial empowerment to millions of customers, but growing data volumes and complex pipelines meant the team needed a single, centralized feature environment that could handle both historical and operational feature pipelines and run large-scale backfills quickly enough to support timely credit decisions.

60x faster compute performance enabling real-time iteration
80% decrease in time and cost per job
New model features shipped in weeks including historical backfills

Solution

Capital One's data and engineering teams built a centralized Feature Hub on Databricks that serves as the single environment for both historical and operational feature pipelines across millions of accounts, delivering a 360-degree view of customer insights to the credit line increase program. Delta Lake holds the curated feature tables, Photon runs as the vectorized compute engine on top of them, and engineers use Databricks APIs to automate job triggers, monitoring, and reruns. Autoscaling clusters match capacity to workload for both large-scale backfills and the daily operational scoring path, while the same Hub feeds both modeling work and live credit-decisioning, removing the prior split between batch and operational feature stores.

Data flow

Account-level data is curated into Delta Lake feature tables, where Photon executes both historical backfills and daily operational pipelines on autoscaling clusters orchestrated through Databricks APIs. The resulting Feature Hub views feed both the modeling workflow and the production credit-line-increase decisioning application from a single environment.

Solution architecture

4 components · 3 layers
  1. Compute
    • Databricks Photon Vectorized compute engine that runs large-scale feature backfills and operational scoring jobs
    • Databricks Autoscaling Clusters Matches compute capacity to workload size for both historical backfills and daily jobs
  2. Storage
    • Delta Lake Storage foundation holding curated feature tables across millions of customer accounts
  3. Orchestration
    • Databricks APIs Automates job triggers, monitoring, and reruns across the feature pipelines

Architecture clues

  • Autoscaling clusters tuned to balance spend and performance
  • Centralized Feature Hub built on Databricks
  • Curates records across millions of accounts for a 360-degree customer view
  • Operational scoring and modeling workloads share the same feature store
  • Single environment for historical and operational feature pipelines

Evidence from the source

We consistently see Databricks performing 60x faster than other available systems.
We have seen an 80% decrease in time and cost per job
developing and deploying new model features only takes weeks, including full historical backfills