Solution
Capital One's data and engineering teams built a centralized Feature Hub on Databricks that serves as the single environment for both historical and operational feature pipelines across millions of accounts, delivering a 360-degree view of customer insights to the credit line increase program. Delta Lake holds the curated feature tables, Photon runs as the vectorized compute engine on top of them, and engineers use Databricks APIs to automate job triggers, monitoring, and reruns. Autoscaling clusters match capacity to workload for both large-scale backfills and the daily operational scoring path, while the same Hub feeds both modeling work and live credit-decisioning, removing the prior split between batch and operational feature stores.
Data flow
Account-level data is curated into Delta Lake feature tables, where Photon executes both historical backfills and daily operational pipelines on autoscaling clusters orchestrated through Databricks APIs. The resulting Feature Hub views feed both the modeling workflow and the production credit-line-increase decisioning application from a single environment.