DB Databricks Media and Entertainment Global

Amagi

Data Lake / Data Warehouse Modernization · Real-time Analytics

Amagi supports 8,000+ streaming channels, 300+ distributors, and 26B annual ad impressions for 45% of the world's top media companies, but its legacy stack (Dataproc for batch, Snowflake for analytics, in-house platforms) created data silos with separate governance models, conflicting numbers across Finance, Operations, and Product, and complex cross-region governance across AWS and GCP.

15x faster time to market for new data products
26 billion annual ad impressions
35% fewer production incidents
45% cost savings

Solution

Amagi rebuilt its data platform as the Amagi Data Platform (ADP) on Databricks, consolidating Dataproc, Snowflake, and multiple in-house systems into a single Lakehouse spanning AWS and GCP. Lakeflow Spark Declarative Pipelines ingest content metadata, asset catalogs, advertising data, service streams, program schedules, and user data into governed Delta Lake tables, where Unity Catalog enforces lineage, access controls, sensitive-data tagging, and cross-cloud compliance. Finance, Operations, and Product teams query the same source of truth through Databricks SQL warehouses, Genie natural-language sessions, and notebooks running on serverless compute, while customer-facing data products for streaming platforms, ad-tech partners, and content owners are derived from the same governed Lakehouse, eliminating reconciliation between systems.

Data flow

Source systems — content metadata, asset catalogs, advertising data, service streams, program schedules, and user data — land in Delta Lake via Lakeflow declarative pipelines. Unity Catalog governs lineage, access, and sensitive-data tagging across the multicloud AWS+GCP footprint. Internal users query the same governed Lakehouse through Databricks SQL warehouses, Genie natural-language sessions, and notebooks on serverless compute, while customer-facing data products (streaming analytics, ad attribution, usage and revenue reporting) are derived from the same source of truth.

Solution architecture

6 components · 4 layers
  1. Compute
    • Lakeflow Spark Declarative Pipelines Ingestion and validation framework that lands data from all source systems into the governed Lakehouse
    • Databricks Serverless Compute Auto-scaling compute that replaces always-on infrastructure across analytics workloads
  2. Storage
    • Delta Lake Lakehouse storage holding the governed source-of-truth tables for content, ads, services, schedules, and user data
  3. Serving
    • Databricks SQL SQL warehouse access for Finance, Operations, and Product analytics on the unified Lakehouse
    • Databricks Genie Natural-language querying for internal stakeholders against governed Lakehouse data
  4. Governance
    • Unity Catalog Cross-cloud governance for lineage, access control, sensitive-data tagging, and compliance across AWS and GCP

Architecture clues

  • Customer-facing data products served from the same platform
  • Internal access via notebooks, Genie, SQL warehouses
  • Multicloud across AWS and GCP
  • Serverless compute for dynamic scaling
  • Unified lakehouse on Databricks with Unity Catalog governance

Evidence from the source

15x Faster time to market for new data products and features
35% Fewer production incidents due to improved data reliability
45% Cost savings after consolidating fragmented data systems