
A leading data infrastructure company's Install Base Data Platform (IBDP), a P1-critical application powering renewal and opportunity decisions, ran on a fragmented, single-threaded legacy pipeline. The challenge involved consolidating five disparate enterprise systems (Oracle, ADLS, Amazon S3, PostgreSQL and Salesforce) into one governed, analytics-ready platform, while eliminating the refresh-driven downtime that blocked live sales access to the data.
IBDP consolidates transactional data from five distinct enterprise systems, all converging into one governed, analytics-ready IBDP platform:
The legacy pipeline had no intermediate layer. Source systems were read directly and refreshed once daily:
The modern implementation introduces a persisted, governed 3-layer path, orchestrated and refreshed every 4 hours:
The legacy IBDP pipeline read directly from source systems with no intermediate layer, refreshing once daily via manual Python/ETL processing. Transactional data stayed fragmented across five disconnected systems with no single governed source of truth, and every refresh cycle triggered 4–5 hours of sales-user idle time on a single database, blocking live access during renewal and opportunity decision-making. Failed jobs had to wait a full day for the next cycle to retry, and with no persisted intermediate layer, root-cause analysis and safe reprocessing were effectively unavailable.
Sigmasoft redesigned IBDP as a metadata-driven, Airflow-orchestrated 3-layer Delta architecture, built and parallelized with PySpark, and introduced a custom high-availability mechanism that removed refresh-driven downtime entirely.
3-Layer Staging → Delta → Dim Architecture
Replaced the single-hop legacy pipeline with a persisted Staging (raw, as-landed), Delta (cleansed, deduplicated) and Dim (business-ready dimensional models) layer design in Delta format. Issues are now caught and cleaned before reaching business reports, every layer can be traced stage by stage for root cause analysis, and any layer can be safely rebuilt independently without re-reading from source systems.
Advantages of the Layered Design
Airflow-Orchestrated, Metadata-Driven Pipeline
Automated, dependency-aware scheduling replaced manual daily triggers. Config-driven ingestion allows new sources to be added without rewriting pipeline code, and failed jobs now retry within the same cycle instead of waiting a full day, cutting day-to-day operational overhead.
PySpark Parallel & Distributed Processing
Transformation jobs are parallelized across the cluster for performance at scale, with downstream reporting reading pre-modeled Dim data instead of raw source data.
Three-Database Lock & Release Mechanism
Designed and proposed to eliminate sales-team downtime during refresh cycles. Refresh, read and write duties rotate continuously across three databases (live reads / refreshing / standing by), so no single database ever blocks live sales access, removing the 4–5 hour idle window entirely.
ELT on PySpark and Databricks, with a 3-layer Delta architecture.
IBDP – a unified, scalable foundation for enterprise data and decision-making