Install Base Data Platform (IBDP) Modernization

Overview

A leading data infrastructure company's Install Base Data Platform (IBDP), a P1-critical application powering renewal and opportunity decisions, ran on a fragmented, single-threaded legacy pipeline. The challenge involved consolidating five disparate enterprise systems (Oracle, ADLS, Amazon S3, PostgreSQL and Salesforce) into one governed, analytics-ready platform, while eliminating the refresh-driven downtime that blocked live sales access to the data.

Data Landscape: Source Systems Integrated

IBDP consolidates transactional data from five distinct enterprise systems, all converging into one governed, analytics-ready IBDP platform:

  • Oracle – core transactional database
  • ADLS – on-premises data lake storage
  • Amazon S3 – cloud object storage
  • PostgreSQL – relational data store
  • SFDC – Salesforce CRM data

Architecture Evolution: From Legacy Pipeline to Modern 3-Layer Architecture

The legacy pipeline had no intermediate layer. Source systems were read directly and refreshed once daily:

  • Legacy: Source Systems → Direct Processing (Python, ETL) → Reports / API

The modern implementation introduces a persisted, governed 3-layer path, orchestrated and refreshed every 4 hours:

  • Modern: Source Systems → Staging Layer → Delta Layer → Dim Layer → Reports / API – Airflow-orchestrated, metadata-driven and parallelized with PySpark

Business Challenge

The legacy IBDP pipeline read directly from source systems with no intermediate layer, refreshing once daily via manual Python/ETL processing. Transactional data stayed fragmented across five disconnected systems with no single governed source of truth, and every refresh cycle triggered 4–5 hours of sales-user idle time on a single database, blocking live access during renewal and opportunity decision-making. Failed jobs had to wait a full day for the next cycle to retry, and with no persisted intermediate layer, root-cause analysis and safe reprocessing were effectively unavailable.

Digital Transformation

Sigmasoft redesigned IBDP as a metadata-driven, Airflow-orchestrated 3-layer Delta architecture, built and parallelized with PySpark, and introduced a custom high-availability mechanism that removed refresh-driven downtime entirely.

Solution Delivered

3-Layer Staging → Delta → Dim Architecture
Replaced the single-hop legacy pipeline with a persisted Staging (raw, as-landed), Delta (cleansed, deduplicated) and Dim (business-ready dimensional models) layer design in Delta format. Issues are now caught and cleaned before reaching business reports, every layer can be traced stage by stage for root cause analysis, and any layer can be safely rebuilt independently without re-reading from source systems.

Advantages of the Layered Design

  • Data quality & trust – issues are caught and cleaned at Staging/Delta before reaching business reports
  • Root cause analysis (RCA) – every layer is persisted, so any downstream issue can be traced back stage by stage
  • Safe reprocessing – any layer can be rebuilt independently without re-reading from source systems
  • Performance at scale – downstream reporting reads pre-modeled Dim data instead of raw source data

Airflow-Orchestrated, Metadata-Driven Pipeline
Automated, dependency-aware scheduling replaced manual daily triggers. Config-driven ingestion allows new sources to be added without rewriting pipeline code, and failed jobs now retry within the same cycle instead of waiting a full day, cutting day-to-day operational overhead.

PySpark Parallel & Distributed Processing
Transformation jobs are parallelized across the cluster for performance at scale, with downstream reporting reading pre-modeled Dim data instead of raw source data.

Three-Database Lock & Release Mechanism
Designed and proposed to eliminate sales-team downtime during refresh cycles. Refresh, read and write duties rotate continuously across three databases (live reads / refreshing / standing by), so no single database ever blocks live sales access, removing the 4–5 hour idle window entirely.

Operational Upgrade: Faster, Smarter, More Resilient

  • Airflow-orchestrated scheduling – automated, dependency-aware scheduling replaces manual daily triggers
  • Metadata-driven pipeline – config-driven ingestion adds new sources without rewriting pipeline code
  • Parallel & distributed processing – PySpark parallelizes transformation jobs across the cluster for speed at scale
  • 6x faster refresh cadence – refresh cycle moved from once daily to every 4 hours
  • Faster failure recovery – a failed job now retries within the same cycle, instead of waiting a full day
  • Less manual intervention – metadata-driven orchestration cuts day-to-day operational overhead

Technology Stack

  • Databricks – unified platform for PySpark processing and the Delta layer
  • Airflow – orchestrates and schedules every pipeline run
  • Delta Format – Staging, Delta and Dim layers
  • Oracle / PostgreSQL – relational transactional sources
  • ADLS – data lake
  • Amazon S3 / SFDC – cloud storage and CRM integration
  • 3-DB Lock & Release – custom availability mechanism

Business Impact

  • Unified data foundation – five fragmented sources replaced with one governed, analytics-ready platform.
  • Modern, scalable pipeline – ELT on PySpark with a 3-layer Delta architecture, built and parallelized with PySpark on Databricks.
  • Uninterrupted sales access – the three-DB lock & release mechanism removed refresh-driven downtime for sales users.
  • 6x faster refresh cadence – moved from once daily to every 4 hours.
  • Elimination of the 4–5 hour sales-user idle window per refresh cycle via the 3-DB lock & release mechanism.
  • Five fragmented source systems (Oracle, ADLS, Amazon S3, PostgreSQL, SFDC) unified into one governed, analytics-ready platform.
  • Faster failure recovery – failed jobs retry within the same cycle instead of waiting a full day.
  • Faster root cause analysis – persisted Staging/Delta/Dim layers let issues be traced back stage by stage.
  • Reliable renewal & opportunity decisions – as a P1 application, IBDP underpins renewal and opportunity decisions with consistent data.
  • Governed and replayable – any layer can be safely rebuilt independently without re-reading from source systems.

Project Ownership & Reporting

ELT on PySpark and Databricks, with a 3-layer Delta architecture.

IBDP – a unified, scalable foundation for enterprise data and decision-making

Client

Leading data infrastructure company