
A leading retail enterprise's analytics platform, powering revenue, fulfilment, pricing and customer decisions, relied on manual, after-the-fact data checks. Data-quality problems (nulls, duplicates, invalid values, broken references, partial loads, schema drift and stale feeds) were routinely discovered by business users after reports were already wrong. The challenge was to detect, explain and safely fix data issues automatically, before they reached the business.
The platform continuously monitors the core retail data domains, all governed in Unity Catalog and stored in Delta Lake, converging into one trusted, analytics-ready layer:
The legacy approach relied on hard-coded SQL checks run after the fact, with no history and no root cause:
The modern implementation introduces a governed, metadata-driven, AI-assisted path, orchestrated as one Databricks Workflow:
Data-quality issues were caught by business users instead of by the platform. Checks were hard-coded per table, so every new rule needed a code change; there was no single DQ score to tell leadership whether data could be trusted; volume drops and metric shifts went unnoticed because nothing learned what “normal” looked like; and when a check failed, engineers spent hours tracing the root cause manually. Bad records flowed into reports and fixes were applied by hand, with no audit trail.
Sigmasoft designed an AI-Powered Data Quality Monitoring Platform on Databricks, built with PySpark, Spark SQL and Delta Lake, governed by Unity Catalog and orchestrated as a single Databricks Workflow that runs from ingestion to monitoring with no manual steps.
Metadata-Driven DQ Rule Engine
Rules live as configuration in a Delta table, not in code. Ten rule types (not-null, uniqueness, regex format, range, allowed values, referential integrity, SQL business rules, volume, freshness and schema contract) are evaluated automatically against every table. Each rule carries its severity, threshold, quarantine flag and safe auto-fix, rolling up into weighted table and overall DQ scores.
Advantages of the Metadata-Driven Design
Automated Profiling & ML Anomaly Detection
Every column is profiled automatically (nulls, distinct values, min/max, mean, top values). Table metrics (row count, null %, duplicate %, revenue and order value) are tracked over time and scored by a statistical Z-score model and a scikit-learn IsolationForest model trained on historical baselines, catching volume drops and metric shifts that no fixed rule would flag.
AI-Based Root-Cause Analysis
An evidence-correlation engine reads the actual rule failures, anomalies and data to explain why issues happened, for example linking an order volume drop to missing load dates and the orphaned payments it caused, or recognizing inflated amounts as tax applied twice. An LLM summary via Databricks ai_query adds an executive summary with prioritized actions.
Quarantine & Safe Automated Remediation
Failing records are copied to a quarantine table with full payloads. Only safe, deterministic fixes are applied automatically (exact-duplicate removal, trim and case normalization, amount recalculation and schema conformance), producing a trusted silver layer while raw bronze data is never modified.
End-to-end AI-powered data quality platform on Databricks: PySpark, Delta Lake, Unity Catalog and Databricks Workflows.
AI-Powered Data Quality – a trusted foundation for enterprise data and decision-making