A pipeline run turns green in Databricks Workflows. The job ran in eight minutes, consumed twelve DBUs, wrote forty thousand records into silver, and exited with status zero. Alerts stay silent.
Three hours later, finance opens their morning report. Two dozen transactions show negative revenue. Several orders have null customer IDs. Unsupported currency codes slip through unflagged.
The pipeline did not fail. Spark executed every transformation in the plan. From the engine’s perspective, the run succeeded.
This scenario plays out weekly. Pipeline success proves code ran without unhandled exceptions. It does not establish that output data is fit for consumption.
In twenty-six years across enterprise data platforms, from DTS and SSIS to Hadoop and modern Lakehouses, I see teams trapped between two extremes:
Big-consulting bloat: Spending $500K on governance slide decks while unmonitored clusters bleed DBUs writing garbage into gold tables.
Tutorial-hell resume padding: Copying syntactic tricks from bootcamps and Kaggle projects without understanding the CFO’s cloud bill or audit liabilities.
When corrupt records reach trusted layers, the business pays twice: first in compute and engineering time to repair the damage, and then in lost trust when reports, models, or operational systems produce results nobody can defend.

Scattering assertions throughout notebooks does not solve that problem. Neither does documenting requirements in a wiki that has no connection to runtime behavior.
A data contract only earns its keep when its computable requirements become enforceable checks.
That is where Databricks Labs DQX becomes interesting.
This week we walk through how Databricks Labs DQX, an open-source framework, helps by translating schema and quality requirements into executable PySpark checks. We are going to check out how DQX parses contracts, generates rules, routes quarantine streams, and adapts when business rules change.
What Is Databricks Labs DQX?
Databricks Labs DQX (Data Quality eXtended) is a Python library built for PySpark workloads on Apache Spark. It operates on both batch DataFrames and Spark Structured Streaming pipelines.
DQX provides several core capabilities:
Over eighty built-in check functions for row-level and dataset-level validation.
Declarative and programmatic rule definitions.
Automatic rule generation from Open Data Contract Standard (ODCS) contract specifications.
Row-level error tagging using structured _errors and _warnings columns.
DataFrame splitting into accepted and quarantined streams.
Summary metric persistence into Delta tables for Lakeview dashboards.
Support and Licensing Boundaries
Understand DQX’s operational boundary before adoption. It is a Databricks Labs project, not a core runtime product.
As documented on databricks-labs-dqx on PyPI:
Labs libraries are provided AS-IS without formal vendor SLAs.
Bugs cannot be submitted through standard Databricks support tickets; issues are tracked on GitHub.
Your team owns testing, deployment, and dependency pinning.
This walkthrough pins DQX 0.16.0 (August 2026):
pip install 'databricks-labs-dqx[datacontract]==0.16.0'
The [datacontract] extra supplies ODCS parsing dependencies. Pinned dependencies protect your pipeline from breaking changes across releases.
The Contract Boundary: Agreement vs. Execution
Treating a validation library as a complete data contract solution is an architectural mistake.
Under the Open Data Contract Standard (ODCS), a contract specifies dataset identity, ownership, schema, SLAs, and governance. A validation engine cannot resolve organizational questions. DQX cannot decide who owns an orphaned record, who gets paged when an enum changes, or whether a pipeline should halt.
DQX executes the computable subset: confirming physical data types, checking column presence, and evaluating row-level or dataset predicates. The contract defines the interface; DQX enforces its verifiable rules.
Establishing a Fair Baseline: Plain PySpark Validation
To evaluate DQX fairly, we establish a competent PySpark baseline. We validate a synthetic orders dataset with five fields: order_id, customer_id, order_total, currency, and order_timestamp.
We enforce three business rules:
Order identifiers must be present (order_id non-null).
Order totals must be nonnegative (order_total >= 0.0), as refunds live in a separate pipeline.
Currency codes must belong to supported currencies (USD, EUR, GBP).
Here is a standard PySpark implementation routing invalid records to quarantine:
from pyspark.sql import functions as F
def validate_orders(df):
allowed_currencies = ["USD", "EUR", "GBP"]
missing_order_id = F.col("order_id").isNull()
negative_total = F.col("order_total") < 0.0
unsupported_currency = (
~F.col("currency").isin(allowed_currencies)
| F.col("currency").isNull()
)
tagged_df = df.withColumn(
"_errors",
F.array_compact(
F.array(
F.when(
missing_order_id,
F.struct(
F.lit("order_id_not_null").alias("check_name"),
F.lit("order_id").alias("column"),
F.lit("Order identifier must not be null").alias("message"),
),
),
F.when(
negative_total,
F.struct(
F.lit("order_total_nonnegative").alias("check_name"),
F.lit("order_total").alias("column"),
F.lit("Order total must be >= 0.0").alias("message"),
),
),
F.when(
unsupported_currency,
F.struct(
F.lit("currency_supported").alias("check_name"),
F.lit("currency").alias("column"),
F.lit(
f"Currency must be in {allowed_currencies}"
).alias("message"),
),
),
)
),
)
valid_df = tagged_df.filter(F.size("_errors") == 0).drop("_errors")
quarantine_df = tagged_df.filter(F.size("_errors") > 0)
return valid_df, quarantine_df
This baseline functions cleanly, but introduces operational friction:
Rule duplication: Allowed currencies are hardcoded in Python. Contract changes require manual script edits.
Plumbing boilerplate: Tagging nested structs and filtering array sizes creates maintenance drag.
Reporting overhead: Summarizing pass rates requires custom aggregation passes.
Schema fragility: Upstream column renames trigger unhandled analysis exceptions during plan resolution.
Contract-Driven Validation with DQX
Now let us implement the same validation logic using an Open Data Contract Standard (ODCS) contract and DQX.
1. The Open Data Contract Standard (ODCS) Contract
Here is a minimal ODCS v3.0.2 contract saved as orders_contract.yaml. The required fields generate null checks; the two explicit business rules use DQX's supported type: custom implementation format.
kind: DataContract
apiVersion: v3.0.2
id: urn:datacontract:checkout:orders
name: Orders Ingestion Contract
version: 1.0.0
status: active
domain: checkout
dataProduct: orders
schema:
- name: raw_orders
physicalType: table
properties:
- name: order_id
logicalType: string
physicalType: string
required: true
- name: customer_id
logicalType: string
physicalType: string
required: true
- name: order_total
logicalType: number
physicalType: double
required: true
- name: currency
logicalType: string
physicalType: string
required: true
- name: order_timestamp
logicalType: timestamp
physicalType: timestamp
required: true
quality:
- type: custom
engine: dqx
description: Refunds are handled separately; order totals cannot be negative.
implementation:
name: order_total_nonnegative
criticality: error
check:
function: is_in_range
arguments:
column: order_total
min_limit: 0.0
- type: custom
engine: dqx
description: Only approved settlement currencies are accepted.
implementation:
name: currency_supported
criticality: error
check:
function: is_in_list
arguments:
column: currency
allowed: ["'USD'", "'EUR'", "'GBP'"]
There are two different types of information here, and separating them matters.
The schema tells us what the dataset is expected to look like.
The quality rules tell us what valid business data means within that schema.
A DOUBLE can hold -35.50.
An order total in this particular business process cannot.
A STRING can contain XYZ.
A settlement currency in this pipeline cannot unless XYZ has been approved.
Data types describe representation. Business rules describe meaning.
Do not confuse the two.
2. Generating Executable Rules
In your pipeline, DQGenerator translates the contract into executable DQX rule metadata:
from databricks.sdk import WorkspaceClient
from databricks.labs.dqx.profiler.generator import DQGenerator
from databricks.labs.dqx.engine import DQEngine
ws = WorkspaceClient()
generator = DQGenerator(workspace_client=ws, spark=spark)
dqx_checks = generator.generate_rules_from_contract(
contract_file="orders_contract.yaml",
contract_format="odcs",
generate_predefined_rules=True,
generate_schema_validation=True,
strict_schema_validation=True,
process_text_rules=False,
default_criticality="error",
)
Setting process_text_rules=False skips LLM-based text rules. In DQX 0.16.0, the quoted values inside allowed are intentional: bare strings are treated as column expressions, not currency literals. Review the generated checks and test them against representative data before deployment.
But, at this point, do not immediately throw those generated rules at production data.
Generated rules are compiled artifacts.
Inspect them. Validate them. Test them.
DQX exposes validation for rule metadata, so I would include this in CI:
status = DQEngine.validate_checks(dqx_checks)
if status.has_errors:
raise ValueError("Generated DQX checks failed validation.")
This is an important mindset shift.
The contract may be the authoritative specification, but the generated rule set is still executable software. Treat it that way!
The generated set should contain nullability checks for required properties, the two named business checks, and a dataset-level has_valid_schema check. Assert that those expected checks exist in CI; do not assume a type: library entry will be compiled into a DQX rule.
3. Validating the Synthetic Dataset
We construct ten synthetic records with valid transactions, a zero-total boundary case, single-rule violations, and one multi-rule violation. Timestamp values match the contract's physical type.
from datetime import datetime, timezone
from pyspark.sql.types import (
StructType, StructField, StringType, DoubleType, TimestampType
)
def ts(hour, minute):
return datetime(2026, 9, 18, hour, minute, tzinfo=timezone.utc)
schema = StructType([
StructField("order_id", StringType(), True),
StructField("customer_id", StringType(), True),
StructField("order_total", DoubleType(), True),
StructField("currency", StringType(), True),
StructField("order_timestamp", TimestampType(), True),
])
orders_data = [
("ord-1001", "cust-501", 120.50, "USD", ts(10, 0)),
("ord-1002", "cust-502", 45.00, "EUR", ts(10, 5)),
("ord-1003", "cust-503", 210.75, "GBP", ts(10, 10)),
("ord-1004", "cust-504", 0.00, "USD", ts(10, 15)),
(None, "cust-505", 89.99, "USD", ts(10, 20)),
("ord-1006", "cust-506", -35.50, "USD", ts(10, 25)),
("ord-1007", "cust-507", 145.00, "JPY", ts(10, 30)),
(None, "cust-508", -75.00, "CAD", ts(10, 35)),
("ord-1009", "cust-509", 99.00, "EUR", ts(10, 40)),
("ord-1010", "cust-510", 520.00, "GBP", ts(10, 45)),
]
raw_orders_df = spark.createDataFrame(orders_data, schema=schema)
For the row-level validation path, separate out the dataset-level schema rule:
engine = DQEngine(ws)
row_checks = [
check
for check in dqx_checks
if check.get("check", {}).get("function") != "has_valid_schema"
]
valid_df, quarantine_df = engine.apply_checks_by_metadata_and_split(
raw_orders_df,
row_checks,
)
DQX’s current APIs support both annotation and split patterns. apply_checks_by_metadata retains diagnostic columns on checked data, while apply_checks_by_metadata_and_split returns good and invalid/diagnostic DataFrames.
For this batch, using only error-level rules, the expected result is straightforward:
6 records pass.
4 records contain errors.
The fourth bad record is particularly useful because it violates multiple requirements simultaneously:
order_id = null
order_total = -75.00
currency = CAD
Instead of receiving a generic “validation failed” message, the diagnostic record can carry structured details describing each violated check.
That matters when somebody eventually asks:
Why was this transaction excluded?
You have an answer tied to the individual record instead of a pipeline log containing only a failure count.
4. Inspection of Expected Failures
Here is how each record in our ten-row batch evaluates against the contract:

DQX records violations in _errors. For Record 8, inspect that column to confirm that the missing order ID, negative total, and unapproved currency are each reported. The exact diagnostic message wording is library-version dependent.
5. Routing Semantics: Errors, Warnings, and Pipeline Halt
Understanding how DQX treats severity levels is essential:
Error criticality (criticality: “error”):
Rows with one or more error-level violations route to quarantine_df. Rows with zero error violations route to valid_df.
Warning criticality (criticality: “warn”):
Warnings populate the _warnings array column. Warning violations do not route rows to quarantine. A record with warnings remains in valid_df, preserving data for downstream processing while retaining audit tags.
Record Accounting:
Because routing evaluates size(_errors) == 0 versus size(_errors) > 0, every input record lands in either valid_df or quarantine_df without duplication or loss.
Pipeline Failure Is Explicit:
Assigning criticality: “error” does not automatically abort your Databricks job. DQX tags and splits data; it does not kill the cluster by default. If your requirements demand failing the pipeline on bad data, you must implement that check explicitly:
valid_df.write.mode("append").saveAsTable("silver.orders_clean")
quarantine_df.write.mode("append").saveAsTable("quarantine.orders_invalid")
if quarantine_df.count() > 0:
print(f"WARNING: {quarantine_df.count()} records quarantined.")
Schema Drift Is a Different Failure Mode
What happens if an upstream producer alters payload structure before delivery?
Row-level validation answers:
Is this record valid?
Schema validation answers:
Is this still the dataset we agreed to receive?
We test this using a separate DataFrame to avoid masking row-level checks:
drifted_schema = StructType([
StructField("order_id", StringType(), True),
StructField("customer_id", StringType(), True),
StructField("order_total", DoubleType(), True),
StructField("currency", StringType(), True),
StructField("order_timestamp", TimestampType(), True),
StructField(
"uncontracted_promo_code",
StringType(),
True,
),
])
drifted_data = [
(
"ord-9999",
"cust-999",
50.00,
"USD",
ts(11, 0),
"FALL2026",
)
]
drifted_df = spark.createDataFrame(drifted_data, drifted_schema)
Now run only the generated schema validation check:
schema_checks = [
check
for check in dqx_checks
if check.get("check", {}).get("function") == "has_valid_schema"
]
schema_result = engine.apply_checks_by_metadata(
drifted_df,
schema_checks,
)
schema_result.select("_errors").show(truncate=False)
With strict schema validation enabled, that structural difference can be detected before you silently normalize a producer’s breaking change into downstream tables.
Whether an extra column should actually fail your pipeline is a different question.
That is a contract decision.
Some domains want strict producer interfaces. Others intentionally allow additive schema evolution.
The tool should enforce the architecture, not make the architecture decision for you.
Demonstrating a Requirement Change
Data contracts evolve. Suppose checkout operations expand into Japan and officially add Japanese Yen (JPY) to approved currencies.
In a traditional codebase, an engineer must locate hardcoded lists across repositories, update strings, and redeploy.
In a contract-driven workflow, the process centers on the specification. Copy the first contract to orders_contract_v2.yaml and change only the currency rule's allowed list:
allowed: ["'USD'", "'EUR'", "'GBP'", "'JPY'"]
We regenerate rules and rerun validation against the original ten records:
dqx_checks_v2 = generator.generate_rules_from_contract(
contract_file="orders_contract_v2.yaml",
contract_format="odcs",
generate_predefined_rules=True,
generate_schema_validation=True,
strict_schema_validation=True,
process_text_rules=False,
default_criticality="error",
)
row_checks_v2 = [c for c in dqx_checks_v2 if c.get("check", {}).get("function") != "has_valid_schema"]
valid_v2_df, quarantine_v2_df = engine.apply_checks_by_metadata_and_split(raw_orders_df, row_checks_v2)
print(f"V2 Accepted: {valid_v2_df.count()} | V2 Quarantined: {quarantine_v2_df.count()}")
Result of the Change
When we run V2:
Accepted count increases from 6 to 7.
Quarantined count decreases from 4 to 3.
Record 7 (ord-1007 with JPY) is now accepted into the clean silver table.
Record 8 remains in quarantine because its currency is CAD and it still violates identifier nullability and nonnegative total constraints.
The pipeline code did not change. The specification drove the updated execution plan. Validate this contract change and the expected 7/3 split in your environment before releasing it.
The Quality Tool Landscape: DQX vs. Alternatives
DQX is not the only validation mechanism on Databricks. Understanding where each tool fits clarifies architecture choices:
Delta Constraints: Enforce hard invariants at storage (such as CHECK (order_total >= 0)). They fail transactions immediately on write rather than routing records.
Lakeflow Pipeline Expectations: Best when pipelines run entirely in Lakeflow and aggregate event log metrics satisfy monitoring needs.
Great Expectations: Best for multi-engine architectures outside Databricks needing documentation-centric validation suites.
dbt Tests: Best when transformations are purely SQL models managed inside dbt.
Databricks Labs DQX: Best for PySpark workloads needing contract-driven generation, row-level diagnostic structs, and automated quarantine routing.
DQX vs. Lakeflow Pipeline Expectations
Comparing DQX to Lakeflow Pipeline Expectations reveals distinct trade-offs across five dimensions:
Execution Context: Lakeflow expectations run exclusively inside Lakeflow Declarative Pipelines. DQX runs across standard Databricks Jobs, interactive notebooks, Python wheels, Structured Streaming, and Lakeflow pipelines.
Rule Management: Lakeflow expectations are embedded directly in transformation code as decorators or SQL constraints. DQX externalizes rules into ODCS YAML files, Unity Catalog tables, Volumes, or Lakebase.
Failure Information: Lakeflow expectations record aggregate counts in event logs without attaching failure metadata to individual rows. DQX appends structured _errors and _warnings struct arrays directly to each record.
Quarantine Routing: Lakeflow requires declaring two separate tables with inverted filter logic to split clean and bad records. DQX provides native splitting via apply_checks_by_metadata_and_split() for metadata rules.
Support & SLAs: Lakeflow expectations are a core product with enterprise support SLAs. DQX is a Databricks Labs project provided AS-IS without vendor SLAs.
Lakeflow Expectations are excellent when your quality rules naturally belong inside a Lakeflow pipeline. Databricks currently supports warn, drop, and fail behaviors and documents explicit quarantine patterns for preserving invalid records.
Databricks also recommends externalizing reusable expectation definitions rather than assuming every rule must be hardcoded into decorators, so the distinction is no longer simply “Lakeflow equals code; DQX equals metadata.”
The more useful distinction is this:
Lakeflow Expectations are pipeline-native quality controls. DQX provides a broader PySpark quality framework with contract-driven rule generation and persistent row-level diagnostics.
There are plenty of architectures where both make sense.
Use Lakeflow for pipeline orchestration and native enforcement.
Use DQX where contract compilation, richer diagnostics, or reusable validation across jobs matters.
Coexistence: Using Both Together
These tools are not mutually exclusive. Because DQX operates on standard PySpark DataFrames, you can execute DQX checks inside a Lakeflow pipeline table definition:
import dlt
from databricks.labs.dqx.engine import DQEngine
from databricks.sdk import WorkspaceClient
@dlt.table(comment="Silver orders validated via ODCS contract rules within Lakeflow")
def silver_orders_validated():
raw_df = dlt.read_stream("bronze_orders")
engine = DQEngine(WorkspaceClient())
return engine.apply_checks_by_metadata(raw_df, row_checks)
This pattern pairs Lakeflow compute orchestration with DQX contract-driven rule generation.
Operational Realities and Team Responsibilities
Contract-driven quality introduces five production responsibilities:
1. Version Pinning and Regression Control
Databricks Labs projects iterate rapidly. Pin specific minor versions in Databricks Asset Bundles (DABs) or cluster init scripts. Never leave databricks-labs-dqx unpinned in automated jobs.
2. Generated Rule Review
Treat generated rules as compiled code. In CI/CD pipelines, inspect generated rule YAML before applying it to production tables. Ensure physical types and default criticalities align with operational requirements.
3. Quarantine Lifecycle and Ownership
A quarantine table without an owner becomes a data graveyard. Establish clear ownership: who gets alerted on threshold spikes, how long dead-letter records live, and who triages root causes.
4. Replay Patterns and Idempotency
When quarantined records are fixed, pipelines must support idempotent reprocessing. Ensure Delta merge keys prevent replayed rows from generating duplicate records downstream.
5. Validation Compute Overhead
Every quality check adds Catalyst expressions. Simple null and range checks add negligible overhead, but complex regex and array operations increase CPU demand. Benchmark check performance against realistic production volumes.
A 30-Day Adoption Roadmap
Adopt DQX incrementally rather than rewriting existing pipelines all at once:
Week 1 (Annotation Mode): Target one table. Run DQX via apply_checks_by_metadata to append _errors and _warnings without dropping records. Review real-world failure patterns.
Week 2 (Quarantine & Metrics): Switch to apply_checks_by_metadata_and_split. Route invalid records to quarantine and persist summary metrics to Delta tables for Lakeview dashboards.
Week 3 (Governed Metadata): Move rules from notebook code into versioned YAML files in a Unity Catalog Volume. Add rule validation to CI/CD.
Week 4 (Contract-Driven Enforcement): Replace standalone YAML rules with an ODCS contract. Use generate_rules_from_contract in production and test quarantine replay procedures.
The Decision Framework: Who Should Adopt DQX?
Should your team adopt DQX and contract-driven quality checks?
Adopt This Pattern If:
Multiple producer teams deliver data into your platform and you need ODCS contracts to govern the interface.
You run standard Databricks Jobs and require row-level quarantine triage.
Your platform prioritizes declarative specifications over hardcoded pipeline logic.
Your team can manage an open-source Databricks Labs dependency.
Skip This Pattern If:
You build small pipelines where the same engineers own ingestion and reporting.
Your entire workload runs in Lakeflow and native expectations satisfy requirements.
Enterprise governance forbids tools without vendor SLAs.
Five lines of plain PySpark satisfy your checks.
The Data With Direction Self-Audit
Run through this five-point audit with your team before writing validation scripts:
Failure Transparency: If invalid records arrive today, does your pipeline fail, drop them silently, or route them to quarantine?
Contract Synchronization: Where do requirements live? If definitions live in documentation while checks live in Python, how do you prevent drift?
Violation Attribution: When an analyst questions a metric, can you trace that record back to specific checks it passed or failed at ingestion?
Tool Coupling: Are your rules locked into one runtime, or can they run across any PySpark batch or streaming job?
Remediation Path: If twenty thousand records land in quarantine tomorrow, do you have a tested, idempotent replay procedure?
In twenty-six years moving from DTS and SSIS to modern Lakehouses and AI assistants like Cursor and OmniGent, the tools change constantly, but the foundational principle never does: code is ephemeral, but contracts govern the boundary between code and cash.
Make the data decision you can defend six months from now.
Strategy over syntax.
Founder & Principal Data Architect, Gambill Data, LLC
Ready to Build Defensible Data Architecture?
Enterprise Advisory: Auditing Databricks platform spend, eliminating DBU compute bleed, or establishing production data contracts? Book an Architecture Strategy Call.
Sources and further reading
- Data Contract Quality Rules GenerationDatabricks Labs DQX
- Quality Checks DefinitionDatabricks Labs DQX
- Quality ChecksDatabricks Labs DQX
Related decision support
Data engineering consulting
Get an independent review of contract ownership, validation rules, quarantine routing, and operational reliability before scaling the pattern.
Review the service