Key takeaways
  • A data quality framework is an operating model: defined dimensions, automated rules in the pipeline, quality scores, and named owners. A one-off cleanup is not a framework.
  • Quality is asserted, observability is discovered. Quality checks verify rules you wrote; observability detects anomalies you did not think to write a rule for. You need both.
  • Push checks upstream. A rule enforced at ingestion catches a bad batch before it reaches fifty dashboards; the same rule at the report catches it after the damage.
  • A quality score per dataset, visible in the catalog, turns an abstract goal into a number owners are accountable for and consumers can act on.

Why cleanups fail and frameworks last

A one-time cleanup treats the symptom. It corrects the values in a dataset without changing the process that produced the bad values, so the defects reappear at the source rate. Six weeks later the dataset is dirty again and the exercise repeats, which trains the organization to believe quality is unachievable.

A data quality framework treats the process. It defines what good means, enforces it automatically on every run, measures the result, and assigns someone to own it. The difference is not effort, it is where the effort goes: into the pipeline and the operating model rather than into a periodic firefight.

The dimensions, briefly

Quality has to be measurable before it can be managed. Six dimensions cover most of what matters: completeness (are required values present), accuracy (do they reflect reality), consistency (do systems agree), timeliness (are they fresh enough), uniqueness (is each entity present once), and validity (do values obey the rules). Each becomes one or more automated tests. The point of naming them is not taxonomy; it is that "improve data quality" is not actionable while "raise completeness of the customer country field above 99 percent" is.

Worked example: invoice data from ERP to reporting

Consider a reporting dataset assembled from supplier invoices, purchase orders and payment records in an ERP. A dashboard uses it to show unpaid invoices. The framework starts with the business decision the report supports: a finance team needs to distinguish an unpaid invoice from one that is missing a payment update. The following rules are illustrative; their thresholds and exception paths must be agreed with the finance owner before implementation.

Check and dimensionWhere it runsWhat happens on failure
Invoice ID, supplier ID and invoice date are present (completeness)At ingestion from the ERPQuarantine the record and send the missing fields to the source-system owner.
The same supplier and invoice ID do not appear twice (uniqueness)Before the reporting table is updatedHold the duplicate for review; do not silently drop either record.
The sum of invoice lines matches the invoice total under the agreed currency and tax rules (consistency)During transformationFlag the invoice as unreconciled and exclude it from an approved-total metric.
Payment status arrives within the reporting window agreed with finance (timeliness)Before the dashboard refreshMark the dataset as stale and show the last successful refresh to consumers.

This is a rule register, not just a list of tests. Each row also needs a named owner, a severity, an agreed threshold and a route for correcting the record at its source. The failed-record count and its cause should remain visible; otherwise a clean-looking dashboard may simply be hiding excluded invoices. The same approach applies to customer master data and event streams, but the rules and owners change with the business decision.

Where checks belong: push upstream

The highest-leverage decision in a quality framework is where the checks run. The instinct is to validate at the end, in the report, because that is where problems are noticed. That is the most expensive place to catch them.

A check at ingestion catches a bad batch before it propagates into every downstream table and dashboard. A check at the report catches it after it has already fed decisions. The same rule, moved upstream, changes from damage assessment to damage prevention. Tools like Great Expectations, dbt tests and Soda make it practical to assert expectations at each stage of the pipeline, and the framework should place the most important checks as close to the source as possible. Data contracts take this to its conclusion, moving the guarantee to the producing system so violations are caught at the boundary.

Scoring: making quality a number

An abstract commitment to quality changes nothing. A quality score does. The framework assigns each dataset a score, computed from its checks across the dimensions, and surfaces it in the data catalog next to ownership and lineage.

That number does two jobs. For the owner, it makes a regression visible and assigns it to someone who can act. For the consumer, it helps compare datasets before building on them. Publish the component measures alongside the score: a high average must not conceal a failed critical reconciliation rule or a stale payment feed. The score is a summary for triage, not permission to ignore failed checks.

Roles: quality needs owners

Technology enforces rules but does not decide them. Whether a value is correct in business terms is a judgment only the domain knows, which is why a quality framework is also a set of roles.

  • Data owners are accountable for the quality of the datasets their domain produces, including the score.
  • Data stewards define the rules, adjudicate edge cases and maintain the business meaning behind the checks.
  • Data consumers report issues and set the quality expectations their use cases require.

Without these roles, measured defects have no home: the score drops and no one is responsible for raising it. This is the same accountability structure that underpins data governance, applied specifically to quality.

For the invoice example, finance defines which discrepancies block a report, the ERP owner corrects source records, and the data team maintains the checks and the exception queue. Record those responsibilities in the catalog. The broader ownership and access model belongs in a data governance framework.

Quality versus observability

Data quality and data observability are often sold together and are not the same thing, and understanding the difference prevents buying one when you need the other.

Quality is assertion. You know completeness matters for a field, so you write a rule that checks it, and the rule passes or fails. It catches the failures you anticipated.

Observability is discovery. It monitors the data's own behavior (row counts, distributions, freshness, schema) and flags anomalies you did not write a rule for: a table that suddenly has half its usual rows, a column whose values shifted. It catches the failures you did not anticipate.

A mature framework uses both: explicit quality rules for the known requirements, and observability for the unknown unknowns. Relying only on rules means you are blind to novel failures; relying only on observability means you have no defined standard to hold data to.

Failure patterns

From the field, the patterns that keep quality out of reach:

1. Cleanup instead of framework. Fixing values without fixing the process, so defects return at the source rate.

2. Checks at the end. Validation only in the final report, catching problems after they have fed decisions.

3. No score. Quality as a slogan with no number, so it is argued about rather than managed.

4. No owners. Measured defects with no one accountable to fix them.

5. Observability mistaken for quality. Anomaly detection bought as if it defined a standard, when it only detects deviation from the data's own past.

Talk through your data quality framework

DNA Solutions helps European enterprises define quality rules, place checks in data pipelines and make failures visible to the people responsible for resolving them. See our Data & Analytics service or talk to us about a dataset and the decision it supports.

DNA Solutions expertise for your project

DNA Solutions unifies disparate data sources into central, high-performance querying environments. We build cloud data lakes, warehouses and streaming pipelines so your teams get real-time access to production information, without vendor lock-in.

Data & Analytics

Related services: Data & Analytics