Dependable data, clearly explained.

Data engineering & analytics

Flores DataCore, front page

Insights Data quality

4 min read October 2026

Data quality checks that matter

Start with the checks that catch real incidents: freshness, volume, keys, and the rules your business already states.

By the Flores DataCore insights desk

Hand checking off items on a handwritten list
Fig. 1A short list that runs every day beats a long list that runs once.

Data quality programs often begin with a long list of dimensions and a scoring spreadsheet, and stall there. A faster route is to start from the incidents you have already had. Most of them fall into a few patterns, and a few cheap checks catch the bulk of them.

Five checks that catch most incidents

Freshness. When was this table last updated, and how current do its readers need it? A daily sales table might warn after 26 hours and fail after 36. Tools such as dbt can run this check against a timestamp column, with separate warning and error thresholds.1

Volume. Compare each load's row count with its usual range. A load that brings in zero rows, or ten times the normal number, is almost always a problem upstream. Allow for seasonality: Mondays, month ends, and holidays each have their own normal.

Uniqueness. Primary keys must be unique. A duplicated order row quietly doubles that order's revenue in every report built on top of it.

Missing values. Fields a process depends on, such as the customer on an order, must be present. Check the ones that matter, not every column.

Business rules. Status values from an agreed list, amounts that are never negative, and every order pointing to a customer that exists. These rules are usually already written down somewhere; the check makes them enforceable.2

One more cheap check belongs at the front door: alert when a source adds, drops, or renames a column. A surprising share of breakages start there.

Five checks, what they catch, and where to run them
CheckCatchesRun it
FreshnessStalled loads, stale dashboardsOn every source table
VolumePartial loads, runaway duplicatesAfter each load
UniquenessDouble countingOn every primary key
Missing valuesOrphaned and partial recordsOn required fields
Business rulesBad statuses, impossible amounts, broken linksIn core models and marts
Fig. 2Five checks, what they catch, and where they belong.

Put checks where data changes hands

You do not need a test on every column of every table. Put checks at the boundaries: when data lands from a source, after staging cleans it, and in the marts that people read. A problem caught at the boundary has not spread yet.

One more check pays for itself on critical metrics: reconcile a daily total against the source system. If the warehouse says 1,204 orders yesterday and the order system says 1,211, someone should know before the morning meeting.

Decide what a failure does

Every check needs a severity. A warning tells the owner. An error stops the models that depend on the failing table, so a bad load does not reach a dashboard.

Every alert should answer four questions without a meeting: what failed, what it affects, who owns it, and what to try first. Lineage answers the second. A runbook answers the last. Send alerts to a named owner, not to a shared channel that nobody reads.

Write the expected response next to each check. The next person on call should know whether to fix the data, rerun the load, or fix the test.

Check the checks

Track two numbers each month: incidents caught by checks, and incidents reported by people. Over time, the second should fall. Watch for noisy tests, too. A test that fails every week and gets ignored is worse than no test, because it teaches everyone to ignore alerts. Fix it or delete it.

Write tests from incidents

After every incident, add the check that would have caught it. Within a few months, your tests describe the failures your data actually has, rather than a generic list from a vendor's checklist.

A starting list for this month

  1. Freshness on every source table that a dashboard depends on.
  2. Unique and not-null tests on every primary key.
  3. Accepted values on status and type fields.
  4. Relationship tests between facts and dimensions.
  5. One reconciliation against the source for each metric leadership uses.

None of this requires a new platform. It requires deciding what correct means and writing it down where a machine can check it every day.

  1. 1

    dbt documentation, "freshness": a source can declare warn_after and error_after thresholds, each with a count and a period, measured against a loaded_at_field timestamp. Back

  2. 2

    dbt's four built-in generic tests map closely to these checks: unique, not_null, accepted_values, and relationships. Each is a query that returns the failing rows. Back

B1More from the desk