Interactive artifact

Semantic dbt tests: data quality checks written in English

Four dbt tests written as plain-English sentences, judged row by row by a small model and scored against a hidden answer key and a regex baseline. Move the threshold, switch the request layout, and read any row.

Title block
Type
Data tool
Drawn
Runs
In your browser, nothing to install
Tags
dbt · data-quality · analytics-engineering · duckdb · interactive

Detail A · Reading the drawing

What it shows

Structural dbt tests (not_null, unique, accepted_values, relationships) check shape. They pass happily on a customer called 'asdf asdf', a return coded late whose comment says the box arrived crushed, or a support ticket with someone's IBAN in it. A semantic test states what's wrong in a sentence, and a model returns the probability that each row fails it.

Every dot on the sheet is a real probability from a logged live run of TypeSafe's Jev over 1,300 synthetic rows: 500 customers, 200 returns, 400 reviews and 200 tickets. The answer key marks defects and hard negatives (rows built to look wrong but aren't). The regex baseline is the same project's hand-written keyword tests, re-run over the same rows. Nothing is judged in your browser.

Detail B · Procedure

How to use it

  1. Start from a preset: the gate run, the packing collapse, the nested fix, a threshold set too low, or the honest misses.
  2. Pick one of the four tests and read it as written in schema.yml.
  3. Drag the threshold and watch precision, recall and the gate change for Jev against the regex baseline.
  4. Switch the request layout to see what batching 8 or 32 rows into one request does to recall, and how the nested layout gets it back.
  5. Click a dot, or a row id under the scorecard, to read the row, its answer-key note, and its probability in every layout.

Detail C · Glossary

Key concepts

Semantic test
A data test whose condition is a sentence rather than SQL, e.g. 'the comment describes a different reason than reason_code', judged per row by a model.
Hard negative
A row that looks like a defect to a keyword rule but isn't: the surname Test, the shop's own phone number in a ticket, a sarcastic 1-star review.
Precision and recall
Precision is the share of flagged rows that are real defects; recall is the share of real defects that got flagged.
Threshold
The probability at or above which a row is stored as a failure. Lower catches more defects and raises more false alarms.
Request packing
Sending several rows in one model request. Faster and cheaper, but if rows share the request's state they distract each other.

All artifacts →