Interactive artifact
Semantic dbt tests: data quality checks written in English
Four dbt tests written as plain-English sentences, judged row by row by a small model and scored against a hidden answer key and a regex baseline. Move the threshold, switch the request layout, and read any row.
- Type
- Data tool
- Drawn
- Runs
- In your browser, nothing to install
- Tags
- dbt · data-quality · analytics-engineering · duckdb · interactive
Detail A · Reading the drawing
What it shows
Structural dbt tests (not_null, unique, accepted_values, relationships) check shape. They pass happily on a customer called 'asdf asdf', a return coded late whose comment says the box arrived crushed, or a support ticket with someone's IBAN in it. A semantic test states what's wrong in a sentence, and a model returns the probability that each row fails it.
Every dot on the sheet is a real probability from a logged live run of TypeSafe's Jev over 1,300 synthetic rows: 500 customers, 200 returns, 400 reviews and 200 tickets. The answer key marks defects and hard negatives (rows built to look wrong but aren't). The regex baseline is the same project's hand-written keyword tests, re-run over the same rows. Nothing is judged in your browser.
Detail B · Procedure
How to use it
- Start from a preset: the gate run, the packing collapse, the nested fix, a threshold set too low, or the honest misses.
- Pick one of the four tests and read it as written in schema.yml.
- Drag the threshold and watch precision, recall and the gate change for Jev against the regex baseline.
- Switch the request layout to see what batching 8 or 32 rows into one request does to recall, and how the nested layout gets it back.
- Click a dot, or a row id under the scorecard, to read the row, its answer-key note, and its probability in every layout.
Detail C · Glossary
Key concepts
- Semantic test
- A data test whose condition is a sentence rather than SQL, e.g. 'the comment describes a different reason than reason_code', judged per row by a model.
- Hard negative
- A row that looks like a defect to a keyword rule but isn't: the surname Test, the shop's own phone number in a ticket, a sarcastic 1-star review.
- Precision and recall
- Precision is the share of flagged rows that are real defects; recall is the share of real defects that got flagged.
- Threshold
- The probability at or above which a row is stored as a failure. Lower catches more defects and raises more false alarms.
- Request packing
- Sending several rows in one model request. Faster and cheaper, but if rows share the request's state they distract each other.