Fig. 1 · Replay of live runs · Data quality

Semantic dbt tests, scored

Four dbt tests written as English sentences, judged row by row by TypeSafe's Jev and checked against a hidden answer key and a regex baseline. Every dot is a real, logged probability. Move the threshold, switch the request layout, and click any dot to read the row.

Rows
1,300
Tests
4
Model
jev-latest
Logged
2026-09-26
Presets
B · Jev probability per row
Loading logged runs…
defect (answer key) hard negative clean Jev wrong
Click a dot, or a row id below, to read the row and its probability in every layout.
C · Scorecard
D · Request layoutswhole run · 4 tests · 1,057 unique states

Requests, tokens, time and cost are for the whole run (all four tests), from the pack benchmark that calls Jev directly with the same states and questions dbt sends. Precision and recall are for the selected test at the current threshold. Click a layout to plot it.