Typed judgments inside data tools: four Jev demos

·8 min read·

Most LLM features in data tools generate text: a summary, a SQL string, a chart spec you then parse and hope about. That's the wrong shape for most of what data engineers actually need, which is a judgment. Does this review ask for a refund? Does this comment match its reason code? Which of these five valid charts answers the question? Those are yes/no, pick-one and how-much questions, and the answer belongs in a column.

TypeSafe↗'s Jev is built for exactly that. They call it a System One model: you send a state and a set of questions, and you get back typed values (a probability, a choice from a list, a position on a rubric) rather than prose. According to TypeSafe's own docs (as of 2026-09-18) it's also cheap enough to run per row: $0.042 per million input tokens, output free, 1,200 requests a minute.

So I built four small demos that put it inside tools data people already use. All four are public in jev-demos↗, MIT-licensed, and each runs offline in a clearly labelled simulated mode if you don't have a key.

The pattern: code proposes, the model judges, code decides#

The four demos share one architecture, and it's the part I'd reuse regardless of the model.

That split keeps the failure modes small. A bad judgment makes one row or one chart wrong; it can't produce invalid SQL or a broken spec, because the model never writes either. It also makes everything cacheable: every demo caches judgments by content (model, question, input), so re-running a query costs nothing and returns the same answer.

The other shared rule is about honesty. Every demo has a SIMULATED mode that is labelled loudly everywhere it shows up, and every number in the READMEs, in these posts and in the artifacts comes from a logged live run. If a run wasn't logged, I don't quote it.

04 · Semantic dbt tests#

The most complete of the four, and the one to start with. Four dbt tests written as English sentences (fails_if: "The customer's comment describes a different main reason than reason_code") sit next to not_null and accepted_values in schema.yml. They catch rows that pass every structural test but are still wrong: a customer called Mickey Mouse, a return coded late whose comment says the box arrived crushed, a support ticket with an IBAN in it.

Against a hidden answer key, the official live run scored precision and recall of 0.96–1.00 on all four tests. A hand-written regex baseline scored between 0.39 and 0.71 on precision. The whole run was 1,300 judgments for $0.017. The most useful finding was a failure: batching rows into one request quietly dropped recall on the two-column test from 1.00 to 0.42. Moving each record into its own question fixed it, at 35 requests instead of 1,057.

Read it: Semantic dbt tests: data quality checks in plain English, or play with the scorecard, which replays every logged probability.

02 · semsql: semantic SQL in DuckDB#

Plain-English judgments as DuckDB functions, so they filter, sort and aggregate like any other column:

WITH candidates AS (SELECT * FROM reviews WHERE stars <= 2)
SELECT id, body, jev_score(body, 'anger') AS anger
FROM candidates
WHERE jev_noul(body, 'The customer is asking for their money back') > 0.8
ORDER BY anger DESC LIMIT 10;

jev_noul returns a probability, jev_score a position on a rubric defined in TOML, jev_choice one of a list of options. On 10,000 synthetic reviews, the keyword filter ILIKE '%refund%' finds 780 rows. The semantic query finds 913 refund requests among the low-star reviews, and 595 of those never contain the word "refund". A keyword filter can't find them; the English sentence can.

Two engineering details carry over to any per-row model call. First, put the model predicate after the cheap filters (the stars <= 2 CTE above), because it's the slowest and priciest part of the query. Second, measure packing before you trust it: semsql eval-pack scores a sample both ways. At 16 rows per request, 0.5% of rows flipped across the 0.8 threshold on a 200-row sample. The full demo query (about 4,400 rows at pack=16) took 19.7 s and $0.022. Re-running it cost 0.04 s and $0.0000, all cache hits.

The rubric lesson is the one I'd pass on: rubric levels should describe situations, not degrees. "Open hostility: insults, all-caps shouting, or explicit threats" gives the model something to match. "Very angry (4/5)" doesn't.

03 · VISUALIZE: the query picks its own chart#

A SQL query ends in VISUALIZE 'how are regions trending'. Code profiles the result and proposes only the charts that are valid for its shape. Each chart type is a rule that describes its candidate in a fixed grammar: form, measure, breakdown, what it reveals. Jev scores every candidate against the intent in one batched request, and code ranks the scores and builds the Vega-Lite spec. For "how are regions trending", a multi-line chart by region scores 1.00, and a single line of total revenue scores 0.57: a partial answer, rated as one.

On the latest run over a 30-case golden set, Jev's top pick was accepted 97% of the time, against 33% for a rules-only "first valid chart" baseline. Warm, run-to-chart-painted takes about 0.25–0.45 s. The caveat matters, so I'll put it next to the number: 20 of the 30 cases were drafted with the AI assistant that helped build the system, and the eval log says to read 97% as "the approach works and the misses are legible", not as a benchmark. What I like most is how legible the misses are. The one run where a stacked bar fell out of the top three traced back to a single clause in its description.

01 · The LinkedIn Cringe-o-Meter#

The unserious one, and the one people actually share. Paste a LinkedIn post draft and eight judgments come back in parallel: humblebrag, engagement bait, fake vulnerability, buzzword density, broetry formatting, the "I'm humbled to announce" opener, toddler life lessons, and unsolicited hustle advice.

There's a real design choice hiding in it. Dimensions where degree matters (how humblebraggy?) use a four-level score. Dimensions where presence matters (is this broetry?) use a yes/no probability. The composite leans on the worst offences rather than averaging them out, because three maxed-out dimensions are cringe even when the other five are clean: 80% the weighted mean of the top values, 20% the overall mean. Weight sliders recompute it in the browser from the cached answers, with no new API call. There are no logged per-post scores for this one, so I'm not quoting any.

What these demos don't claim#

After the semsql demo went round on LinkedIn, someone asked, fairly, whether I'd benchmarked Jev against frontier LLMs on accuracy, latency and cost. I hadn't. From working with both, I expect Jev to be far cheaper and faster per judgment. But that's an expectation, not a measurement, and accuracy is the open question. Every baseline in these demos is a keyword, regex or rules baseline, because those are what you'd otherwise ship. None of them is a general-purpose LLM.

That benchmark is designed (same rows, interleaved requests, temperature 0, equal prompt effort, parse failures counted, prices looked up on the day) and not yet run. When it is, the results go here, whichever way they land. Until then I'm not claiming Jev beats a frontier model at anything.

Two more limits apply to all four demos. The data is synthetic, generated by seeded scripts, so the hard cases are the ones someone thought to plant. And they're demos, not packages: localhost, no auth, no deployment story.

Where to go next#

If you do analytics engineering, start with the semantic dbt tests post and its scorecard. If you're thinking about where models belong in a data stack in general, the architecture above is the transferable part. It's the same "code for the deterministic, model for the judgment" split I use everywhere, laid out in how I build AI-native. The dbt tests sit naturally next to column-level lineage as a data quality layer, and semsql runs on the same DuckDB I use for everything else.

Disclosure: I'm an independent user of Jev with no commercial relationship with TypeSafe. Numbers come from logged live runs; simulated-mode output is never quoted.

Related

How I build AI-native

Building AI-native isn't about writing code faster. It's about thinking in systems and moving your engineering effort to where it survives contact with production. Here's the method, and the work that proves it.

5 min read

Building agentic workflows with Claude Code

An LLM can do almost anything once. Getting it to do the right thing every time is an engineering problem. Here's the architecture I use, drawn from a skill that audits websites for GDPR end to end.

7 min read

Semantic dbt tests: data quality checks in plain English

not_null and accepted_values will happily pass a customer called Mickey Mouse, a return coded late whose comment says the box arrived crushed, and a ticket with someone's IBAN in it. So I wrote four dbt tests as English sentences and scored them against an answer key.

12 min read