Skip to main content

Jev for Analytics: 5 Use Cases on Real Data

2026/10/08

Mehdi Ouazza shows how to use MotherDuck's prompt_jev() to turn text columns into columns you can GROUP BY. He starts with 100,000 public consumer complaints, then walks through five use cases on real data: AI agent logs, Hacker News sentiment, sales call signals, cloud job ads, and remote job ads. Each one comes down to a short SQL query over a text column.

What Jev is and where it fits

Jev is a model from TypeSafe that returns a fixed answer instead of free text. You give it a question and a list of options, and it returns one option with a confidence score. Mehdi frames it with the System 1 and System 2 split from Thinking, Fast and Slow. An LLM handles slow, open-ended work. Jev handles quick decisions: classifying, filtering, ranking, and labeling. Compared with training a classic ML classifier, there is no labeling, training, or deployment to maintain, because you ask in plain English. In his comparison on 100,000 short news articles, prompt_jev() ran about 30x faster than GPT-5 nano through MotherDuck's prompt() function, at a much lower cost. He suggests running your own benchmark.

Classifying 100,000 complaints

The hands-on demo uses a public MotherDuck share of CFPB consumer complaints, where each row has a free-text narrative. Mehdi asks one question, what is the customer asking for, with three options: refund, explanation, or other. In the MotherDuck UI, three rows came back in under three seconds with a label and confidence score each. Scoring all 100,000 rows took a bit more than a minute. He saves the labels to a table, keeps rows at 80% confidence or higher, and queries it with plain SQL. The top-requests query runs in about 100 milliseconds. Explanation came out on top, with correcting or removing credit report entries and refunds or charge reversals also in the top five. Rows with low confidence go to manual review: one complaint about unauthorized credit inquiries came back labeled identity theft at 0.53.

Five use cases on real data

AI agent logs. MotherDuck traces the demo agent at motherduck.com/try, and each trace is one long text column. Mehdi asks Jev whether the task completed, whether it failed, which error code appeared, how many tool calls it took, and where the agent wasted calls, such as querying the schema three times when one SQL query would do.

Hacker News social listening. The full Hacker News history, from 2006 to the day before, sits in MotherDuck. Mehdi asks which AI company a post is about, adds context to cut false positives from names like Gemini (also an internet protocol) and Claude (also a first name), then asks for sentiment and the main complaint. In the resulting Dive, OpenAI leads the conversation and Anthropic takes a significant share, though he notes that is not market share. The top complaint about Anthropic is coding and developer tools, and answer quality is a leading complaint about OpenAI and Google.

Sales call signals. Recorded calls are long text columns too. Mehdi asks questions like how large the prospect's data and team are, answered in buckets, so calls can be sorted by customer size without needing exact numbers.

Cloud job ads. For each job ad, one question: which cloud is it primarily built on, with options for AWS, Azure, GCP, multi-cloud, not clear, or mentioned only in passing. The Dive shows Europe leaning Azure while the US varies by state, with a Google Cloud cluster in Arkansas. Hiring location often tracks company headquarters.

Remote job ads. "Is this job remote?" returns a Boolean with a confidence score. "Where does the person work?" returns remote, hybrid, on-site, or not stated.

Tips for your first experiment

Give Jev an escape option such as "not clear" or "other." Without one it has to pick a wrong label, and your counts skew. Put context in the question, such as which companies count as AI labs, to cut false positives.

You can ask several questions about the same column in one query. Per Mehdi, TypeSafe runs them in parallel, so cost and runtime barely grow. Filter on the confidence column (he uses 80% or higher) and send the rest to manual review.

If you do not know your categories yet, ask an LLM to propose labels from a sample of rows, review them by hand, then have prompt_jev() score the full table. The prompt_jev() docs cover the syntax.

FAQS

Jev returns one option from a list you provide, plus a confidence score, instead of free text. That makes it fast and cheap enough to run over a whole table. In Mehdi's comparison, prompt_jev() ran about 30x faster than GPT-5 nano on 100,000 short news articles. Use an LLM for open-ended work and Jev for classifying, filtering, ranking, and labeling.

prompt_jev() is a MotherDuck feature. A community DuckDB extension can call Jev, but you need your own Jev account. From a DuckDB install on Windows or Linux, you can attach md: and call prompt_jev(), or call Jev from a Python function with your own API key.

Pricing is based on input tokens, and the AI function section of the MotherDuck pricing docs has the details. Asking several questions about the same text column in one query stays relatively cheap, according to TypeSafe.

It can. Every result includes a confidence score, so filter on it (Mehdi uses 80% or higher) and review the rest by hand. You can also send low-confidence rows to an LLM as a fallback, as Jacob Matson suggested in chat.

Choice picks one option from a list you provide. The yes/no mode (noul) returns a Boolean with a confidence score. Score returns a value you can use in ORDER BY. The docs also cover batching, and the syntax is in the prompt_jev() docs.

Related Videos

"Virtual Workshop: Build a Data Agent in 60 Minutes" video thumbnail

51:58

2026-10-01

Build a Data Agent in 60 Minutes

Jacob Matson walks through a working data chat agent he built: MCP tools, an agentic loop, a tuned system prompt, hard-coded read-only guardrails, and telemetry. He demos it live against a real dataset, walks through the actual codebase, then shows the same backend re-platformed into a Slack bot.

Stream

MotherDuck Features

AI ML and LLMs

Data Pipelines

"We Classified 100,000 Rows in 40 Seconds: Introducing prompt_jev()" video thumbnail

49:02

2026-09-29

prompt_jev(): SQL-Native Text Classification

MotherDuck's new prompt_jev() function runs text classification directly from SQL using Jev, a decision-focused model from typesafe.ai. In a live demo, Jacob Matson, Dumky de Wilde, and Hamilton Ulmer showed it filtering, disambiguating, and re-ranking real data without leaving SQL. The number they kept coming back to: 100,000 rows classified at 89% accuracy in 40 seconds, compared to 32 minutes and $37.58 for a comparable LLM call.

Stream

MotherDuck Features

AI ML and LLMs

Data Pipelines