Skip to main content

We Classified 100,000 Rows in 40 Seconds: Introducing prompt_jev()

2026/09/29

MotherDuck's new prompt_jev() function runs text classification directly from SQL using Jev, a decision-focused model from typesafe.ai. In a live demo, Jacob Matson, Dumky de Wilde, and Hamilton Ulmer showed it filtering, disambiguating, and re-ranking real data without leaving SQL. The number they kept coming back to: 100,000 rows classified at 89% accuracy in 40 seconds, compared to 32 minutes and $37.58 for a comparable LLM call.

The gap between LLM calls and trained classifiers

If you need to pull structure out of messy text — support tickets, sales calls, product descriptions — you've historically had two choices. Prompt an LLM for every row, which is flexible but slow and expensive at volume. Or train a classifier, which is fast and cheap to run but needs a hand-labeled training set and ongoing ML work. Jev does something different: instead of generating a paragraph of reasoning, it returns a typed decision. That gets it to classifier speed without giving up the convenience of a plain-language prompt.

What prompt_jev() actually returns

The function supports three output types. noul answers a yes/no question and returns a confidence score between 0 and 1. choice picks from a list you define, up to 255 categories. score gives a categorical rating, useful for triaging support-ticket urgency or re-ranking search results. All three read from the same cached text field, so asking multiple questions about a single row doesn't multiply your calls.

Live demo patterns from a real job-postings table

Dumky de Wilde ran prompt_jev() against a table of job descriptions. Semantic filtering caught "a remote job based in London" in ways a plain ILIKE '%London%' never could. A choice classification disambiguated the word "Go," separating the programming language from governance roles and go-to-market postings — exactly the kind of thing keyword matching falls apart on. A staged pipeline used Jev cheaply to flag which rows were worth sending to a more expensive model, cutting a 100-row batch down to 33 rows that actually needed the pricier extraction call. Pairing MotherDuck's embedding function for initial similarity search with Jev as a re-ranker let them discard results that were technically close in vector space but not actually relevant.

The benchmark, and what's next

On 100,000 rows from the AG News dataset, Jev hit 89% accuracy in 40 seconds, about 2,500 rows per second. GPT-5.6 Terra scored 88% accuracy in roughly 32 minutes at $37.58 retail; Jev cost $0.50. Hamilton Ulmer said preliminary numbers put Jev at roughly 50x faster and about 1% of the cost for comparable classification work. The team shipped a further 2x speed improvement the same week as the stream. Eight days after launch, customers were already running it in production: tagging product and event data, scoring support tickets, analyzing sales calls, monitoring agent traces and developer token usage. It runs anywhere you'd write SQL, including on a schedule inside a MotherDuck Flight.

Data handling: Jev runs on typesafe.ai's servers under a zero-data-retention policy, not on MotherDuck's infrastructure. The function is in preview, currently available in US East 1 and US West 2, with EU support targeted for October 6 pending GDPR sign-off. HIPAA and similar compliance work is still in progress — confirm your own requirements before sending regulated data through it.

FAQS

Yes. prompt_jev() calls out to typesafe.ai's hosted Jev model rather than running inside MotherDuck. Typesafe holds the data only long enough to generate a response, under a zero-retention policy, and MotherDuck says it doesn't believe they train on customer data. The function is in preview, so treat it like any new third-party integration and confirm it fits your compliance requirements, including HIPAA, before sending regulated data through it.

noul asks a yes/no question and returns a confidence score between 0 and 1. choice picks from a defined list of up to 255 categories. score gives a categorical rating — useful for ranking support-ticket urgency as low, medium, or high, or re-ranking search results. All three read from the same cached text field, so asking several questions about one row doesn't multiply your calls.

Not the way you'd fine-tune a traditional classifier. Jev has a 32,000-token context window, so the team's guidance is to pack more context into the prompt itself — reference examples, category definitions, business rules — rather than expecting a training step. If you have a large, well-labeled proprietary dataset, a fine-tuned BERT may still outperform it.

If you already have labeled data unique to your business, or a high-stakes classification task like medical imaging, a model you train yourself will do better. prompt_jev() makes sense when "very good" is good enough and you don't want to build and maintain a custom model.

Yes. It's a regular SQL function, so it works inside a MotherDuck Flight like any other query. Customers are already running it continuously — classifying sales calls as they come in, for example.

Related Videos