---
title: "Build a Data Agent in 60 Minutes | MotherDuck | MotherDuck"
description: "See how Jacob Matson architected a working data chat agent: MCP tools, an agentic loop, guardrails, and telemetry, demoed live and re-platformed into a Slack bot."
canonical: "https://motherduck.com/videos/build-a-data-agent-workshop/"
---

[BACK TO VIDEOS](/videos/)

[Stream](/videos/?category=Stream#videos-and-webinars-library)[MotherDuck Features](/videos/?category=MotherDuck%20Features#videos-and-webinars-library)[AI ML and LLMs](/videos/?category=AI%20ML%20and%20LLMs#videos-and-webinars-library)[Data Pipelines](/videos/?category=Data%20Pipelines#videos-and-webinars-library)

# Virtual Workshop: Build a Data Agent in 60 Minutes

2026/10/01

Jacob Matson walks through a working data chat agent he built: MCP tools, an agentic loop, a tuned system prompt, hard-coded read-only guardrails, and telemetry. He demos it live against a real dataset, walks through the actual codebase and PR history, then shows the same backend re-platformed into a Slack bot.

## Getting access to data

An agent needs four things: data, a credential to reach it (read-only, so nobody's waiting on a lock or worried about a stray delete), MCP tools so the model can actually query the database instead of just talking about it, and somewhere to run it. The demo, Data Chat Mini, is a Next.js app on Vercel, deployed with a single `vercel --prod`. The model is Gemini 3 Flash via OpenRouter, chosen for speed and cost, not leaderboard position. Figuring out what a user wants and writing narrow-scope SQL doesn't need a frontier reasoning model.

## Making it behave

The loop is a plain while-loop with a hard cap of 40 iterations so a complex question can't spin forever. The system prompt tells the model to explore the schema instead of guessing, flags DuckDB-vs-Postgres quirks it was tripping on, and bans raw HTML output in favor of a strict JSON chart schema. That schema is enforced through mviz, a small charting library Jacob built because every LLM renders visualizations differently. Claude does it inline now, GPT does something else, and so on. The guardrails are read-only access only, a tool allow-list instead of the full MCP surface, and full query logging. Logging is off by default in the public demo but flipped on with a one-line change in production. Build trust by showing your work, not by telling people to trust it.

## The live demo

Jacob queried an NBA dataset against the already-running app and got charts back in seconds, walked through the actual pull requests behind it, then showed the same backend running as a Slack bot (Quackbot, running on Kimi via Modal) he'd built ahead of the stream. Once the MCP/loop/prompt pattern exists, wiring it to a new chat surface is fast. Chat history and saved context stay local to the browser in this demo, so it never has to hold or be accountable for anyone's data. [MotherDuck Guides](https://motherduck.com/docs/key-tasks/guides/) is the recommended path when you want context shared across users in production. Full source for Data Chat Mini (and Quackbot, the Slack-bot version) is public at [github.com/motherduckdb/labs/tree/main/projects/data-chat-mini](https://github.com/motherduckdb/labs/tree/main/projects/data-chat-mini).

## Why SQL beats a semantic layer here

MotherDuck's own research found that adding a non-SQL query language like DAX, MDX, or similar semantic-layer DSLs makes an agent slower, more expensive, and less accurate than just writing SQL. There are decades of SQL in the training corpora these models learned from, and nowhere near as much in any given semantic-layer language. Jacob's take: a semantic layer is really just an extremely overfit context layer. MotherDuck Guides is where that business context, the joins and metric definitions, actually belongs, separate from execution and display.

TABLE OF CONTENTS

- Getting access to data
- Making it behave
- The live demo
- Why SQL beats a semantic layer here

Start using MotherDuck now!

[Try 7 Days Free](https://auth.motherduck.com/authorize?app_source=web&response_type=code&client_id=bza3KWQpxRAFlTlRFXUo29AOg9xD7zcp&redirect_uri=https%3A%2F%2Fapp.motherduck.com%2F&state=STATE&auth_flow=signup&screen_hint=signup&ext-ph_distinct_id=2ec145c7-df84-4f4d-8100-0aac77496082)

## FAQS

### Does this agent have write access to my database?

No. Everything runs read-only. MotherDuck has separate read-write and read-only query endpoints, and you can limit access further with the role you grant. Jacob's advice: default to read-only for agentic work unless you have a specific reason not to.

### Can users upload their own files to chat with?

Not in this demo. MCP doesn't have a file-upload interface, so you'd need to build that separately. The agent can already read from S3 or attach other MotherDuck databases on its own.

### Is chat memory or saved context shared across users?

No, it's local to each user's browser. That's deliberate — the public demo never has to hold or be accountable for anyone's data. If you want shared context across a team in production, MotherDuck Guides is the recommended path.

### Why Gemini 3 Flash instead of a more powerful model?

Narrow-scope SQL generation doesn't need a frontier reasoning model. It needs to be fast and cheap. Jacob swaps models through OpenRouter and expects similarly lightweight models to work fine. The model matters less than good schema context and a tight system prompt.

### How is the Slack bot (Quackbot) built on top of the same system?

A Slack app exposes an HTTP endpoint. A small server on Modal or Fly.io relays messages between Slack and the LLM, reusing the same MCP tools as the web demo. It's a thin integration layer, not a rebuild — which is the whole point of designing the agent around MCP.

## Related Videos

[49:02](/videos/prompt-jev-sql-text-classification/)[2026-09-29](/videos/prompt-jev-sql-text-classification/)

### [prompt_jev(): SQL-Native Text Classification](/videos/prompt-jev-sql-text-classification/)

MotherDuck's new prompt_jev() function runs text classification directly from SQL using Jev, a decision-focused model from typesafe.ai. In a live demo, Jacob Matson, Dumky de Wilde, and Hamilton Ulmer showed it filtering, disambiguating, and re-ranking real data without leaving SQL. The number they kept coming back to: 100,000 rows classified at 89% accuracy in 40 seconds, compared to 32 minutes and $37.58 for a comparable LLM call.

Stream

MotherDuck Features

AI ML and LLMs

Data Pipelines

[37:22](/videos/motherduck-cli-vs-mcp/)[2026-09-16](/videos/motherduck-cli-vs-mcp/)

### [MotherDuck CLI: query, pipelines, dashboards for your agents and CI](/videos/motherduck-cli-vs-mcp/)

One CLI to query MotherDuck, publish Python pipelines, and host dashboards from the terminal. See when it beats the MCP server, and when it doesn't.

Stream

MotherDuck Features

AI, ML and LLMs

Data Pipelines

[47:43](/videos/spark-iceberg-motherduck-pipeline/)[2026-09-10](/videos/spark-iceberg-motherduck-pipeline/)

### [Data takes Flight: Transforming Data with Spark and MotherDuck via Iceberg](/videos/spark-iceberg-motherduck-pipeline/)

See a live cross-engine pipeline: Spark transcribes audio and writes to Iceberg, then MotherDuck attaches that catalog directly as a queryable database.

Stream

Data Pipelines

dbt

Ecosystem

[View all](/videos/)