Skip to main content

AI Needs Data, Data Needs AI: What You'll Learn at Data Outpost

AIAgentsSql

- 9 min read

Data Outpost: 2 days, November 4-5, 2026, San Francisco

AI is rebuilding how companies work, but we all know that slapping a chatbot on top won't cut it. Data systems are the way that even the most powerful frontier model knows anything about your business. The intersection of the two is filled with complexity: what does it mean when agents are "part of the team"? What does it take for an AI system to be trustworthy? How can we build insight-generating systems that scale?

Data Outpost is an AI and data conference where you'll learn from down-to-earth experts and visionaries alike. It's a small format event, so you get to actually talk to these folks and collaborate with them.

Workshops are on Wednesday, November 4th, and the main day of talks and panels is Thursday, November 5, 2026, both at Convene 100 Stockton in San Francisco.

Sound good? Register for Data Outpost here! Need more convincing? Read on!

We'll cover a quick tour of what you'll learn from our sessions plus my take about what I'm most excited about for each!

Read-Only Was Nice While It Lasted

Agents used to read your data and hand back answers. Now they act on it, and once they do, someone has to decide what they're trusted with.

Data + AI Product Leadership: The Data Stack After Agents

Tristan Handy
Stefanie Tignor
Jake Hannan
Leanne Fretz
Alexis Weill

Panel: Tristan Handy (President & Co-Founder, Fivetran + dbt Labs, moderator), Stefanie Tignor (Head of Data, Clay), Jake Hannan (VP, Data & Revenue Engineering, Sigma), Leanne Fretz (Head of Data, Mercury), and Alexis Weill (Head of Data Science, Perplexity). On the agenda

dbt, Clay, Sigma, Mercury, and Perplexity take on what a data stack looks like when agents write, respond, and assemble their own context at a pace far faster than we humans do. If your agents are starting to do more than read, you'll hear how 5 companies are tackling that challenge and the benefits they are seeing from it.

My take: Data and AI are linked together now, just like software and AI. Agents writing data makes a lot of sense as the next step to me.

Finding a Seat at the Table for Your Agent

Michael Rhee
Talk: Michael Rhee, Senior Manager, Data Engineering, Fleetio. On the agenda

Michael's team built a capable data engineering agent, but found that creating it was only half the challenge. They wanted to make the agent a "real" team member.

Fitting an agent into their workflows raised harder questions around trust and accountability. You'll leave with the framework his team uses for agent oversight.

My take: If you have questions about the details, so do I! This approach is definitely at the frontier and I want to know more.

Getting It to Work Is the Easy Part

This is the most AI-flavored theme of the day. Even here, the hard problems show up after the demo.

AI at Scale Inside Real Companies

Tomasz Tunguz
Sudeep Das
Tosh Rayadhurgam

Panel: Tomasz Tunguz (General Partner, Theory Ventures, moderator), Sudeep Das (Head of New Verticals AI, DoorDash), and Tosh Rayadhurgam (Head of Advanced AI, Stripe). On the agenda

Getting an AI feature to work is the easy part. Keeping it working for millions of users, under real latency and cost constraints, is decidedly harder.

The heads of AI at Stripe and DoorDash will share how they evaluate models, decide what to build versus buy, and structure teams around systems that are never quite finished. If you have a working prototype and a roadmap that says "production," this is your panel.

My take: Scaling AI systems is absolutely an in-demand skillset. Demos are one thing, but it's a lot tougher to actually build a full platform.

What Does "Active Customer" Mean, Anyway?

If you want AI to give people the right answers, it first needs to understand what the business means. Both talks here are about getting that meaning in front of the AI.

Automating the creation of the semantic layer

Sam Redfern
Talk: Sam Redfern, Staff Data Scientist, Canva. On the agenda

Do you have 3 definitions of "active customer" and just one expert who knows the difference? You are not alone.

Sam at Canva tackled this problem with semantic leaking, a novel technique that uses LLMs to deliberately leak context between technical artifacts and the language of the business. It surfaces the terms people actually use, generates realistic questions for each domain, and finds the concepts and relationships your models need to support.

Then Sam turns that same context against the system. Ambiguous, adversarial, and deliberately unanswerable questions test whether the AI answers what it can, refuses what it can't, and explains the difference.

My take: Agents should answer most questions, but sometimes the agents really ought to escalate to the data folks! Teaching them how to decide is powerful.

Delivering Self-Serve Intelligence in a Composable System of Record

Arnav Mishra
Talk: Arnav Mishra, Co-Founder & CTO, Doss. On the agenda

Picture an operator asking why an order is late. The answer is in the ERP somewhere, but getting to it means knowing the data model, stitching together reports, and asking a technical team for help.

Doss is building a system of record whose schema reflects how each customer actually runs their business, with interfaces and agents on top. Arnav will show how they connect operational data, business context, permissions, and agent-driven exploration, where that flexibility helps, and how to make a system of record actually useful.

My take: I'm a self-service-maximalist, so a talk about serving customers who know their business better than their data is right up my alley!

The Harness Around the Model

Which matters more for AI-written SQL: the model, or everything around it? Julia's talk has a clear answer.

The SQL Agent Harness We Didn’t Mean to Build

Julia Silge
Talk: Julia Silge, Engineering Manager, Posit PBC. On the agenda

Natural language to SQL rests on an appealing assumption: the model is what makes the query correct. Julia argues the result depends far less on the model than on the harness around it, meaning what the model can see, verify, and run.

Her team built SQL tools in Positron (a data science IDE from the team behind RStudio) for people: schema browsing, completions, an interactive console, and a data explorer. Then they pointed an AI assistant at that same environment and realized they had already built the harness.

You'll learn what a good SQL authoring harness requires, and why bolting on an "AI mode" is so often the wrong shape for this kind of work.

My take: For SQL, a more advanced harness beats a more expensive LLM any day. Learning what a good harness needs is a very valuable skill!

Understand (and Own) Your Data First

AI can turn documents into tables faster than ever. These 2 talks cover what has to be true about your data for that speed to pay off.

Understanding comes before modeling: a data engineer's philosophy for the age of AI extraction

Lucie Thimus
Talk: Lucie Thimus, Data Engineer, Opto Investments. On the agenda

Thanks to agents, regulatory documents, insurance claims, and bills can become tables in a heartbeat. By itself, that won't amplify your team.

Lucie built the data platform at Opto Investments on 1 principle: you don't model what you don't understand yet. You keep the data, study it, and let it show you what it is before you impose structure. That is how you reach scale.

When AI extraction lands on a well-understood data model, every new document and source compounds. Skip that work, and every new thing breaks something.

You'll see the infrastructure behind that approach, including what it looks like when the data keeps surprising you and there's no playbook.

My take: Documents are the final boss of unstructured data. Lucie is getting real value out of them, and it takes both data and AI expertise.

Your Compliance Team Was Right: Why Regulated Firms Must Own Their Data+AI Stack

Upal Saha
Talk: Upal Saha, Co-Founder & CTO, bem. On the agenda

For banks, insurers, healthcare organizations, and law firms, sending sensitive data to third-party AI tools can be more than risky - it can be illegal. But sitting out AI isn't an option.

Upal has deployed document intelligence pipelines inside regulated Fortune 500 environments. He'll cover the rules driving this (privilege, HIPAA, GLBA, and data residency) and the architecture patterns that make it feasible for lean data teams.

He'll also show where to draw the line between building and buying self-hosted components, so compliance, legal, and engineering all sign off. You'll leave with a concrete framework for bringing AI to your data instead of shipping your data to AI.

My take: For something as critical as security but as new as AI, you want to learn from someone who has been there.

See You at the Outpost!

AI gets better when the data work underneath it is done well, whether that's a context layer, a SQL harness, or a data model that fits the domain. And that data work gets more valuable once agents can use it.

Stay tuned, because a few more sessions are still landing! Jordan Tigani (Co-Founder & CEO, MotherDuck) kicks off the day with the opening keynote. Joe Reis (co-author of O'Reilly's Fundamentals of Data Engineering), Ameya Bhatawdekar (VP, Field CTO, Braintrust), and Benn Stancil (former Co-founder and Chief Analytics Officer at Mode) are on the agenda too, with topics coming soon. Plus, Ben Holtzman (Head of Data, AheadComputing) will share real life stories of building an AI-centered self-service data platform. Keep an eye on the agenda for the details.

Want to get your hands dirty first? The day before, November 4, has 11 hands-on workshops where you can build a data agent, scale DuckLake and Iceberg for agent workloads, stream real-time data with change data capture, and even build an MCP server that reserves a real mystery gift!

If these abstracts read like your team's whiteboard, come ask your questions in person. Register for Data Outpost, and come a day early for the workshops if you want to get hands-on.

See you in San Francisco!

Subscribe to motherduck blog

PREVIOUS POSTS