Skip to main content

Data Engineering Is the Cache Invalidation Problem ft. Josh Wills

How toAgentsAI

- 20 min read

In this episode, we dive into how much models actually matter: more like a preference, like whether you use an iPhone or Android, than their capabilities. We talk about bare-metal disks when running Kubernetes, then move to agents and how data engineering becomes like being an emergency pilot in a plane, and how we automate away the boring work with no more database migrations or unit tests.

We go into spiritual territory, where he feels in his gut whether a pipeline will work. Then we finish with where every data engineer or software engineer eventually wants to end up, as the field is evolving so fast: on a farm with beautiful flowers and no screen. This is interview #5. I hope you enjoy it.

This article is structured into four parts: (1) how Kubernetes and disk abstractions are key for his work, (2) how he works with the models and agents, and how much the model actually matters, (3) what tooling he uses and how agents ease his life at work, and (4) whether data engineering is automated away with agents.

Introducing the Guest: #5 Josh Wills

Josh Wills is a software and data engineer with 25 years of experience. At Google, he worked on the ad auction system and led the analytics infrastructure behind Google+. As head of data science at Cloudera, he founded the Apache Crunch project and co-authored O'Reilly's Advanced Analytics with Spark, and he went on to build and lead Slack's data engineering team. He's also the original creator of the dbt-duckdb adapter.

Since 2024, Josh has been a Member of Technical Staff at DatologyAI, which helps teams train better models, faster and smaller. There he builds tools that automatically select the best data for training deep learning models, along with the distributed pipelines that process it at trillion-token pretraining scale. We go into more detail on what it takes to work with datasets that gigantic.

TIP: Fun Fact

He basically retired in 2022, and "the AI stuff got him back" 😉

DuckDB Yes, but also Kubernetes Disk Abstractions

As many of you might know, Josh was the initial creator of dbt-duckdb, so I asked him if he still uses dbt and, more importantly, how much DuckDB.

His answer was that everyone uses it: "Everybody uses DuckDB in some capacity. I feel like anyone who knows about DuckDB uses DuckDB for something. Everyone has their own specific use case for it". He is using it less for his specific pipelines at work, as they have huge datasets, text data, not classical datasets: e.g., 30 TB is tiny. He uses Spark or Ray. He does much of the processing on text and multimodal data (time series, images, etc.), not much classical relational data in his day-to-day.

The use case for DuckDB he uses a lot is reading Parquet metadata that sits on S3, as the above data pipelines "throw off enormous amounts of Parquet metadata, which is still substantial but much smaller, on the order of 100 GBs". Information about runs, when a dataset was last used to train a model, how much he pays to store it in S3 for indefinite periods. He usually asks his agents to read that on his dev box on EC2, which he runs these commands on, to cut down on transfer times, just for speed.

NOTE: DuckDB got integrated into dbt v2

Interestingly, dbt-duckdb just got an update. The dbt adapter for DuckDB received its first pull request on August 27, 2021, and in the meantime has 1.4k stars on GitHub. Since then, dbt users have been able to install one Python package (dbt-duckdb, via pip), point it at a file (a local DuckDB database), and have a working project (models building into tables and views), without having to sign up for (and pay for) servers or warehouses. With the latest release, DuckDB now runs inside dbt with the dbt Fusion engine. See DuckDB Now Ships inside dbt v2.

"We Lost a Whole Generation" of Data Infra People, Retired 2022, AI Pulled Him Back

When talking about what has changed since he worked at Google as a staff engineer (2007–2011) and led data engineering at Slack (2015–2019), and how data engineering workloads differ today with AI, he said Snowflake is great. Still, we do not have many people anymore who know how the data is laid out on disk.

But not for his size of workload. Josh works with such big data that we normies can barely fathom, and that's why he uses lots of Kubernetes jobs and workloads. He likes Apache Celeborn (Pronunciation: spelled Celeborn, pronounced "Keleborn"), which abstracts the disks (not needing to know where they physically come from), because in Kubernetes you still do that manually, as Kubernetes is not really cloud-abstracted, he says:

Kubernetes is fantastic in that it's in theory abstracted away from an individual cloud, but in practice it's not, because where does your disk come from? If you're on AWS or GCP you can use an EBS-like service, but if you're running bare metal, that doesn't exist the same way.

Celeborn has something called push-based shuffle write, where mapper nodes skip local disk storage, ideal for cloud-native architectures, which lets you separate storage and compute, so

The Spark jobs don't need to worry about where their disk comes from anymore, from a shuffle perspective. There's the shuffle service, no matter how much data you're going to shuffle. You have your local disk, and then you have your shuffle service for all your other disk needs.

NOTE: Nobody Remembers How to Run a Spark Cluster

Josh says the hardest thing is:

That there are relatively few people in places that need to think about this stuff anymore, at all. The whole data community during the late teens kind of went Snowflake happy. And I used Snowflake at my last company, it's a fantastic product. But we have sort of lost a whole generation. There's not many people left anymore who still remember how to run a large, effectively open source, Spark cluster themselves.

How Much Does the Model Matter to You?

When asked about the models and how much they matter, Josh said for him it's like the iMessage blue bubbles vs. green bubbles. Android phones are fantastic, but he wants the blue bubbles. And Josh's prediction is that all the models will become competent, so the real question becomes:

which one you like and enjoy hanging out with

Josh Wills tweet: wtf did anthropic do to opus 5, this is like gpt 3.5-level stupid Source: Josh's tweet

Above, Josh explains how he disliked the Opus 5 model, and continues to do so. In this interview, he added that it's also the way it communicates. An interesting insight, he said, as someone who works on building models all day:

the fact that Opus 5 is so bad, it deeply indicates to me that we literally have no idea what we're doing at all when it comes to training these models.

The good thing: changing might be easier than ever1, and he disliked the previous 5.0 so much that he changed his workflow and switched to OpenAI. He commented that:

with no drop in productivity, it's kind of insane.

TIP

One trick I learned: use the caveman skill, and load it at the beginning. I use only the one MD file of it, and then every model sounds the same, and even uses fewer tokens and is less verbose. It also needs less brainpower to read through.

How to Code with Agents These Days

Next, we talked about working with agents. He jokingly came back to a Tweet where he wants to start a reinforcement learning (RL) environments company, where he does the teaching:

Josh Wills tweet about an RL environments company teaching models to write data pipelines

Josh says:

[..] a joke that I want to start my own RL environments company where it's just me teaching the models how not to write absolutely garbage data pipelines. Because I use agents to write all of my code, with very, very few exceptions. And yet I still find myself deeply unhappy with the data pipelines it writes.

TIP: His favorite utility is still htop

He said back in 2022 that he loves htop. Asked about his favorite tools, he says that htop is fantastic, and he still uses htop for a lot of stuff he does. In case you don't know, you see all processes and CPUs of your machine at a glance. See htop - an interactive process viewer.

How to Get Fast Feedback Loops with Agents?

Josh says AI wins where there's a large training model and fast responses, because if it's not in the training data, it's slow and it needs to do a lot of stuff. He says: "Any sort of environment in which the models have highly reliable training data with fast feedback mechanisms, they're just going to dominate. Absolutely going to dominate. You see this with math, you see this with front-end engineering, any RL environment where there is a fast feedback mechanism."

And that's exactly what's missing in large-scale data pipelines today:

What's tricky about data pipelines, at least at a large scale, is that you don't actually have that fast feedback mechanism available to you, because running a giant pipeline can take a long time and a lot of computers.

Fast feedback loops are also what Josh was already referring to almost four years ago, before any of the AI hype. He liked fast feedback loops, and quoted Erik Bernhardsson:

Erik Bernhardsson quote: make the feedback loops fast
NOTE

Interestingly, today Jev came out to a lot of praise online, and it's based on Reinforcement Learning For Calibrated Decisions (RLCD), which is basically what Josh wanted.

Feeling the Tuples Flowing through the Pipelines, a Gut Feeling

Josh, with his years of experience, has built a gut feeling:

Through years and years of suffering and writing gigantic pipelines, I can look at a pipeline and it kind of feels like something to me. I can feel the tuples flowing through it, because I have a mental model of how this is going to go.

He explains how SQL databases have a great fast feedback mechanism with their query explain plans. They tell you if it's a good or a bad query. But for a Spark job, it's not yet exactly that.2

Josh describes an aesthetic aspect of data pipelines:

I have, weirdly, a kind of aesthetic sensibility when it comes to data pipelines. A data pipeline can be beautiful to me. It can be elegant, well designed, thoughtful. That makes me want to look, in the way other people would visit an art museum, or read math proofs because they find them beautiful.

That's where he still sees a lot of "ugliness" in what the agents write. They were not yet beautiful when we talked a couple of weeks back. Josh said, "Some of the best programmers I've worked with were musicians by training". And he continues:

One distinguished engineer at Google, his master's degree is in the violin, and he was the best coder I think I'd ever seen, because his code was just like music. It was beautiful.

I have a similar affinity for writing text, and good writing has a musical quality to it that certain people can hear. You develop an ear for good writing after a while. It sounds like something in your head, like poetry or prose. Josh says that "This is something we have not figured out how to RL in any kind of way".

The Cost Aspect, Literally Lighting Money on Fire

When talking about cost, Josh says that:

"the models might waste time on the most pedantic, stupid checks for errors that are simply not possible given my knowledge of the structure of the upstream data sources." "Literally lighting money on fire"

He continues that the pipelines are huge and compute costs a lot. And models might write a pipeline for average-sized data, but not have the knowledge or the big training data on "gigantic pipelines": "It will work, it's correct. It's just stupidly cost inefficient".

He has seen this problem at companies getting rid of their data teams, as LLMs could vibe code a job in Databricks. He said it like this: "You can easily vibe code a job in Databricks that costs like two hundred thousand dollars a day to run, that a competent data engineer, or at least one who supervised it, would never have written."

And it's not the model's fault, but a function of the training data that has not investigated or run a lot of "gigantic data pipelines for multiple hours across hundreds of machines".

That's just not the way the reinforcement learning environments are structured right now:

It's this fast feedback loop: tiny data, is it correct? Tiny front end, does it render the CSS correctly?

Josh makes a great science analogy:

People talk a lot about how science has the same problem. Labs, biology, physics, these are all real things where you literally have to wait for the culture to grow. Data engineering has a lightweight version of this. We are not as immediately amenable to this kind of feedback style right now.

He found the models to be fantastic at SQL back when we held the interview. "They're really, really good," he said. But for his kind of work (the raw Ray pipeline, the Spark pipeline), there is no explain plan. You're still building the thing from scratch, as he said.

The Job of 'Emergency Exception Handler'

He mentions that in his job, because the models haven't learned everything yet, he is basically like a pilot who usually flies on autopilot, but in case of emergency, he is there to fix it or take over.

What Is the Tooling He Uses & How Agents Ease His Life at Work

As initially pointed out, he disliked Opus and switched to Codex, so now he is using Codex all the time, but he used Claude for a long time.

When asked how he works or reviews code, he said: mostly giving feedback, like a staff engineer does. His life as a senior technical IC has been trending towards that for a very long time, he says. He primarily gives feedback. That's his primary job. Interviewing people, reviewing design documents, and less hands-on, with very few exceptions.

He likes to work on high-demand, high-urgency work that forces him to create. This is back to the 'emergency exception handler':

If there's a broken-ass data pipeline and it has to run right now, because a customer has a training job coming up and the data has to be ready for it. That's the primary value I provide.

What Josh Likes about Agents: No Feelings and the Caffeine Dial

Josh also revealed that he likes working with agents because agents have no feelings. He can give agents feedback on demand, force things to be created, he says:

I don't have to be very nice to the agents.

He can just say things:

I don't have to care about their feelings or their career development, because they're going to cease to exist when I'm done with this coding session. I can be like: this is bad and you should feel bad. Ctrl-C, you're dead, I'm going to start over with someone else.

As someone who calls himself "a fairly mediocre manager, because I don't generally like caring about these things", this works well for him 🙂3.

The other interesting thing he says:

I can dial up and down the work on demand as I see fit. If I'm in a moment where I need five agents going, I can do it. If it gets to be too much, I dial it back. In the same way that I try to maintain a certain level of caffeine at any given moment of the day, I can maintain a certain pace to match my energy, to match what I'm up to right now. Am I super excited, am I kind of bored? That's a virtue of this new world that was not true before.

It's like the caffeine dose.

No More Unit Tests!

He also likes to automate most of the boring work. Who wants to do one more database migration? Or fix "flaky unit tests"? He likes to do the higher-level strategic stuff, which is how people end up becoming investors or product managers, he says.

Even more so:

I have an agent whose whole job is just to watch for flaky tests and fix them. That's it, that's all it does. Database migrations? Absolutely not. I have an agent that does database migrations. I have automated away the shittiest, lowest-value, most unpleasant parts of my life, and I am here for it.

For him, that's a massive quality increase. And he can focus very hard on things he likes, and where the leverage is high, to make hard decisions.

Decisions Made Are What Is Left, and Communicating Why Is Key

In regard to code review, one thing he is trying to be militant about in the world where agents do everything is:

For every project, I'm focusing really hard on the decisions I am making.

He mentions a concrete example from Datology where he's building an index over datasets to figure out which parts of a dataset are reusable, wholesale or partially, someplace else:

  • Option A: query the metadata catalog at read time and figure out the useful assets.
  • Option B: write an explicit index up front: "this asset exists in this index and it is usable for this purpose"

He says:

Either choice is valid, there are pros and cons. I'm making an explicit trade-off here, and those are the decisions I need to communicate. The details of how the index works are much less important.

He makes a great point that:

It's really the decisions we need to communicate. The decisions build the shared understanding, the choices we make, the trade-offs. That's still the meat of engineering work.

I also think this very thing is the most important, and it can turn into a bad spiral if you are not careful, as someone vented in this tweet. Plus, in my opinion, agents force you to focus on architecture, and that's why they're handier for seniors.

Is Data Engineering Automated Away with Agents?

Josh says clearly that data engineering is not going away. For the same reason it hasn't gone away for so long: the fundamental problems won't go away. We're still solving the problem of materializing views and caching!

I don't see data engineering going away in any way, shape or form, even in this world. Aside from the fact that the RL environment problem makes it unusually difficult for the agents to be good at running these very long pipelines, the fundamental problem of data engineering is the same problem it's been my whole career, and 25 years before I started working. It's the materialized view problem.

The questions are: "What data should I pre-process and create now to make queries faster later on, and what stuff is not worth doing, what should I just do in the moment?" These are the fundamental problems.

He says it will be a "problem for the superintelligence", because:

It's not a fixed thing, it's going to continuously move, and the right trade-off is going to change all the time.

Josh says, aside from reinforcement learning:

It's the two hard problems in computer science, right? Cache invalidation and naming things. And the agents still suck at them because they're legit hard.

"Data engineering is fundamentally the cache invalidation problem", he continues, "just done in a batch-processing context".

NOTE: Even Hannes Started to Use Claude Code

Hannes Mühleisen said in The Past, Present, and Future of DuckDB, DuckLabs that he now uses Claude Code for DuckDB, since the Quack Protocol launch in May 2026 or so. So he hasn't written a lot of code since then, but he still looks at things.

Humility and Burnout

Lastly, we talked about humility and how not to get burned out.

It helps to stay humble, as he says he is very good at it, telling his family:

I am simultaneously an idiot and also among the best in the world at what I do.

I say that we all probably feel some kind of ambivalence, as the AI can do part of our job and we feel unneeded. That's where humility helps us. Josh also adds that he is kind of grateful that he can still be at the keyboard and program. It's "tremendous fun" for him, and he loves to be here for this moment in time.

I asked if he does not get burned out when agents do not generate code the way he does it. Or does he split into "code I don't mind" vs "code I want perfect"? He says that his advantage is having done it for a "really, really long time":

Old me tolerates things that younger me would never, ever, ever tolerate at all. I am used to cycles of feedback with somewhat "stubborn" people.

He said that once, the agent couldn't find a problem he was telling it about: "And I could not, for the life of me, get the agent to write the pipeline in the way it would do this". He was telling the agent "you cannot use caching this way", and it couldn't help itself. He had to take over.

This again is the perfect example of the exception handling situation. It's "kind of my job these days," Josh says, to "identify these situations where, for whatever reason, the model literally isn't capable of doing what I am asking". But he says that the models are getting better every day:

The reality is, I don't encounter that many of these problems anymore.

Conclusion

One main takeaway from this conversation is that data engineering is the cache invalidation problem, just done in a batch-processing context. That's important: when done right, agents dominate because of a fast feedback loop, like SQL with its explain plans or math. Gigantic Spark and Ray pipelines that run for hours across hundreds of machines don't have that cycle yet. That's why Josh still ends up as the emergency exception handler, and why others can vibe code a Databricks job that costs $200k a day.

What's interesting is how little the model itself matters to him. It's more about design and feel, e.g., blue bubbles vs. green bubbles: pick the one you enjoy hanging out with, switch when you don't like it anymore, or hand off the flaky tests and database migrations to an agent. What's left is what was always the hard part of data engineering: making the decisions and communicating the trade-offs.

We got a bit spiritual, with Josh feeling the tuples flow through a pipeline, and ended where many of us do as the field moves this fast: dreaming of a farm with beautiful flowers and no screen 🙂. Josh admits he periodically flirts with it too: "Maybe I'm just going to retire again and just go play chess. I'm not made of stone".

I hope you enjoyed this episode of the interview series. Check out the others, starting with Mark Freeman, Chris Riccomini, Wes McKinney and Maxime Beauchemin.


If you read this far and liked it, also check out the talk with Mehdi about the "Coupling Problem in Data Engineering" or check Josh's recent talk Josh Wills on the End of Coding at Pebblebed.

1.He also has no shame about hopping back when the model changes. With the latest model, he went back to Opus 5.5.
2.The same impression someone had on Reddit: "agents are absolutely dogshit AI at optimizing Spark jobs (or any sort of high complexity pipeline). They are laughably bad, in my experience."
3.I haven't worked under him, but I'm sure that he is a great manager nonetheless!

Subscribe to motherduck blog