The MAD Podcast with Matt Turck · AI Research & Frontier Labs · October 2026
Pavlo is describing a side project in which he tracks every database system he knows about. He notes the count is a floor, since people can turn off the co-signing that marks a commit as agent-written, though most don't. It comes up as he argues that agents can now build most of a database system with enough guidance.
Yes. So to give one anecdote, I teach a course on database minimum systems at Carnegie Mellon University. A year ago, the agents couldn't do our entire project. So the projects would be like, we give you a scaffolding web database system and you have to implement the indexes, the query engine, and things like that. It could do some of it, not all of it. I think it was Opus 4, whatever Anthropic put out last year, then that just opened the floodgate and the agent basically do all our assignments without very little prompting. And of course, there is a lot of training data for it because all our projects are open source. They're all on GitHub, not just students at Carnegie Mellon University, but also students outside of the university. We let them use it. So there's a lot of training data for them to implement things. And so I think agents basically could implement anything, you know, you would want to build a database system now. You know, obviously you have to prompt it the right way and hold its hand and make sure you generate the right design or produce the implementation based on the design that you want. At the end of the day, I think these agents are very capable to be able to do this.
So they could build an entire database because so my experience of building a database as a venture investor is that it's a 10-year journey of pain where nothing much happened for three years and you have some of the smartest people in the world getting together to solve what seems each time like insurmountable problems. So we're now saying that you can build the whole thing. Yeah,
the old adage from database systems is that it takes 10 years of a data system. You can build the first 90% in three years and then the remaining 10% takes the next 70 years or seven years. So yeah, no, I think that the agents are very capable, you know, with enough tokens, of course, and then with enough guidance, people can, you can build Vibecode entire data building system. And there's certainly companies that are doing this now. And pretty much every single database company is using agents to help develop things. So, one of my, again, I love databases. One of my side projects is the database of databases, dbdv.io. And one of the things we do now is we keep track of every single data system that I know about. And for the open source ones, every single night we pull down all the latest commits on GitHub and then we track to see which ones are actually being co-signed by Claude or Codex and things like that. And obviously, people turn that feature off. You don't know whether it's actually been generated from an agent, but most people don't do that. And at this point, I think like 60%, over 60% of the open source database systems are being have commits coming from agents.
And so, does that be on the writing also apply to the running of it? So, going to that cell-driving database concept, that's something that was a big project of yours 10 years ago, I believe. Roughly, we've been, yeah. Yes. So, walk us through that journey. What was not possible then that has become possible today?
Yeah. So, when I started at Carnegie Mellon University, one of the things I did was my first years, I would go visit companies and sort of see what sort of challenges they were facing with databases. And the overarching theme I saw over and over again was like just running these systems, maintaining them, and optimize them was a huge struggle. And this is not a huge revelation for me. Like, people have been trying to do this for decades. I mean, since the creation of the relational model and relational databases in the 1970s, people have been trying to do auto-tuning for indexes, partitioning keys, sharding keys, and tuning knobs and so forth. Microsoft Research did a lot of work in the early 2000s in this auto admin project. They had a bunch of tools allowed to manage and can optimize data systems automatically. And
for context that there's because there's thousands and thousands of possible configuration of a database system. Right. So,