AP

Andy Pavlo

Things Andy Says on Podcasts

Where to Find Them

Andy Pavlo has been a guest on The Data Stack Show (2 times) , The GeekNarrator , The Analytics Engineering Podcast , The MAD Podcast with Matt Turck and Data Engineering Podcast .

Recently: “What Happens When Billions of AI Agents Hit Your Database? (Andy Pavlo)” on The MAD Podcast with Matt Turck (October 2026); “AI Powered Database optimisation with Andy Pavlo, Ottertune” on The GeekNarrator (January 2024); “The State of Databases Today (w/ Andy Pavlo)” on The Analytics Engineering Podcast (September 2023); “135: Database Knob Tuning and AI with Andy Pavlo and Dana Van Aken of OtterTune” on The Data Stack Show (April 2023); “The PRQL: Database Tuning and Optimization with Andy Pavlo and Dana Van Aken of OtterTune” on The Data Stack Show (April 2023); “Make Database Performance Optimization A Playful Experience With OtterTune” on Data Engineering Podcast (June 2021).

What They Said

“Relative to AI agents everything looks stagnant, because in the history of computer science there's been nothing like that before. It's like taking a cheetah, giving it a bunch of cocaine and putting it in a Ferrari. The amount of speed that people are developing these things is insane. So everything looks slow or dead or stagnant to that.” — Andy Pavlo, The MAD Podcast with Matt Turck

The host suggests the database market has slowed down after the burst of activity in the 2010s. Pavlo answers that the field has been through consolidation cycles before, and that it only looks stagnant by comparison with agents. He goes on to say he doesn't expect agents to force anyone to throw out the relational model.

The MAD Podcast with Matt Turck · 2026-10-08 Permalink → Listen →
The MAD Podcast with Matt Turck Around 1:06:16 into the episode
Andy Pavlo

I mean, you can always build new data systems for new hardware. But my track record on this is terrible in terms of like, we've done a bunch of research on experimental hardware and it always gets canceled. Or like, it's even not that experimental. It's like, you know, Intel had this optane persistent memory stuff. We did my, we did a bunch of research on building systems for that. Because like, if you assume now your DRAM is persistent, you know, like you pull the plug and you don't lose anything, that changes how you fundamentally build a data system. We did a bunch of work on that. Then Intel killed that product line. We were doing another research on processing and memory hardware. So think of like DRAM sticks with like CPU cores directly on the DIM. So the data system now can say, okay, instead of pulling things from memory, printing into my CPU caches, and then I can compute things on them, I'll just send the query down to the DIMM itself and run it there. We were doing a bunch of work on this on this thing called Opmem that got bought by Qualcomm and got killed last year. So that didn't work out. So there's always a bunch of work you can do on data systems or new hardware. I would say that actually, I don't know the answer, right? This is one of the things I'm trying to figure out at Click House is like, you know, we talked about this in the very beginning. Like, are agenic workloads significantly different than what humans or what existing applications do now? And if so, why or how? And could you, how would you change maybe the development of a data system to take better advantage of this? That remains to be seen how that works. I think there's always a bunch of problems in query optimization. I think they're interesting. That remains the hardest part about database systems. Incremental memorialized views, another big challenge. Again, but these are not like things that no one else has thought of. People are trying to do these things for decades. So I think the agenic stuff is probably the most interesting and relevant thing to me right now: what changes with these workloads? What changes in the system architecture? And to be honest, I don't know the answer. I just haven't seen it yet.

Matt Turck

I'm asking because it generally feels like the database market is a specific moment in its history, meaning that there was an explosion of activity in the 2010s. There's SQL versus NoSQL, that all evolution. Then there was the emergence of Databricks, Snowflake, and now Click House. But it seems that things have slowed down a little bit in terms of explosion. I guess, you know,

Andy Pavlo

stagnant. Yeah, no, but we've been to this trend before, right? There was a lot of activity relational databases, 1970s, 1980s. And then the 1990s, again, people sort of, you know, the market sort of solidified around these major enterprises, the Oracles, the, you know, the IBMs, Teradatives. And then, you know, if you, if your only viewpoint of databases were from those kind of companies, then yeah, it looked like it's been stagnant for years. But as you said, a lot of activity into 2000, 2010s. I mean, relative to AI Agents everything looks stagnant because there's in the history of computer science, there's been nothing like that before. Like, it's just like, you know, taking a cheetah, giving a bunch of cocaine and putting in a Ferrari. Like the amount of speed that people are developing these things is insane. So everything looks slow or dead or stagnant to that. But at the end of the day, I think the volume is going to matter a lot. I think that one interesting question is, to your point, Of like cost and efficiency, like squeaking out the best you can, the best performance you can get for the hardware that you have, trying to reduce that cost, that's always an interesting challenge that could pursue. But like the end of the day, I don't think there's going to be a massive change in what data looks like that requires us to throw everything away that we've known about databases. In the same way, you wouldn't come up with a new notion of arithmetic or mathematics to replace one plus one equals two. The relational model itself is the foundation of how you want to represent data. And it's just you can vary the implementations of that. And so there's certainly a lot of work to make that these systems more efficient for this. But I don't think you're going to throw everything away. And like, you know, the agents need some kind of database system that you've never even conceived of now. At the end of the day, it doesn't make sense. So I don't think that part changes. I think, like I said, always new hardware. I think there's certainly improvements that could be done for SQL. There's always going to be people trying to replace SQL. I think that might be a lost cause. Although SQL could end up being like how in the same way that people don't write assembly anymore, SQL might end up being like that because text to SQL works so well. I think the agent stuff is super interesting. And I think getting better performance is always going to be a from my perspective about fun challenges and things we can pursue. You can imagine a crazy world where you say, I don't need a general purpose data system anymore for every single application. I want to vibe code exactly a data system that does this for this one thing can then be hyper specialized. You kind of do this now with code generation or just time compilation for some aspects of queries. And some systems do that. Click House, Postgres, Umbra from the Germans. But like hyper specialization and making that be sustainable might be a bigger research question going forward. So

Matt Turck

it's this, I don't know if it's a paradoxical kind of situation, but on the one hand, it's a bit of a stagnant industry right now in terms of evolution. At the same time, as we've hopefully established through the conversation, the layer itself is as important as ever, which is one of the reasons why your friend Larry Ellison is the always close to the richest man in the world or was it. He's down

Andy Pavlo

as of today. He's back in eighth. I mean,

Matt Turck

Oracle is a lot more than just a database company, but it's still the core program. I mean, the

Speaker names from our own diarization · position estimated from where the line sits in the episode
“For the open source ones, every single night we pull down all the latest commits on GitHub and then we track to see which ones are actually being co-signed by Claude or Codex ... At this point, I think over 60% of the open source database systems have commits coming from agents.” — Andy Pavlo, The MAD Podcast with Matt Turck

Pavlo is describing a side project in which he tracks every database system he knows about. He notes the count is a floor, since people can turn off the co-signing that marks a commit as agent-written, though most don't. It comes up as he argues that agents can now build most of a database system with enough guidance.

The MAD Podcast with Matt Turck · 2026-10-08 Permalink → Listen →
The MAD Podcast with Matt Turck Around 34:47 into the episode
Andy Pavlo

Yes. So to give one anecdote, I teach a course on database minimum systems at Carnegie Mellon University. A year ago, the agents couldn't do our entire project. So the projects would be like, we give you a scaffolding web database system and you have to implement the indexes, the query engine, and things like that. It could do some of it, not all of it. I think it was Opus 4, whatever Anthropic put out last year, then that just opened the floodgate and the agent basically do all our assignments without very little prompting. And of course, there is a lot of training data for it because all our projects are open source. They're all on GitHub, not just students at Carnegie Mellon University, but also students outside of the university. We let them use it. So there's a lot of training data for them to implement things. And so I think agents basically could implement anything, you know, you would want to build a database system now. You know, obviously you have to prompt it the right way and hold its hand and make sure you generate the right design or produce the implementation based on the design that you want. At the end of the day, I think these agents are very capable to be able to do this.

Matt Turck

So they could build an entire database because so my experience of building a database as a venture investor is that it's a 10-year journey of pain where nothing much happened for three years and you have some of the smartest people in the world getting together to solve what seems each time like insurmountable problems. So we're now saying that you can build the whole thing. Yeah,

Andy Pavlo

the old adage from database systems is that it takes 10 years of a data system. You can build the first 90% in three years and then the remaining 10% takes the next 70 years or seven years. So yeah, no, I think that the agents are very capable, you know, with enough tokens, of course, and then with enough guidance, people can, you can build Vibecode entire data building system. And there's certainly companies that are doing this now. And pretty much every single database company is using agents to help develop things. So, one of my, again, I love databases. One of my side projects is the database of databases, dbdv.io. And one of the things we do now is we keep track of every single data system that I know about. And for the open source ones, every single night we pull down all the latest commits on GitHub and then we track to see which ones are actually being co-signed by Claude or Codex and things like that. And obviously, people turn that feature off. You don't know whether it's actually been generated from an agent, but most people don't do that. And at this point, I think like 60%, over 60% of the open source database systems are being have commits coming from agents.

Matt Turck

And so, does that be on the writing also apply to the running of it? So, going to that cell-driving database concept, that's something that was a big project of yours 10 years ago, I believe. Roughly, we've been, yeah. Yes. So, walk us through that journey. What was not possible then that has become possible today?

Andy Pavlo

Yeah. So, when I started at Carnegie Mellon University, one of the things I did was my first years, I would go visit companies and sort of see what sort of challenges they were facing with databases. And the overarching theme I saw over and over again was like just running these systems, maintaining them, and optimize them was a huge struggle. And this is not a huge revelation for me. Like, people have been trying to do this for decades. I mean, since the creation of the relational model and relational databases in the 1970s, people have been trying to do auto-tuning for indexes, partitioning keys, sharding keys, and tuning knobs and so forth. Microsoft Research did a lot of work in the early 2000s in this auto admin project. They had a bunch of tools allowed to manage and can optimize data systems automatically. And

Matt Turck

for context that there's because there's thousands and thousands of possible configuration of a database system. Right. So,

Speaker names from our own diarization · position estimated from where the line sits in the episode
“Graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.” — Andy Pavlo, The MAD Podcast with Matt Turck

Pavlo says he has a standing bet that graph databases will not overtake the relational market by 2030, and that he will wear an "I love graph databases" shirt on his ID photos if he loses. His argument is that relational engines with the right optimizations already do graph traversals efficiently, and that the SQL standard now supports property graph queries.

The MAD Podcast with Matt Turck · 2026-10-08 Permalink → Listen →
The MAD Podcast with Matt Turck Around 1:02:20 into the episode
Andy Pavlo

of the things that happened before when I was a professor is like, I would always get, you always hear rumors about who's doing well, not doing well through a combination of like either the investors or like former employees or students that maybe go to internships or whatever, like or like maybe interview some places and they come back. So you get sort of bits of information from everyone to kind of piece together what the data landscape looks like. Probably not being in Clickhouse and now I see everything. Like as an investor, you see everything too. The one vector database company that I know is doing very well is TurboPuffer. And they are hyper-specialized in doing vector search at a cost performance ratio that's much better than everyone else. So I don't think that the vector databases are going to go away. I think that they'll evolve in two ways. They'll have to become either a sort of a general purpose system like a Postgres, like a MySQL, where they become the system of record where you're storing the original tuples plus the embeddings of the vectors for them. Or they become like an Elasticsearch where there's a separate system where they have a copy of the data that's being pulled from the operational side. And in that case, they can live sort of comfortably as being this additional thing you add on. And if you want the raw best performance of vector search, You some cases you may have to go to one of these specialized systems. So I don't think that's that's going to go away. I just don't think I've seen predictions of like, oh, Postgres is going to die at the hands of a vector database. That's not happening. That's not happening.

Matt Turck

Very much the opposite, right? Yes. Graph databases, we mentioned at the beginning. So, you know, not to become them, but like Neo4j has been around for 20 years now. Sure, yes. And this was supposed to be the moment, right, for graph databases. So what's happening there?

Andy Pavlo

I say I have an outstanding bet with somebody on Hacker News where they said that by the year 2030, the graph database market was going to be was going to overcome be larger than the relational database market. And if this becomes true, then I will wear a shirt that says I love graph databases and I will use that as my driver's license, my university ID. I'll put on my website to the day I die, right? I'm pretty comfortable. It's 2026. We got four years ago. This is not happening. No, it's always been a niche market. And I think that because my perspective on the research side, the research shows that if you do certain things in implementing the engine, which Clickhouse does do, DuckDB does some of this as well. There's things you can do that allow you to do the traversals of graphs, which essentially just joins, soft joins on the table. You can implement those things very, very efficiently and you can easily outperform Neo4j. And that's kind of like kicking. That's like Neo4j is like saying you're faster than Neo4j is like saying I'm faster than somebody maybe like that's you know that's in a wheelchair, right? You can run fast. That's it's it's a low blow. So but I'm just saying all the graph databases, I think like even the best ones, you're just not going to relate to a system that has a bunch of these optimizations that are in the research and actually appearing in some of these systems now. You're just, you're going to lose. What you will lose on against the graph database if you're doing the graph traversal with the client side and the server side, meaning like I, you know, I got to figure out what the next node I want to go look at. I go back to the client and that size the next node to go traverse. If you're doing that back and forth, yeah, they'll beat, they'll beat you guys. But like I said, the SQL standard now supports property graph queries. Oracle has this, right? They were a big pusher of this, this extension of SQL. It allows you to do that traversal on the server side. So like graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.

Matt Turck

GPU databases. Yes. I know you have a special interest there. There was a cycle when a generation appeared, then it went away. There seems to be a renewal. What is a GPU database and what is your prediction? So

Andy Pavlo

GPU database is a data center system where the execution end for queries is offloaded to a GPU running on the PCIe or running in the same box or another box. So, I mean, the history of people trying to build accelerators for data systems goes back to the beginning of data systems. In the 1970s, they were called database machines. So people would build specialized hardware to run sorting and query execution operators. And that obviously died out in the early 1980s because by the time it takes you to like design and fab new specialized hardware, Intel or Motorola will put out the next CPU or the hardware got better and just the gains you were getting went away. So hardware accelerators for databases basically died out in the 1980s. There wasn't a lot of activity in the 90s, 2000s, you saw sort of the rise of people trying to do FPGAs for databases. And every so often that comes back now. Some of the cloud vendors do a little bit of these things, but usually like to filter things on the NIC on the network side of things coming in. So for GPU databases, again, there was a bunch of systems in the 2010s that were trying this. We did a seminar series at the university where we invited all the GPU database guys come to give talks about what they were doing, why they were faster than existing systems. And the big challenge at the time was with those systems, you had to put the entire database inside the memory of the GPU. Because if you had to go back up through PCIe, it was just way too slow. Andy Pavlo bunch of those startups sort of fizzled out. Some of them are still around, but they're sort of specialized for doing visualizations. And then there was in the last year or so, NVIDIA has basically gobbled up a bunch of these GPU database companies that were kind of like struggling along. And I was an advisor for one of them called Voltron. But they also picked up HeavyDB. And so NVIDIA is all going all in on this now. So it remains to be seen whether the idea that you're going to build a CPU-only database system, like long-term, whether that's going to still hold. I've heard getting mixed reports like this is public. Microsoft has, you know, their, they have offerings now in the cloud that can be accelerated, you know, for your data system can be accelerated GPUs for analytics. Another major database company that I can't say who they are, they looked at the economics of GPUs and decided it wasn't worth it. So, one of my former students now is a professor at University of Wisconsin. They're now on leave at NVIDIA. They have a project called SirisDB, which is not necessarily a new data system, but it's a layer in between an existing system and like the CUDA. And so, it supports taking DuckDB queries and running that down on the GPU. I think they can do this in Doris or Star Rocks and DataFusion. And so, at Click House, we've been potentially looking at this as well, but it's research. I don't like it, it's interesting to see whether how much you have to do translation between how Click House expects things and how CUDA wants things to be, the data layout and so forth. How do you organize memory or share memory between these different components? TBA remains to be seen whether this actually makes sense. But certainly, there's a lot of research energy behind this. And publicly, I can say this: like NVIDIA is obviously pushing this because it'll sell more GPUs, right? Because it's hard enough to get new CPUs. Everyone's compute-bound, or memory is hard to get. The computing hardware is very expensive, hard to get now. And GPUs, of all the things, is the most expensive hardware to get. And now you can say your entire database is going to run with GPU. I don't know if that makes sense, at least in the short term. But if the performance improvements are quite significant, and some of the research shows that it is, maybe it makes sense.

Matt Turck

Is there an emerging category or maybe a niche somewhere within a category that people don't talk about enough yet? Of databases? Yeah.

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In