AP Andy Pavlo On The MAD Podcast with Matt Turck

“Graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.”

The MAD Podcast with Matt Turck · AI Research & Frontier Labs · October 2026

“Graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.” — Andy Pavlo, The MAD Podcast with Matt Turck

Pavlo says he has a standing bet that graph databases will not overtake the relational market by 2030, and that he will wear an "I love graph databases" shirt on his ID photos if he loses. His argument is that relational engines with the right optimizations already do graph traversals efficiently, and that the SQL standard now supports property graph queries.

Transcript

The MAD Podcast with Matt Turck Around 1:02:20 into the episode
Andy Pavlo

of the things that happened before when I was a professor is like, I would always get, you always hear rumors about who's doing well, not doing well through a combination of like either the investors or like former employees or students that maybe go to internships or whatever, like or like maybe interview some places and they come back. So you get sort of bits of information from everyone to kind of piece together what the data landscape looks like. Probably not being in Clickhouse and now I see everything. Like as an investor, you see everything too. The one vector database company that I know is doing very well is TurboPuffer. And they are hyper-specialized in doing vector search at a cost performance ratio that's much better than everyone else. So I don't think that the vector databases are going to go away. I think that they'll evolve in two ways. They'll have to become either a sort of a general purpose system like a Postgres, like a MySQL, where they become the system of record where you're storing the original tuples plus the embeddings of the vectors for them. Or they become like an Elasticsearch where there's a separate system where they have a copy of the data that's being pulled from the operational side. And in that case, they can live sort of comfortably as being this additional thing you add on. And if you want the raw best performance of vector search, You some cases you may have to go to one of these specialized systems. So I don't think that's that's going to go away. I just don't think I've seen predictions of like, oh, Postgres is going to die at the hands of a vector database. That's not happening. That's not happening.

Matt Turck

Very much the opposite, right? Yes. Graph databases, we mentioned at the beginning. So, you know, not to become them, but like Neo4j has been around for 20 years now. Sure, yes. And this was supposed to be the moment, right, for graph databases. So what's happening there?

Andy Pavlo

I say I have an outstanding bet with somebody on Hacker News where they said that by the year 2030, the graph database market was going to be was going to overcome be larger than the relational database market. And if this becomes true, then I will wear a shirt that says I love graph databases and I will use that as my driver's license, my university ID. I'll put on my website to the day I die, right? I'm pretty comfortable. It's 2026. We got four years ago. This is not happening. No, it's always been a niche market. And I think that because my perspective on the research side, the research shows that if you do certain things in implementing the engine, which Clickhouse does do, DuckDB does some of this as well. There's things you can do that allow you to do the traversals of graphs, which essentially just joins, soft joins on the table. You can implement those things very, very efficiently and you can easily outperform Neo4j. And that's kind of like kicking. That's like Neo4j is like saying you're faster than Neo4j is like saying I'm faster than somebody maybe like that's you know that's in a wheelchair, right? You can run fast. That's it's it's a low blow. So but I'm just saying all the graph databases, I think like even the best ones, you're just not going to relate to a system that has a bunch of these optimizations that are in the research and actually appearing in some of these systems now. You're just, you're going to lose. What you will lose on against the graph database if you're doing the graph traversal with the client side and the server side, meaning like I, you know, I got to figure out what the next node I want to go look at. I go back to the client and that size the next node to go traverse. If you're doing that back and forth, yeah, they'll beat, they'll beat you guys. But like I said, the SQL standard now supports property graph queries. Oracle has this, right? They were a big pusher of this, this extension of SQL. It allows you to do that traversal on the server side. So like graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.

Matt Turck

GPU databases. Yes. I know you have a special interest there. There was a cycle when a generation appeared, then it went away. There seems to be a renewal. What is a GPU database and what is your prediction? So

Andy Pavlo

GPU database is a data center system where the execution end for queries is offloaded to a GPU running on the PCIe or running in the same box or another box. So, I mean, the history of people trying to build accelerators for data systems goes back to the beginning of data systems. In the 1970s, they were called database machines. So people would build specialized hardware to run sorting and query execution operators. And that obviously died out in the early 1980s because by the time it takes you to like design and fab new specialized hardware, Intel or Motorola will put out the next CPU or the hardware got better and just the gains you were getting went away. So hardware accelerators for databases basically died out in the 1980s. There wasn't a lot of activity in the 90s, 2000s, you saw sort of the rise of people trying to do FPGAs for databases. And every so often that comes back now. Some of the cloud vendors do a little bit of these things, but usually like to filter things on the NIC on the network side of things coming in. So for GPU databases, again, there was a bunch of systems in the 2010s that were trying this. We did a seminar series at the university where we invited all the GPU database guys come to give talks about what they were doing, why they were faster than existing systems. And the big challenge at the time was with those systems, you had to put the entire database inside the memory of the GPU. Because if you had to go back up through PCIe, it was just way too slow. Andy Pavlo bunch of those startups sort of fizzled out. Some of them are still around, but they're sort of specialized for doing visualizations. And then there was in the last year or so, NVIDIA has basically gobbled up a bunch of these GPU database companies that were kind of like struggling along. And I was an advisor for one of them called Voltron. But they also picked up HeavyDB. And so NVIDIA is all going all in on this now. So it remains to be seen whether the idea that you're going to build a CPU-only database system, like long-term, whether that's going to still hold. I've heard getting mixed reports like this is public. Microsoft has, you know, their, they have offerings now in the cloud that can be accelerated, you know, for your data system can be accelerated GPUs for analytics. Another major database company that I can't say who they are, they looked at the economics of GPUs and decided it wasn't worth it. So, one of my former students now is a professor at University of Wisconsin. They're now on leave at NVIDIA. They have a project called SirisDB, which is not necessarily a new data system, but it's a layer in between an existing system and like the CUDA. And so, it supports taking DuckDB queries and running that down on the GPU. I think they can do this in Doris or Star Rocks and DataFusion. And so, at Click House, we've been potentially looking at this as well, but it's research. I don't like it, it's interesting to see whether how much you have to do translation between how Click House expects things and how CUDA wants things to be, the data layout and so forth. How do you organize memory or share memory between these different components? TBA remains to be seen whether this actually makes sense. But certainly, there's a lot of research energy behind this. And publicly, I can say this: like NVIDIA is obviously pushing this because it'll sell more GPUs, right? Because it's hard enough to get new CPUs. Everyone's compute-bound, or memory is hard to get. The computing hardware is very expensive, hard to get now. And GPUs, of all the things, is the most expensive hardware to get. And now you can say your entire database is going to run with GPU. I don't know if that makes sense, at least in the short term. But if the performance improvements are quite significant, and some of the research shows that it is, maybe it makes sense.

Matt Turck

Is there an emerging category or maybe a niche somewhere within a category that people don't talk about enough yet? Of databases? Yeah.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from The MAD Podcast with Matt Turck