Hands-On Engineering Podcasts · August 2026

“My opinion there is you should never connect directly to Postgres.” — Shaun Thomas, Postgres FM

The episode is about estimating work_mem, and it detours into connection limits: Nikolay Samokhvalov notes that RDS defaults to thousands of max connections and that plenty of customers run with no pooler at all. Thomas answers with a flat rule rather than a tuning tip. He goes straight back to sampling active connections afterwards.

Transcript

Postgres FM Around 05:17 into the episode
Shaun Thomas

No, I don't. What will RDS do? Max

Nikolay Samokhvalov

connections 5,000. Even for small clusters. Yes. And they expect that everyone will probably use RDS proxy, but we observe a lot of customers who come to us, they don't. No pitcher bouncer, no RDS proxy, only some poolers are on application side. And those guys who run application nodes, they put it to Kubernetes with auto-scaling. So nobody knows which probability of reaching those 5,000 connections.

Shaun Thomas

Yeah, you could have a whole other conversation on pooling. My opinion there is you should never connect directly to Postgres.

Nikolay Samokhvalov

Right. Well, yeah, pooler should be inside. Yeah, I agree. But this is the reality. And RDS is the most popular one. And people come with 16 V CPUs, 25 or 50 hundred Max connections.

Shaun Thomas

Yeah. At that point, you get to an area where you have to say sample your number of active connections and then it's a little more involved, but it's roughly the same idea.

Nikolay Samokhvalov

Yeah, so this is great. And this is like for many years it was my number one approach. Basically, we tune based on observations. Feedback loop, right? But I really want something more like not looking into actual thing, but predicting something, right? Because sometimes you launch new service or sometimes you expect some growth. And observing production just doesn't feel super. It's very practical, of course. And it can work either way as well. For example, if we take your formula, divide by five, and even drop that five, just no additional dividing at all. And then we say, you know what, like we still need more memory. And we just observe that there's a lot of page cache is huge. Because workman, it's not allocated immediately, right? It's allocated in chunks, like it's gradually. So workman is a limit for one operation inside query. But it doesn't say that run a query, each operation will have workman, maybe less. And if we observe from reality that we actually use much less, like we can overcommit here, right? This is what we do. Like we say we go beyond theoretical limits because we see that practically we don't have out-of-memory risks. Theoretically we have them, but practically no. And I know such clusters, huge clusters, they are running just ignoring this formula. And it's okay because majority of queries, they actually use a tiny amount of workman. What do you think about this? Overall, it's like a bag of very different ideas. And how to find a better path to some recipe.

Speaker names from our own diarization · position estimated from where the line sits in the episode