The TWIML AI Podcast · AI Builders · October 2026
Almeida was contrasting self-driving cars, which he sees as a win for careful engineering of reliable parts, with how most AI products get built. He added straight away that he doesn't mean it as an attack on those startups: with models as unreliable as they are, keeping a human in the loop is the right move.
don't believe in running from questions. So I will try to get them both. But actually, I'll just do the first one. I will ramble. So I want to focus up on that one. I think a lot of what AI does is super freaking sick. I use coding agents. I don't even code that much anymore, but I use coding agents every day. I use chatbots every day. Holy smokes, would it be a lot harder to run a company without chatbots? But they are.
In case it wasn't obvious with the Tesla example, I didn't say the punchline, which was the car was driving itself at the time. Oh, right. Yeah.
Self-driving, actually, I think it's a little bit different. And we can talk about the distinction between self-driving. That was only sci-fi. Oh, cool. My argument with sci-fi is especially related to LLMs, which I think tends to be what people mean by AI these days, because it seems to be like the greatest compression of intelligence while having the least utility. So, and also, it's quite, I wouldn't say it's like universally hated, but there's a lot of dislike for AI, which is tragic. And I think that we would be remiss as a field to not own up to why that is the case. So, there's like a lot of like dark sides to AI too, and I could like name a subset of them, but I think that that's worth talking about. So, number one, I think self-driving is super duper cool. And also, I see self-driving as a ginormous engineering win, not necessarily an AI win. I use Waymo as like the gold standard here. No offense to Tesla fans, but like they seem to be safer. And also, the way they do it is by building a la software and engineering, like reliable systems that they understand the pieces of that have their own like decomposed abstracted parts and make sure those are like unbelievably good while programming in the behavior of the policies you want such that this can generalize outside of distribution. So, this is super freaking cool, but none of this is happening in AI right now. When in AI, what you tend to get is a lot more, let's just make a demo. And you know, the my cynical loop here is make a demo, raise like a Cedar Series A, say that you're going to make it reliable, never end up making that reliable, pivot into a human-in-the-loop version of this thing instead of actually automating the task. And it's not, I don't want to hate on them, to be clear, that is actually the right move to do given the current AI climate because the models are not reliable enough. So, even you know, I try to be as unbiased as I possibly can. I obviously can't be perfect, but I would like to have like the same type of critical lens to LLM demos as Jev demos. And I, you know, like, I think that a lot of them are super cool, especially when they have sick ideas. And also, I can't really vouch for them because they solve work when they actually reliably can be run in the background. And LLMs are extremely not there. I actually don't know of anyone who like actually runs like an LLM as a separate dependency because you can't get abstraction with an LM, right? It can just break arbitrarily, and you need to have a abstraction needs to leak in order to see, like, oh, why did you talk about biology? Oh, I'm sorry, my claw can't talk about biology. It fell back to another model or something like that. And that is extremely sensible for a first-party app, but super nonsensical, in my opinion, for something developers are meant to build on top of. So, I think that there's that lack of usefulness, I still think, is present there, despite demos looking like sci-fi, because demos will always be significantly ahead of the curve. Opening, I have seen, has been showing like customer service demos since 2020, I believe. That is now roughly six years ago, and customer service is still not solved. And maybe because the work is hard, as you say, but it really looks a lot easier than Millennium Prize problems in math, because I don't know anyone who could solve a Millennium Prize problem in math. I feel like I know everyone, and I don't know everyone who could solve customer service, but everyone I know probably could do that job better than an LM can, or even drive-throughs. People keep trying and failing to automate drive-thrus, and like there's no like good answer for that other than maybe drive-thrus are harder than unsolved math.
It's interesting to me that you call out Waymo versus Tesla in part because in the context of this conversation about like bitter lesson-pilled, right? Because, in one sense, like what differentiates Waymo is that they're not taking a bitter, less and pilled approach to it, whereas Tesla's is more bitter-less and pilled. Like, we're gonna, you know, I'm characterizing their approach at a 30,000-foot view, but like, we're gonna collect a lot of data and throw that into a big model and let it do end-to-end. Whereas Tesla is we're gonna like classically engineer this thing, break this thing into a lot of different modules, use all the sensors we can to get as much data as we can. I've had conversations with Drago on the podcast about like how they approach it and how it's so fundamentally not like you know, we're aiming for this end-to-end single model to rule the entire driving experience.
So, I want to emphasize that that is just the MLP. Part of it, but I don't think that's the limit of what we should be building. And I'm actually extremely pro software and software engineering. And I actually think that I hope that, and I want to do everything I can, that Jev will usher in another golden era of software engineering, kind of like with the early internet creative energy of like building crazy things. And for ML, you need data, you need to focus on the right task, you need compute, you need algorithms. But ML is meant to be one part of a larger system, in my opinion, in general. And ML should be scoped to the part of the system that it can really do well. And in general, the way to get higher reliability is to zoom in. You know, like would Waymo be as successful in San Francisco if they tried to do the whole thing end-to-end? Honestly, I would wager not. Like, maybe this is going to be better in five years, but in order to actually ship something now that is useful and reliable and safe, I think you need engineering to study all of these things, study the properties of every single ML component. And when things break, be able to debug which part caused it and fix that. And that's very hard to do in a one model rules them all type setting.
So where did the idea for Jev come from? Like, was this like, hey, I'm going to set up TypeSafe and I'm running towards this thing that I already knew about? Did it evolve? Did you pivot into it? Like, how did it happen?