KG

Keerthana Gopalakrishnan

Things Keerthana Says on Podcasts

Where to Find Them

Keerthana Gopalakrishnan has been a guest on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (4 times) .

Recently: “One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids” on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (October 2026); “Gemini Robotics – AI for the Physical World, with Keerthana Gopalakrishnan and Ted Xiao of Google DeepMind” on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (May 2025); “Robotics Research Update, with Keerthana Gopalakrishnan and Ted Xiao of Google DeepMind” on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (April 2024); “E12: The Robot Revolution with Keerthana Gopalakrishnan of Google Robotics” on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (March 2023).

What They Said

“I think still GPT-2. … Robotics weirdly is still very subject to cross-embodiment, right? If it just works on your robot with your specific setup, is it really a generic brain?” — Keerthana Gopalakrishnan, The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

The host asks where robotics sits if you score the field in "GPTs" of progress. Gopalakrishnan says GPT-3 would need few-shot learning to work well across many tasks, and a model that carries over to robots it wasn't built for.

The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis · 2026-10-03 Permalink → Listen →
The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis Around 16:31 into the episode
Keerthana Gopalakrishnan

Yeah, I think it's going to be a spectrum, definitely. Being able to ICL or in context show a robot how to do a task and then it doing it definitely reduces the time to deployment and the time to pick up the task, right? And you don't need any specific fine-tuning. You don't basically need to train it. You can do it in context. And so that definitely is a very exciting development. Although you can think of it as video prompting, right? Even the older models, like when I say, let's say, pick up the object, right? And it is a very unseen object. What am I doing? I'm sending in language and then I'm getting a very general behavior out of it that I did not show it before. So that's generalization. Now, instead of prompting with text, you can prompt it with an image. You can say, here's the thing, and then draw a circle around the thing that you want manipulated. And that's prompting with image. And this is prompting with video in some sense. And so I would imagine, I think the ICL results currently are on the spectrum about how do you prompt a foundation model. And you can prompt it with video, you can prompt it with language, you can prompt it with image. And then the question is: so now what is the generalization that you can get? Now, these models are also very subject to kind of memorization in some sense. So if you really show it exactly, this is what to do, they will copy it. But the question is, can they generalize, right? Can you now, if you change the scene and stuff, can they do, can they now adapt? And then the second question is, so for the same test, and the second question is, what is the level of difficulty of the task that you are prompting for? A lot of pick and plays, like the models have a lot of data and it's also like easier to comprehend. But can you show a robot to tie a trash bag? And then can the robot tie a trash bag after you show it? So I think it's still fairly early and we need to see how that evolves.

Nathan Labenz

If you had to score, this is a silly question, but it might be useful for just calibrating. If you discore robotics today on one to six GPTs, are we like in earlier conversations, I think we were, yeah, we're maybe hitting like a GPT-2 kind of moment. I think of GPT-3 as being really simple. Synonymous with in-context learning, but we're following a somewhat different path with robotics, obviously, where we have like instruction following built in, maybe in many cases, even before like good in-context learning happened. So that just exposes the flaw in my question. But if you have to analogize to how many GPTs we are along the way in robotics, where would you score the field today?

Keerthana Gopalakrishnan

I think still GPT-2. And here's why. I think for GPT-3, we need, firstly, few shot learning to work really well for a lot of different tasks. And secondly, also, robotics weirdly is still very subject to cross-embodiment, right? If it just works on your robot with your specific setup, is it really a generic brain? Now, can I put that brain on my humanoid or my another robot or some other robot that I just bring in? It should be, and if it's completely helpless in that setting, is that a generic brain? I think GPT never had this problem, right? Like we, my phone or your phone, my computer, Mac, Linux, it doesn't matter where you run it. It kind of behaves. You can expect this very similar behavior. But here, I think we are very subject to which robots that you act on. And so I think there is a lot of work still needed to be done to make very generic brains that can count discount those factors out.

Nathan Labenz

Hey, we'll continue our interview in a moment after a word from our sponsors.

Speaker 3

Today's episode is brought to you by Athena, the executive assistant company on a mission to improve how people work and live. If you want to increase your impact, you have to free up your time. And that's what Athena does best. They match you with a dedicated, full-time, top 1% executive assistant who can take over your inbox, calendar, travel, and everything else that's quietly eating up your week. Athena is SOC2 Type 2 certified, so you can rest easy knowing that your sensitive data is in good hands. And as a former AI advisor to the company, I can personally vouch for how much they've invested in AI tools and training. In fact, one of the very best AI users I've ever met is an Athena client who delegated the exploration of AI tools and the development of AI workflows to his EA. Athena clients report saving an average of 15 hours a week, and the average client refers more than two friends a year. That, to me, checks out. I was an Athena client while running my startup, and to this day, I continue to refer friends. Go to Athena.com/slash cognitive right now and get matched with your EA. That's athena.com/slash cognitive. Give yourself back a few hours this week. Go to athena.com/slash cognitive and see who they'd pair you with this month. Today's episode is sponsored by Parallel, where agents find answers. Most engineers today closely follow new model releases, but don't pay nearly as much attention to their agent's most important tool, web search. If you're like me, your agents are often conducting hundreds of searches per day. And while the unit cost is small, over time, this does start to add up. Before starting to use Parallel, I calculated that my agent search bill would total roughly $500 this year. The good news is that I just had my agents conduct a systematic test, and I found that Parallels fast mode, which costs just $1 per thousand queries, worked just as well as my previous default provider. And switching to it will save me roughly 80% of my search bill going forward. Parallels infrastructure is enterprise-grade, and their suite of APIs offers a range of Pareto optimal options that allow you to choose the right balance of quality, cost, and speed for your needs. So whether you're building voice agents that need 200 millisecond latency or long horizon agents that need deep research, adding Parallel to your agent's toolkit is a no-brainer. Get started for free at parallel.ai slash TCR. That's parallel.ai slash TCR.

Speaker 4

Then maybe that's a perfect transition to

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In