BP Bharat Patel On ChinaTalk

“I awarded four direct to phase two to companies to generate synthetic data. And then the first thing that the company said, hey, do you have representative data? And I'm like, oh, crap.”

ChinaTalk · Geopolitics & World Affairs · October 2026

“I awarded four direct to phase two to companies to generate synthetic data. And then the first thing that the company said, hey, do you have representative data? And I'm like, oh, crap.” — Bharat Patel, ChinaTalk

The guest was describing an Army effort from around 2022 to pre-train AI models on synthetic data before the real systems were fielded. The lesson he drew is that synthetic data depends on real data to copy, so it can add to operational data and fill in edge cases but cannot yet replace it.

Transcript

ChinaTalk Around 28:43 into the episode
Speaker 1

yep. So on that, this is where it gets very interesting because wartime quick reaction capabilities they pop up like this and we can get them on contract very quickly. And that's because, you know, some policies are just kind of waived off. But now we're in peacetime. Peacetime, there's a lot of federal acquisition regulations, competitive prototyping, competitive source selection activities that you just have to follow. Again, a lot of that starts to shift when we're at QR quick reaction capabilities and you have to field stuff that will save or deter our enemies and threats. So a lot of things happen much quicker in wartime. And then there's a lot more resources behind it. Not just your regular service budgets, but then you get contingency operations resources that allow you to, like, no kidding, just move a ton faster. Contracting happens a lot faster. And then all of a sudden, there's a lot less concern about modular system architectures. There's a lot of less concern about switching out, being able to move different components. There's more concern about let's move things fast and we will figure out the seams later, but as long as we could save lives.

Speaker 2

Yeah. Let's come back to data. Synthetic data. People are excited about it. People are worried about it. How useful is this for the sort of targeting autonomy? Physical real world stuff that we're talking about in this context.

Speaker 1

That is, so I, my synthetic data journey has gone, you know, like this. I would say my synthetic data journey started, I think, around 2022 when I was supporting the Titan program. Tactical intelligence, targeting access known, ground station, next generation capability. One of my theories or hopes really was to try to accelerate the AI. So before we even fielded some of these capabilities, I'm like, hey, let's see if we could pre-train models. Let's see if we could do some things. So I launched, I was able to launch partnering with the Small Business Innovative Research Program. I'm lucky I was the first one to use this direct to phase two concept. I awarded four direct to phase two to companies to generate synthetic data. And then the first thing that the company said, hey, do you have representative data? And I'm like, oh, crap. So going back to data is very important. And synthetic data is really dependent on the actual data. So then they could replicate that data. Right. So we need to have enough representative data in order for synthetic data to actually work. And then there's definitely variations of maturity across our industry partners when it comes to synthetic data. But when it comes to, you know, autonomous vehicles, commercial industry, they are investing a ton in synthetic data. NVIDIA is doing some things. So the technology is getting there. But when applied to sensors that the Army has, that's different than commercial sensors that is on either airplanes or on some of these new autonomous vehicles. The military sensors could be two or three years older and it could be like different, completely different, different look angles, different biomes, different experiences. So just trying to be able to give enough of that data to a synthetic data company for them to kind of replicate those scenarios. So we aren't quite there yet from can you build AI and autonomy off of simply synthetic data? I don't think we're quite there yet. I think synthetic data augments operationally relevant data or actual data. So the combination of relevant data and synthetic data will get you a little bit more performance. And then synthetic data, the way I like to use synthetic data is try and figure out those edge cases that you don't necessarily have operationally relevant data for. So then replicate these edge cases that allow us to, again, see what the model does against these weird situations that we just haven't seen before.

Speaker 2

So another difference between training Tesla autopilot and your tank model is like no one is walking around the street trying to like expressly trip up Tesla into driving into people on the street, right? Which is not necessarily the case if you have a human being try to trick a robot trying to shoot them, right? So talk a little bit about the sort of adversary dynamics when it comes to data and what folks should, I don't know, expect or be prepared for in the coming years. Is data poisoning something to be particularly worried about here? Or how do you define that phrase?

Speaker 1

Oh, man, in many ways. I think we should be worried about it 100%. So in my mind, there's like two scenarios there. There's the active poisoning that is happening right now. So right now, there's no question that the open source environment, open source technologies, open source databases, open source frameworks for models, they are out there to help accelerate AI. But there's also people that are contributing to the negative aspects of that in terms of like just putting in bad data for fun, right? You know, sometimes I equate it to not exactly, but I equate sometimes I'm like a jerk and depends on my mood. And I use Waze and then not here. Cop's not there, not there, not a crowd, even though he's right there, right? I'm just being a jerk.

Speaker 2

Just like graffiti, right? Yeah.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from ChinaTalk