Responsible AI: Inside SRI’s race to make generative models safe

Key takeaways:
- Generative AI is everywhere, the challenge is getting people to trust it.
- SRI is using mechanistic interpretability, essentially neuroscience for machines, to understand and correct AI behavior.
- The most promising defense against AI cyberattacks is making them too expensive.
- SRI’s AI Center celebrates 60 years since its founding.
Generative AI has gotten remarkably good at many things. Ask Eric Yeh, a computer scientist in SRI’s Artificial Intelligence Center (AIC), and he’ll tell you about the time his team trained a model to turn architectural sketches into fully rendered building designs that can be used in real projects. The deeper lesson, Yeh adds, is one the industry is still learning: impressive output doesn’t automatically translate into adoption.
Yeh’s team turned the technology into a patented AI toolkit called AiCorb and are starting to see some uptake in usage. What AI can do and what people are willing to trust is the throughline that connects what SRI’s AI Center is working on.
Teaching machines to imagine
One of the applications is what the team calls counterfactual exploration: using generative models as “proxies” for complex real-world systems, like a delivery drone deciding whether to fly through a thunderstorm. Rather than test every scenario in the real world, the system can generate plausible alternatives and attach a confidence level to each one. Autonomous vehicles face the same problem; the open road is an environment that a test suite can’t fully cover.
Decision-making
A second research thread is more philosophical. Mechanistic interpretability, sort of like neuroscience, aims to understand what’s happening in the LLM as it reasons. If researchers can identify the internal signals tied to a bad decision, they could nudge the model away from it before it acts.
What makes SRI’s approach distinct, Yeh says, isn’t any single technique — it’s the mix of people in the room, varied disciplines, and six decades of research. AI Center researchers collaborate with SRI’s formal-assurance specialists, who mathematically verify whether a system will behave as expected, and its computer security researchers, who build threat models around how systems can break. “We work with computer systems, and everyone’s working with AI,” he says. “The thing is how you frame the problem — so you can control an LLM’s behavior and get guarantees.”
Making cyberattacks too expensive to run
That security lens matters more as AI systems start acting autonomously. There are bad actors, some operating from jurisdictions beyond the reach of U.S. law enforcement, who are building malicious AI tools. Rather than chase a defense that may never close the gap, SRI is figuring out how to raise the cost of an attack: feeding an adversarial AI system bad information, so it burns more compute trying to succeed. “If you can make it economically infeasible for them to apply these techniques,” Yeh says, “the balance of power shifts.”
Old rules for a new era
For people worried about deepfakes, hallucinations, misinformation, digital twins, and AI-driven scams, Yeh’s advice is refreshingly low-tech: be careful what you put online and even more careful about what you believe when it comes to you.
It’s a fitting principle SRI has championed for decades. Longtime SRI scientist Peter G. Neumann, who passed away earlier this year after spending more than 50 years at SRI in the Computer Science Lab, spent his career arguing that computer systems can never be fully trusted by default — a rule the industry relearns with each new wave of technology.
“The fundamental principles — they’re still the same,” Yeh says. “The tools keep changing, but really, the discipline doesn’t.”
SRI is advancing security & defense technologies
to address next-generation threats and challenges
SRI’s Artificial Intelligence Center conducts applied research in generative AI,
autonomous systems, and AI safety and security.