Latent Space’s latest episode brings together Zico Kolter and Gray Swan CEO Matt Fredrikson to discuss how red-teaming is changing after Anthropic’s Mythos work. Their central point is that AI security is not simply cybersecurity with a model attached.
Modern model systems can fail through jailbreaks, tool misuse, policy gaps, and emergent behaviors that do not map neatly onto older threat models. That makes evaluation, adversarial testing, and mitigation design a more specialized discipline.
The conversation is useful context for teams deploying agents or high-capability models: security needs to be designed around model behavior, not bolted on after a standard application review.