Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
AI securityRed teamingPrompt injectionAutomated red teamingAdversarial attacksAI interpretabilityAI agentsEnterprise AIAI policy enforcementAI risk assessmentAI underwritingAgent identityAI defense modelsMechanistic interpretabilityAI safety research
This episode features Zico Kolter and Matt Fredrikson from Grey Swan discussing AI security challenges, particularly around adversarial attacks and indirect prompt injection in large language models and AI agents. They explain their approach to red teaming, automated vulnerability detection, and defense mechanisms like their SIGNAL filter model, emphasizing the evolving landscape of AI safety in enterprise deployments.