Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers
Modern frontier AI models deploy safety classifiers. These are second AI models that sit between the user and the frontier model, evaluating every request in real time. If a request is flagged as harmful, the classifier blocks it before the model can respond. Significant investment and safety model expertise have made these classifiers effective. Anticipating how adversaries circumvent these systems is a security problem that requires different expertise.