Red Team
A group that simulates real-world attacks against an organization or AI system to test and improve its defenses.
A red team↗ is a group, internal staff or hired specialists, that simulates real-world attacks against an organization or AI system specifically to test and improve its defenses, acting as an adversary would rather than following a scripted checklist. Their goal is to find the gaps a defensive team might miss by thinking creatively about how a real attacker would actually try to break in.
Red team exercises are often run against a "blue team↗" (the defenders), and the resulting "purple team" collaboration, comparing what the attackers did against what the defenders detected, is one of the most effective ways to sharpen real-world security readiness. In AI safety specifically, red-teaming has taken on a similar meaning: deliberately trying to make a model behave badly (leak data, produce harmful content, bypass safety rules) to find and fix weaknesses before real users encounter them.