The agentic AI playing field was heavily tilted toward offense, so researchers began using red team agents to help teach their blue counterparts.
July 29, 2026
Testing AI-based security systems can be tough amid growing fears about agents cheating, hallucinating, and escaping containment, but a group of researchers believe they've found a way to better measure the effectiveness of agentic defenders.
Earlier this year, Dreadnode, an AI offensive security startup, released two open source tools designed to help users evaluate the security agents deployed in their networks. The first is DreadGOAD, a reproducible Active Directory training environment that's designed to replicate "the messy deployments still common in large organizations," according to the company.
The second is Ares, an agentic red team-blue team system designed to test and study offensive and defensive effectiveness. When Ares is deployed in DreadGOAD, the red team agents discover hosts, identify weaknesses, escalate privileges, and ultimately compromise the environment. The blue team agents, meanwhile, are tasked with analyzing telemetry, triaging alerts, investigating the activity, and hopefully containing it.









