With 96% of senior cybersecurity leaders reporting that attacks driven by artificial intelligence have become a persistent threat to networks and infrastructure, penetration testing and red teaming remain important tools for identifying vulnerabilities, strengthening defenses, protecting data and reducing risk.
At the same time, AI is altering the nature of pen testing, changing the types of skills and abilities that successful red teamers need, as virtual chatbots and frontier AI models change the cybersecurity landscape by increasing the size of the attack surface organizations need to defend.
Consider a recent Microsoft announcement focusing on the company’s AI Red Team and detailing new ways to address security issues that come with the development and release of frontier AI large language models (LLMs) that can quickly uncover vulnerabilities in applications that attackers could theoretically exploit more quickly.
To counter this, Microsoft announced an External Red Team Alliance (EXTRA), an extension of Microsoft’s AI Red Team designed to support and expand external expertise to advance AI safety and security testing. At the same time, Microsoft is creating a network of specialists who can participate directly in red teaming focused on specialized areas requiring deeper expertise.
“AI red teaming is becoming more interdisciplinary, multilingual, and globally distributed. The expertise needed to identify meaningful failure modes increasingly lives across universities, independent research communities, and regional specialists,” according to Microsoft.
At the recent Black Hat conference in Las Vegas, industry experts noted that issues ranging from LLM integration security to agentic AI-driven offensive and defensive mechanics are driving a new approach to AI red teaming, while also reshaping notions of what pen testing and offensive security mean for organizations.
“AI has fundamentally changed the velocity of red team engagements,” Ryan Powell, vice president for IT and security at ThreatDown, told Dice. “The speed of vulnerability discovery, the scope achievable, and the breadth of coverage are unlike anything possible before. That means testing a company's defenses isn't just about whether they can stop an attack anymore; it's about whether they can keep up with one.”
How AI Is Changing the Role Red Teaming Plays
Penetration testing and red teaming have served a specific purpose for years. In this capacity, red teams would simulate real-world threats, probing systems for weaknesses and revealing them to an organization so cybersecurity and IT professionals could address the vulnerabilities before adversaries exploited them, reducing risk.
Deploying AI agents and platforms within corporate networks – some authorized and some not – has changed what red teams need to look for and how they approach their assignments, said Aviv Nahum, co-founder and CEO of Above Security.
“Enterprises now run AI agents with employee-level access, sometimes greater, but those agents are monitored like service accounts: no behavioral baseline, no oversight program, rarely deprovisioned,” Nahum told Dice. “That's a coverage gap, and a modern red team exercise must include it. If you're only testing whether a human can be phished, you're testing last decade's enterprise.”
Allistair Greeves, director of red team operations at Bugcrowd, observed that the term red teaming has moved far beyond its original definition of covert, objective-based compromise of an organization and now encompasses AI safety evaluations, bias testing, prompt injection research, and broad vulnerability assessments.
“These are legitimate disciplines but are not the same thing,” Greeves added.
At the same time, Greeves noted that AI technologies have transformed how red teams operate and deploy their tradecraft.
“Reconnaissance phases that used to take 10 days of effort are now complete in one day through AI-augmented automation layered on deterministic tooling, raising leads to the red teamers and providing customers with a level of insight and intelligence that is unprecedented,” Greeves told Dice. “In the same allocated time, we can show results in more findings and evidence of root causes.”
The use of AI also allows red teamers and pen testers to research obscure blogs and techniques more extensively, conduct target acquisition, construct contextualized phishing pretexts from gathered intelligence, query threat actor databases to build realistic scenarios, write bespoke malware, implants and tools, and identify patterns across large datasets that were previously cumbersome to navigate, Greeves said.
At the same time, organizations that want to deploy red teams throughout their network to test for weaknesses are increasingly asking for more than a list of vulnerabilities and fixes.
“Organizations want to know what an autonomous hacking agent looks like on their network, what indicators of compromise and traffic patterns it generates, whether their detection stacks, which were tuned for human adversaries, would notice an intrusion,” Greeves noted. “Others want us to test whether their own AI tooling, [Microsoft] Copilot deployments or internal agents that are hooked up to sensitive databases can be used by a rogue agent or be leveraged for mass data exfiltration. Thus, the red teamer of 2026 must operate as an adversary and infrastructure architect, AI and tactics, techniques, and procedures [TTP] operator, and be AI-assisted while being an assessor and critiquer simultaneously.”
Adapting to AI Red Teaming
As attackers deploy more sophisticated techniques, red teamers and pen testers need to adapt. For example, with static attack libraries and single-prompt jailbreaks obsolete, real-world attackers typically chain their attacks instead of running them one at a time, said Gal Moyal, who leads the CTO Office at Noma Security.
Red teamers need to adjust.
“Effective red teaming now requires dynamic, multi-turn testing that can simulate this by compounding multiple techniques into a single attack sequence, raising the attack success rates until your application's real weaknesses are surfaced. However, finding vulnerabilities is only half the battle for security teams,” Moyal told Dice.
Other experts note that red teamers’ skill sets need to expand as AI changes both the tools used to uncover vulnerabilities and the ways attackers use those same technologies.
“The next great red teamers will think like behavioral analysts, not just exploit developers. Social engineering now includes engineering the AI – understanding how models reason, where they fail, and how an agent's permissions can be bent to a purpose no one intended,” Nahum said. “The encouraging part: unlike humans, agents have a finite set of possible actions, which makes them testable in a way people never were. The red teamers who learn to map that action space – and to think about intent and business context, not just access – will be the ones who stay ahead.”
Bugcrowd’s Greeves noted that all red teamers still need curiosity, broad exposure, an understanding of TTP, and deep knowledge of technologies, processes, and people.
“There is also the need for operational security and risk considerations to demonstrate impact without introducing significant risk, as well as the likelihood of attack paths that may previously be unknown to the organization or underestimated,” Greeves added.