Hey hackers - agent red teaming outgrew the prompt fuzzer somewhere between the first tool call and the first exploit chain that never touched the user's input.
The scoring harness now decides what "secure" means, and modern agent red teaming lives or dies by what that harness can see. In 2026 it decided wrong in public, twice.
TL;DR: Microsoft's PyRIT passed the CSA/OWASP evaluation at 95 percent automation for core red-teaming workflows but only 40 percent for system-level agent validation - it scores what the model says, not what the agent did. An agentic red teaming platform compressed weeks of library-driven work into three hours: 674 attacks, 7,727 trials, 85 percent attack success against Llama Scout, 232 critical findings.
AutoDojo broke the static-defense fantasy - a filter holding a 0 percent static success rate fell to 28 percent once attacks adapted, 64 percent on action-open tasks. Neighbors: tooling, scanner-passing kits, what leaks.
The scenario layer also made the economics explicit. RapidResponse brute-forces every technique against every objective and burns budget doing it; TextAdaptive picks the next technique per objective from historical attack success rates stored in PyRIT's own database with an epsilon-greedy selector, cutting the search from techniques-times-objectives to a three-attempt cap. Every scan an organization runs sharpens the next one - agent red teaming with a memory, the way an attacker already works.
The CSA and OWASP AI Exchange evaluation put numbers on where that tooling stops. Ninety-five percent-plus automation coverage for prompt execution, scoring, logging, reporting. Eighty percent support for prompt-level agentic testing - role escalation, memory poisoning scenarios, goal manipulation, all scored against textual responses.
Forty percent for system-level validation, because PyRIT does not natively observe tool invocations, state changes, or persistence. It is a force multiplier with a documented edge, and the edge is exactly where agents live.
The case study against Llama Scout ran 681 assessments, 674 attacks, and 7,727 trials in roughly three hours of wall-clock time, landing an 85 percent attack success rate, 401 full jailbreaks, and 232 critical-severity findings, with compliance mapping exported without a human writing a line of workflow code.
The numbers around it explain why the shift happened. Automated approaches hit 69.5 percent success against 47.6 percent for manual across 214,271 recorded attempts by 1,674 participants; AIRTBench's 70 black-box challenges show frontier models solving up to 61 percent with efficiency advantages beyond 5,000-to-1 over human operators on the hard tasks. The operator's job moved from writing exploits to choosing what to probe - and the report comes back tagged, scored, and evidence-backed in an afternoon.
That compression cuts both ways. A system that can run 674 attacks before lunch can also generate 674 confidently-scored findings of which the interesting fraction is small. The wall-clock chart looks like a productivity win and reads like a queue.
This is not a tooling deficiency to shrug at; it is the measurement gap that every agent red teaming program inherits if it starts from prompt-level tooling. The real risk lives in the gap between what the agent says and what the agent can do.
CSA's recommended workflow draws the fix as a pipeline: test refusals with the prompt harness, export the logs, wire agent-runtime telemetry that records actual tool calls, then compare what was said against what was done. Teams that skip the last step publish green dashboards over agents nobody watched.
The same gap scales into the platforms. When an agentic evaluator scores 674 attacks in three hours, its scorers are still mostly judging text and traces - deterministic scorers catch tool invocations and resource changes, LLM-based judges read semantics, and the two disagree systematically.
The RIFT-Bench team found LLM-based attack-success rates consistently lower than deterministic ones, because the deterministic evaluators are tailored to concrete execution signals. An agent red teaming program that quotes only the softer number is grading itself on a curve.
The adaptive loop, a cheap black-box optimizer using a frontier model to iterate injections against the deployed defense, raised attack success against nearly every defense tested across three task suites and five models.
The sharpest cells read like punchlines. PIGuard, a DeBERTa injection classifier, held static success at 0.0 percent and fell to 28 percent overall - 64 percent on action-open tasks where the user's request delegates the action itself to attacker-controlled content. DataFilter went from 12.6 to 33.4 percent, nearly triple. ProtectAI doubled from 7.2 to 15.4.
The authors' phrasing is the takeaway: static-string robustness overstates practical security, and on action-open tasks the injection can pose as ordinary data, which is a structural limit of any defense keyed to instruction-like text.
Two public results made the same point from the field. Nasr et al.'s The Attacker Moves Second broke twelve published defenses at over 90 percent success. And a vendor's own AgentDojo writeup in August 2026 reported a per-message guard taking 46.5 percent down to 0.7 percent while trace-level and containment defenses sat at 47.2 and 43.1 percent against a 46.5 percent baseline - statistical no-ops.
Then came the discovery: one arm was wired after the tool-execution loop had finished, observing attacks that had already succeeded. Forty-six offline tests passed green; none constructed a real pipeline. Placement was the difference between 46.5 percent and 0.7 percent - an agent red teaming lesson no scorer can score.
The probe adapts to the system it is pointed at, which is the whole point: attack suites that only work on the implementation they were written against are demonstrations, not evaluations.
The evaluation substrate matters as much as the attacks. Execution traces, not model outputs: tool invocations, resource consumption, environment state. A multi-agent workflow that routes a poisoned memory into a tool call scores on the trace, and the CSA guide's catalog - memory poisoning, tool misuse, role escalation, goal manipulation, multi-agent attacks - only becomes testable once the harness can see state transitions.
That is also where memory poisoning testing crosses from prompt exercise into systems testing: persistence is a state change, and state changes live in telemetry, not transcripts.
White-box limits are real - RIFT-Bench assumes code access and leans on an observability stack - but the direction is fixed. Output-only scoring is where the field started, and every 2026 framework that matters moved toward traces. System-level agent red teaming has one direction of travel. The question for any program buying or building an evaluator in 2026 is which side of that line its numbers come from.
The scorecard itself needs three bands to stay honest: model-behavior metrics (refusal, compliance, exploit success on text), system metrics (tool-call violations, state changes, egress attempts from ingested content, blast radius per identity), and utility metrics (task completion under attack, because a defense that bricks the agent traded a security incident for a business one).
Report the bands separately. The 2026 embarrassments - the flat-lined defenses, the 0 percent filters, the scores that could not see the firewall - all came from collapsing the bands into one number that meant less than it looked.
Human triage stays in the loop at the top: the 674-attack runs generate findings faster than a team can read them, and severity without context is just a sorted list.
The tools are better than they have ever been, and the failures that made the year's headlines were measurement failures - scoring sentences while agents moved state, wiring detectors after the loop, quoting static numbers against adaptive adversaries.
What holds is the say/do split as doctrine: score what the model says, separately score what the agent did, separately score what it costs the product, and never merge the columns. Add the adaptive re-run as the fourth band, because a defense that only survives the benchmark it was tuned on has not been tested - it has been rehearsed. And keep the human on strategy: automation bought the throughput, and judgment is the part the CSA report refused to automate away.
Run the unchanged scenario against your production stack this week. Whatever the diff says, you will finally be looking at the agent instead of its sentences.
The scoring harness now decides what "secure" means, and modern agent red teaming lives or dies by what that harness can see. In 2026 it decided wrong in public, twice.
TL;DR: Microsoft's PyRIT passed the CSA/OWASP evaluation at 95 percent automation for core red-teaming workflows but only 40 percent for system-level agent validation - it scores what the model says, not what the agent did. An agentic red teaming platform compressed weeks of library-driven work into three hours: 674 attacks, 7,727 trials, 85 percent attack success against Llama Scout, 232 critical findings.
AutoDojo broke the static-defense fantasy - a filter holding a 0 percent static success rate fell to 28 percent once attacks adapted, 64 percent on action-open tasks. Neighbors: tooling, scanner-passing kits, what leaks.
The tooling grew up
PyRIT has run in 100-plus Microsoft AI Red Team operations - Copilots, the Phi-3 release cycle that spent six weeks and a thousand-plus prompts across 15 harm categories - and by 2026 it stopped being a library you assemble by hand. The July 2026 Scenarios release turned the compositions every team kept rebuilding into pre-packaged playbooks: one command, repeatable across model releases, diffable against last quarter's results, tagged with a run ID so a teammate can rerun the exact scan.The scenario layer also made the economics explicit. RapidResponse brute-forces every technique against every objective and burns budget doing it; TextAdaptive picks the next technique per objective from historical attack success rates stored in PyRIT's own database with an epsilon-greedy selector, cutting the search from techniques-times-objectives to a three-attempt cap. Every scan an organization runs sharpens the next one - agent red teaming with a memory, the way an attacker already works.
The CSA and OWASP AI Exchange evaluation put numbers on where that tooling stops. Ninety-five percent-plus automation coverage for prompt execution, scoring, logging, reporting. Eighty percent support for prompt-level agentic testing - role escalation, memory poisoning scenarios, goal manipulation, all scored against textual responses.
Forty percent for system-level validation, because PyRIT does not natively observe tool invocations, state changes, or persistence. It is a force multiplier with a documented edge, and the edge is exactly where agents live.
| Capability | Coverage | What it cannot see |
|---|---|---|
| Core red-team automation | 95%+ | strategy, interpretation (human-owned) |
| Prompt-level agentic tests | 80% | actual state changes |
| System-level agent validation | 40% | tool calls, permissions, memory writes |
The agent took the console
The May 2026 Dreadnode paper is the other half of the story: an agentic red teaming system that accepts natural-language objectives and orchestrates the whole campaign - 45-plus attack strategies, 450-plus transforms, 130-plus scorers across jailbreak detection, credential leakage, exfiltration, MCP security, multi-agent security.The case study against Llama Scout ran 681 assessments, 674 attacks, and 7,727 trials in roughly three hours of wall-clock time, landing an 85 percent attack success rate, 401 full jailbreaks, and 232 critical-severity findings, with compliance mapping exported without a human writing a line of workflow code.
The numbers around it explain why the shift happened. Automated approaches hit 69.5 percent success against 47.6 percent for manual across 214,271 recorded attempts by 1,674 participants; AIRTBench's 70 black-box challenges show frontier models solving up to 61 percent with efficiency advantages beyond 5,000-to-1 over human operators on the hard tasks. The operator's job moved from writing exploits to choosing what to probe - and the report comes back tagged, scored, and evidence-backed in an afternoon.
That compression cuts both ways. A system that can run 674 attacks before lunch can also generate 674 confidently-scored findings of which the interesting fraction is small. The wall-clock chart looks like a productivity win and reads like a queue.
What the score cannot see
The CSA evaluation's sharpest example is a firewall. PyRIT can prompt an agent to disable it and capture the model's response - "I will disable the firewall" - and score that response as compliance or refusal. What it cannot do is verify whether the agent actually called the firewall API, modified a permission, or triggered a downstream workflow. The published test plan computes refusal rate, compliance rate, ambiguity rate, exploit success rate - all measured on sentences.This is not a tooling deficiency to shrug at; it is the measurement gap that every agent red teaming program inherits if it starts from prompt-level tooling. The real risk lives in the gap between what the agent says and what the agent can do.
CSA's recommended workflow draws the fix as a pipeline: test refusals with the prompt harness, export the logs, wire agent-runtime telemetry that records actual tool calls, then compare what was said against what was done. Teams that skip the last step publish green dashboards over agents nobody watched.
The same gap scales into the platforms. When an agentic evaluator scores 674 attacks in three hours, its scorers are still mostly judging text and traces - deterministic scorers catch tool invocations and resource changes, LLM-based judges read semantics, and the two disagree systematically.
The RIFT-Bench team found LLM-based attack-success rates consistently lower than deterministic ones, because the deterministic evaluators are tailored to concrete execution signals. An agent red teaming program that quotes only the softer number is grading itself on a curve.
Split every score into three channels: did the model comply in text, did the tool call execute, did the environment state change. Report all three, never the first alone.
Wire the harness to runtime telemetry before the first campaign - tool-call logs, permission changes, memory writes. PyRIT's own documentation tells you its visibility ends at the response; believe it.
Diff utility under attack too. AgentDojo showed 10-25 percent absolute utility loss under attack with strong correlation to benign utility; an agent that stops working is a finding, not a pass.
Re-run last quarter's scenario unchanged after a model update. If the numbers move, the harness measures the model; if they do not, check whether the harness still executes at all - wiring bugs publish flat lines that read as resilience.
Wire the harness to runtime telemetry before the first campaign - tool-call logs, permission changes, memory writes. PyRIT's own documentation tells you its visibility ends at the response; believe it.
Diff utility under attack too. AgentDojo showed 10-25 percent absolute utility loss under attack with strong correlation to benign utility; an agent that stops working is a finding, not a pass.
Re-run last quarter's scenario unchanged after a model update. If the numbers move, the harness measures the model; if they do not, check whether the harness still executes at all - wiring bugs publish flat lines that read as resilience.
Static benches die on contact
AgentDojo built the field's favorite proving ground - 97 tasks, 629 security test cases, a dynamic tool-calling environment instead of a single forward pass - and the field promptly froze it into a static fixture. Agent red teaming inherited a leaderboard habit. AutoDojo's June 2026 paper is the obituary for that practice: benchmark distributions of attacks are fixed, so defenses tuned against them measure overfitting, not robustness.The adaptive loop, a cheap black-box optimizer using a frontier model to iterate injections against the deployed defense, raised attack success against nearly every defense tested across three task suites and five models.
The sharpest cells read like punchlines. PIGuard, a DeBERTa injection classifier, held static success at 0.0 percent and fell to 28 percent overall - 64 percent on action-open tasks where the user's request delegates the action itself to attacker-controlled content. DataFilter went from 12.6 to 33.4 percent, nearly triple. ProtectAI doubled from 7.2 to 15.4.
The authors' phrasing is the takeaway: static-string robustness overstates practical security, and on action-open tasks the injection can pose as ordinary data, which is a structural limit of any defense keyed to instruction-like text.
Two public results made the same point from the field. Nasr et al.'s The Attacker Moves Second broke twelve published defenses at over 90 percent success. And a vendor's own AgentDojo writeup in August 2026 reported a per-message guard taking 46.5 percent down to 0.7 percent while trace-level and containment defenses sat at 47.2 and 43.1 percent against a 46.5 percent baseline - statistical no-ops.
Then came the discovery: one arm was wired after the tool-execution loop had finished, observing attacks that had already succeeded. Forty-six offline tests passed green; none constructed a real pipeline. Placement was the difference between 46.5 percent and 0.7 percent - an agent red teaming lesson no scorer can score.
System-level evaluation, finally
RIFT-Bench's June 2026 design shows what the next generation measures. Instead of shipping another fixed attack list, it extracts each target's structure into a hierarchical code-grounded representation, then instantiates 105 reusable probes against that structure - 45 agentic systems, more than 10,000 distinct attack tests across five attack surfaces, including a multi-surface backdoor injected through one surface and activated through another.The probe adapts to the system it is pointed at, which is the whole point: attack suites that only work on the implementation they were written against are demonstrations, not evaluations.
The evaluation substrate matters as much as the attacks. Execution traces, not model outputs: tool invocations, resource consumption, environment state. A multi-agent workflow that routes a poisoned memory into a tool call scores on the trace, and the CSA guide's catalog - memory poisoning, tool misuse, role escalation, goal manipulation, multi-agent attacks - only becomes testable once the harness can see state transitions.
That is also where memory poisoning testing crosses from prompt exercise into systems testing: persistence is a state change, and state changes live in telemetry, not transcripts.
White-box limits are real - RIFT-Bench assumes code access and leans on an observability stack - but the direction is fixed. Output-only scoring is where the field started, and every 2026 framework that matters moved toward traces. System-level agent red teaming has one direction of travel. The question for any program buying or building an evaluator in 2026 is which side of that line its numbers come from.
Running the cadence
Continuous agent red teaming, not ceremonial. PyRIT's CI/CD integration exists precisely because model updates, prompt changes, guardrail tuning, and application releases are all events that invalidate yesterday's score - run the scenario suite on each, diff against the previous run, and treat movement in either direction as a finding. The scenario layer's tagged run IDs make this the default instead of an aspiration.The scorecard itself needs three bands to stay honest: model-behavior metrics (refusal, compliance, exploit success on text), system metrics (tool-call violations, state changes, egress attempts from ingested content, blast radius per identity), and utility metrics (task completion under attack, because a defense that bricks the agent traded a security incident for a business one).
Report the bands separately. The 2026 embarrassments - the flat-lined defenses, the 0 percent filters, the scores that could not see the firewall - all came from collapsing the bands into one number that meant less than it looked.
| Band | Primary signal | Failure it exposes |
|---|---|---|
| Model behavior | refusal / exploit success on text | guardrail regressions |
| System behavior | tool calls, state changes, egress | the say/do gap |
| Utility under attack | task completion, latency | defenses that disable the product |
| Adaptive re-run | same suite, optimized attacks | static-bench overfitting |
The stance that holds
Agent red teaming in 2026 earned its seat the hard way: three-hour campaigns that used to take weeks, scenario libraries that outlive notebook rot, adaptive attackers exposing every defense that scored 0 percent on a static string, and a published capability matrix admitting that the flagship toolkit sees 40 percent of what matters.The tools are better than they have ever been, and the failures that made the year's headlines were measurement failures - scoring sentences while agents moved state, wiring detectors after the loop, quoting static numbers against adaptive adversaries.
What holds is the say/do split as doctrine: score what the model says, separately score what the agent did, separately score what it costs the product, and never merge the columns. Add the adaptive re-run as the fourth band, because a defense that only survives the benchmark it was tuned on has not been tested - it has been rehearsed. And keep the human on strategy: automation bought the throughput, and judgment is the part the CSA report refused to automate away.
Run the unchanged scenario against your production stack this week. Whatever the diff says, you will finally be looking at the agent instead of its sentences.