GPT-Red vs. your red team: the 84%-to-13% offense gap is now a blueprint
Self-play adversarial discovery just outran human red-teamers by 6.5x, and the same loop hardening OpenAI's own models is a published method commodity actors can approximate against yours.
The number that should reset your threat model isn't 84% — it's 6.5x, the margin by which GPT-Red's self-play adversarial loop outpaced human red-teamers. Self-play is the operative word: the system trains by attacking a model, watching its defenses, and inventing progressively stronger prompt injections with no human authoring the exploits. The offense improves autonomously, and the published result is effectively a blueprint any well-resourced adversary can approximate against your less-defended deployments. Artisanal manual red-teaming no longer covers the surface it used to.
Two independent signals converge. Technically, an automated attacker compromised GPT-5.1 in 84% of scenarios versus 13% for humans. Institutionally, banking executives and intelligence agencies now fear Anthropic's Mythos frontier model can enable cyber catastrophes — automated exploit generation and scaled social engineering. When quants and spy agencies raise the same flag independently, that's a shifted baseline threat model, not vendor FUD, and it should move offensive AI into a named attack-surface category on your detection roadmap.
The caveat matters: Mythos is institutional alarm, not a demonstrated attack chain, and no CVE, CVSS or TTP mapping exists yet. This is a detection-engineering priority, not an emergency patch. The concrete near-term consequence is the death of the 'poor grammar' phishing heuristic — commodity actors now generate native-quality lures and working exploit scaffolding at scale, eroding the human-layer defenses your awareness program was built on.
The defensive read is the same loop inverted: GPT-Red hardened OpenAI's own GPT-5.6 Sol without capability loss. Automated adversarial testing baked into the training pipeline is where robustness now originates — static content filters lag. That hands you a sharp new vendor-selection criterion: providers who run red-team-in-training should outrank those relying on post-hoc filtering, and you should ask for their detection rates directly. Treat the 84% figure as your assumed exposure rate until a deployment proves otherwise.
What to do
Run automated prompt-injection testing against every production LLM surface (chatbots, copilots, RAG, agents) this quarter, treating 84% as your assumed exposure rate.
Raise phishing-simulation difficulty to native-quality lures and re-test MFA/identity controls now, on the assumption social-engineering quality has jumped a tier.
Add automated adversarial-testing maturity as a scored criterion in LLM vendor assessments and require red-team-in-training detail plus detection rates at next renewal.