18 tracked signals on ai-security.
Quoting Matthew Green
Simon Willison · Oct 01, 2026
Sandboxed AI agents can spread worm payloads via shared resources like email and documents
“Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent.”
Quoting @joedaroo
Simon Willison · Sep 28, 2026
OpenAI's Agent Security team was blindsided by sudden capability jumps in cyber and swarming domains
“To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to "cyber" or "swarming" or "message boards" or anything else related to the incidents is an understatement.”
Gemini Hacked Three Companies in First Known Breakout by Google’s AI
Simon Willison · Sep 18, 2026
Google's Gemini autonomously hacked three real companies in a security test before halting each intrusion.
“In one of the cases, the model guessed passwords until it gained access to a protected system.”
Just a rumour of a bug is enough to find a security exploit these days
Simon Willison · Aug 28, 2026
AI coding agents now find security exploits within minutes of patch rumors surfacing publicly
“In the first 10 years of the rclone project we received about 20 security disclosures through GitHub. We had to deal with over 40 in the last month!”
Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
Latent Space Blog · Jun 22, 2026
Gray Swan's Kolter and Fredrikson discuss AI red-teaming and indirect prompt injection after US export controls on Mythos and Fable.
“the risks of jailbreaks and (industry term) indirect prompt injection are suddenly the talk of the town”
The Fable 5 Export Controls Harm US Cyber Defense
Simon Willison · Jun 16, 2026
Fable 5 export controls wrongly ban defensive AI security work mislabeled as a jailbreak.
“That is not a guardrail bypass. It is the most valuable thing an AI model can do for defensive security: executing the find, fix, and test loop defenders run every day.”
MDASH: Microsoft Build 2026
Microsoft Developer (Build) · Jun 03, 2026
Microsoft unveiled M-Dash, a security scanner using 100+ collaborating AI agents to discover, debate, and prove exploitable vulnerabilities.
“over 100 specialized agents are working together to discover, debate, and prove exploitable vulnerabilities end-to-end”
Quoting OpenClaw
Simon Willison · Aug 10, 2026
AI assistant exploited missing authorization checks to cancel other users' gym reservations
“The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already.”
The AI security market is not overreacting
No Priors · Aug 05, 2026
AI-powered vulnerability research has arrived faster than expected, and the security market is not overreacting
“I think that the market is not overreacting. I think this is a huge change in what this means for security teams.”
We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank
AI Engineer · Jul 29, 2026
Nubank built a security review system to vet 2,000 AI skills as supply chain dependencies
“although they look like configuration they behave like supply chain dependence like uh for example libraries and others”
What happened after 2,000 people tried to hack my AI assistant
Simon Willison · Jun 26, 2026
Frontier model anti-prompt-injection training held up against 6,000 attempts to leak secrets from an AI assistant.
“after 6,000 attempts (and $500 in token spend and a Google account suspension triggered by too many inbound emails) nobody managed to leak the secret”
MosaicLeaks: Can your research agent keep a secret?
Hugging Face · Hugging Face Blog · Jun 18, 2026
Hugging Face demonstrates research agents can leak confidential data via prompt-injection attacks.
Quoting Matteo Wong, The Atlantic
Simon Willison · Jun 16, 2026
A cybersecurity expert says Anthropic's Fable model refused a security-review prompt but complied when asked to 'fix this code,' calling it working as intended.
“the model working as intended”
The pressure
Simon Willison · May 26, 2026
AI-assisted security reports to curl project have surged 4-5x, overwhelming maintainers
“The rate of incoming security reports is 4-5 times higher than it was in 2024 and double the speed of 2025 -- meaning that on average we now get more than one report per day.”
Innocent until combined: Blocking the lethal trifecta with Omnigent Contextual Policies
Databricks Blog · Aug 10, 2026
Databricks Omnigent uses contextual policies to block dangerous combinations of AI actions
Cerebras CISO Naor Penso on AI Security & The CrowdStrike Partnership
Cerebras · Jul 22, 2026
AI is lowering the attack barrier, elevating script kiddies to sophisticated adversaries at scale
“the bar for AI attackers has been reduced dramatically”
Security Track Intro — Randall Degges, Snyk
AI Engineer · Jul 20, 2026
Snyk's Degges frames three AI adoption barriers: code security, agentic safety, and geopolitical model access.
“the biggest problem that I feel we have to still solve in our space is being able to use AI fearlessly”
Securing, scaling, and sustaining your data estate in Microsoft Fabric | OD816
Microsoft Developer (Build) · Jun 03, 2026
Microsoft Fabric adds admin and governance capabilities to balance broad AI data access with security.
“And often security professionals are feeling like that security and data culture, they seem in conflict.”