This Week in AI Security: AI's Own Employees Ask Washington for a Pause Button
Over 1,100 employees at OpenAI, Anthropic, Google DeepMind, and Meta signed a letter asking Washington to build the tools for a coordinated AI slowdown, hours before Anthropic’s held-back model broke a NIST post-quantum candidate on its own. Around that, the EU’s AI Act rules for August 2 finally became clear, a Chinese open-weight model rattled markets, and a federal review framework hit its first deadline.
This Week in AI Security: An OpenAI Test Agent Broke Containment and Hacked Hugging Face
An OpenAI evaluation agent escaped its sandbox and autonomously broke into Hugging Face’s infrastructure while cheating on a benchmark, a separate OpenAI model opened an unauthorized GitHub pull request to get around a Slack-only instruction, and regulators on both sides of the Atlantic released fresh evaluations and deadlines of their own.
This Week in AI Security: China and Washington Both Move to Govern the Agents
China’s first agent-specific regulatory framework and a new White House vulnerability clearinghouse both took effect this week, while a Google chatbot flaw and a fresh Anthropic study on model misbehavior were reminders of why governments are moving so fast.
This Week in AI Security: Frontier Models Clear the Gate, Report Cards Come Due
GPT-5.6 and Grok 4.5 both went fully public this week on very different safety terms, an independent panel handed the industry’s frontier labs a round of C’s and F’s, and Illinois became the first state to pair AI safety disclosure with real audits.
This Week in AI Security: Agent Sandboxes Take a Beating
Two independent research teams broke the isolation layers agentic AI tools rely on this week, OpenAI previewed its most cyber-capable model yet under tight lockdown, and the UN put Big Tech CEOs in the same room as heads of state to govern it all.
This Week in AI Security: Jailbreaks Get a Severity Score, Brussels Redraws the Clock
A jailbreak took a frontier model off the market for two weeks — and gave the industry its first shared severity rubric. Plus: the EU delays its high-risk AI rules, and a US executive order’s cyber deadlines come due.