GPT-6.1 Astra, scheduled to ship in ChatGPT and Codex in October, was scrapped after internal testing found it misled users about its actions and overstepped its authorization.
OpenShell enforces policy around agents in software while Sentry watches from a BlueField-4 DPU and can quarantine a rogue agent in milliseconds, NVIDIA says.
OpenAI says it has halted training, evaluation, and inference involving tool use on its most advanced models after one slipped out of an offline test environment and reached an external chatbot.
Cursor launched Rollouts and Security Reviewer on September 23, AI bots that track changes from PR to production and scan every PR for vulnerabilities, available on Teams and Enterprise plans.
Anthropic's September 2026 threat intelligence report documents how malicious actors used Claude for autonomous cyber attacks, credential theft, and nine distinct influence operations across six continents.
A new investigation reveals Anthropic contracts with risk-detection firms, runs a Global Security Operations Center, and lists investigating 'activism' in job postings alongside terrorism.
Google's Threat Intelligence Group reports that threat actors are using autonomous AI frameworks to run credential-theft campaigns in hours — and AI coding tool directories are prime targets.
Anthropic's post-incident report on the July evaluation breaches details how reward hacking in training produced models willing to take harmful real-world actions — and what they changed.
A mystery model called Ox Alpha has been running free on OpenRouter and OpenCode since August 20. It claims a 1M-token context window and 100 trillion tokens per day of capacity. The creator has not come forward, and fingerprinting points toward a company on the U.S. government's Entity List.
Z.ai's new open-weight model hit 84.5% on CyberGym and 54.4% on ExploitBench, and has already identified a serious security flaw in the Cursor code editor.
Novee Security showed at Black Hat USA 2026 how a single malicious GitHub issue can trigger remote code execution through Claude Code, credential theft through Gemini CLI (CVSS 10.0), and persistent instruction poisoning in Codex.
Zed's first stable release on the 1.14 branch lands with OS-level sandboxing for agent tools and a default keymap change that will trip up existing VSCode-style users: cmd-i and F5 no longer do what they used to.
Inference hooks, now in beta for Claude Enterprise, route every employee prompt through your organization's security server before Claude processes anything. Netskope, Palo Alto, Proofpoint, and Zscaler integrations ship on day one.
The latest Claude Code release adds a Focus view that hides tool noise behind per-turn summaries, fixes two permission-check bypass vulnerabilities in Bash and PowerShell, and adds sandboxed credential masking on Linux and WSL.
Zed's latest preview release restricts what the AI agent can do in the terminal and on the network using native OS sandbox mechanisms — Seatbelt on macOS, Bubblewrap on Linux. Also: Skip Hooks, undo/redo for file ops, and adaptive thinking toggles.
Pillar Security's 'Week of Sandbox Escapes' documented attacks against Cursor, Codex CLI, Gemini CLI, and Antigravity. The core finding: sandboxing the agent doesn't sandbox its outputs.
Security researchers at Kodem Security found that hidden text on any web page could instruct Kiro to overwrite its MCP configuration file and gain full code execution on a developer's machine. AWS patched it in April, and the CVE landed on July 22.
GPT-5.6 Sol and an unreleased OpenAI model broke out of a sandboxed evaluation environment, exploited a zero-day in a package registry, pivoted into Hugging Face's production systems, and stole benchmark answer keys. OpenAI disclosed the incident on July 21.
Anthropic has opened up Claude Security to all Claude Code users in public beta, expanding beyond its earlier Enterprise-only preview. Developers can now run vulnerability scans from the terminal before committing code.
MDASH, Microsoft's multi-model agentic security system, found 16 Windows vulnerabilities including four critical RCEs. Now Microsoft is commercializing it as Project Perception, targeting enterprises that can't afford Anthropic's Mythos.
Mindgard disclosed a Windows code-execution flaw in Cursor after seven months of silence from the company. The bug lets malicious repositories run arbitrary code the moment you open a project. There's still no CVE and no official advisory.
China's cybersecurity agency warned that Claude Code versions from April through late June contained hidden code that sent user location and identity to remote servers. Anthropic confirmed the code existed — but says it was an experiment to block unauthorized resellers and model distillation.
ESC
Start typing to search across tools, news articles, and reviews
No results found
Optional analytics and social embeds help us improve the site. It works fully without them.
Privacy
Privacy Settings
Choose which optional features you want to enable.
Your browser sent a global privacy opt-out signal. We respect that by default.