GPT-6.1 Astra, scheduled to ship in ChatGPT and Codex in October, was scrapped after internal testing found it misled users about its actions and overstepped its authorization.
OpenAI unveiled more than 20 products at DevDay 2026, led by always-on Dots agents, Codex Cloud and Security Cloud, the GPT-6.1 Sol model, and a paid Ultrafast speed tier.
Anthropic's second Claude 5.5 model posts a 70.6% Terminal-Bench coding score at unchanged pricing, though independent testing shows the per-task savings depend on the effort setting.
Anthropic says Claude Fable 5.1 answered a public physics challenge by computing the nine-loop hexagon amplitude in N=4 super-Yang-Mills, working mostly unsupervised for days at an end-user cost of roughly $1,000 to $2,000.
Anthropic now bills Claude API requests refused by its safety classifiers before any output when tagged bio, frontier_llm, or reasoning_extraction, aiming to make large-scale safeguard probing expensive.
In an interview with China Media Group, Elon Musk conceded that Grok 4.7 is not as good as Anthropic's Opus 5.5, and said his AI company probably won't catch the frontier until next year.
OpenAI says it has halted training, evaluation, and inference involving tool use on its most advanced models after one slipped out of an offline test environment and reached an external chatbot.
MiniMax's official account confirmed M3.1-Flash-Preview went live inside MiniMax Code on September 27, a day after users spotted it in the product, with no benchmarks and no published price.
OpenAI launched GPT-6 Sol and GPT-6 Luna on September 22 at roughly half the price of GPT-5.6, within hours of Anthropic cutting Opus 5.5 pricing and cache reads, escalating the coding-model price war.
Claude Opus 5.5 cuts API prices 20 percent, cache reads 60 percent, and boundary-circumvention attempts 85 percent, launching the same night as OpenAI's cheaper GPT-6 Sol and Luna.
Meta's Connect developer sessions made the Model API generally available globally, shipped a 30B open-weights model that runs on a Mac Mini, took Muse Code out of beta, and opened a $1M hackathon.
Project HydraFusion, now in experimental preview for Copilot CLI, picks a mix of models for each coding job at runtime — cutting estimated costs by 67% while matching frontier model quality on benchmarks.
The world's dominant AI chip maker is acquiring the platform where 18 million developers find, share, and deploy open-source models. The deal closes in early 2027 — pending regulatory review.
DeepSeek's new V4.1 Flash model beats V4 Pro on benchmarks while running at Flash pricing. Starting September 14, V4 Pro API requests automatically redirect to it.
Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code, and Claude Opus 4.7 are all being cut from Copilot on October 2. The Kimi replacement costs more and is off by default for enterprise customers.
Meta's agentic coding model gets stronger benchmarks, 20% fewer tool calls, and a new Contributor tier at $0.10 per million input tokens — 21x cheaper than the standard price, in exchange for your prompts and completions being used for training.
OpenAI's most capable model yet hits 100% on ExploitBench, 98.6% on ARC-AGI-3, and operates software faster than any previous version. Greg Brockman says welcome to the AGI era.
Alibaba released a new checkpoint of Qwen3.8-Max on September 2 with targeted improvements to engineering-scale coding, multi-tool agent orchestration, and document understanding.
Anthropic's latest frontier model improves on Fable 5 for long-running agentic work, cuts cache read costs to a quarter of the old rate, and introduces three API-breaking changes that developers need to handle before migrating.
Gemini 3.1 Pro, Claude Opus 4.5, Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6, and Raptor Mini are all being deprecated across every Copilot surface. Here's what to switch to.
A mystery model called Ox Alpha has been running free on OpenRouter and OpenCode since August 20. It claims a 1M-token context window and 100 trillion tokens per day of capacity. The creator has not come forward, and fingerprinting points toward a company on the U.S. government's Entity List.
Z.ai's new open-weight model hit 84.5% on CyberGym and 54.4% on ExploitBench, and has already identified a serious security flaw in the Cursor code editor.
Two new models landed in GitHub Copilot this week: Kimi K3 from Moonshot AI at $3 per million tokens, and MAI-Code-1.1-Flash, Microsoft's own code model with native image understanding. Agent Plugins 1.0 also went generally available.
The new entry-level model outpaces Claude Sonnet 5 and GPT-5.6 Terra on several tests, ships at $0.75 per million input tokens through year-end, and powers Google's Gemini Spark agent.
ESC
Start typing to search across tools, news articles, and reviews
No results found
Optional analytics and social embeds help us improve the site. It works fully without them.
Privacy
Privacy Settings
Choose which optional features you want to enable.
Your browser sent a global privacy opt-out signal. We respect that by default.