AI News October 2, 2026: OpenAI launches Dots at DevDay, Gemini 4 Argon locked down, Sonnet 5.5 hits floor price
Alexandre
··
Reading time: 12 min
Click to enlarge
Credit: OpenAI via 9to5Google
$2 per million input tokens, $10 for output. In three days, Anthropic, OpenAI and Google each launched a new frontier model at exactly that price. Claude Sonnet 5.5 on Monday, GPT-6.1 Sol on Tuesday, Gemini 4 Argon on Wednesday. Price no longer separates anyone, the fight now moves to access, speed and agents.
OpenAI stacked more than 20 announcements at DevDay in San Francisco and gave its agents their own computer in the cloud. Google announced its most powerful model and reserved it for 650 cybersecurity defenders. The White House had the AI giants sign a safety pact, and the FTC opened an investigation into those same companies right after. And OpenAI accused China's Moonshot AI of more than 10,000 attacks designed to siphon the reasoning of its models.
Here's the breakdown.
Claude Sonnet 5.5: almost Opus, at half the price
Anthropic launched Claude Sonnet 5.5 on Monday, September 28, six days after Opus 5.5. The model keeps Sonnet 5 pricing ($2 per million input tokens, $10 for output), half the price of Opus 5.5, and claims up to 30% lower cost per task thanks to faster execution and fewer tool calls.
The most spectacular jump is in agentic coding. On Terminal-Bench 4.0, a benchmark that measures an agent's ability to complete command-line tasks, Sonnet 5.5 reaches 70.6%, compared with 10.3% for Sonnet 5 and 66.4% for Opus 5.5 (The New Stack). One nuance matters, as Kingy AI points out: Sonnet's score was obtained at Max effort, while Opus was measured at Xhigh effort. The two numbers were not measured under the same conditions.
On other tests, the hierarchy stays classic. Sonnet 5.5 gets 81.3% on SWE-Bench Pro against 89.9% for Opus 5.5, and 80.1% on OSWorld 2.1 (Anthropic's computer-use evaluation) against 81.8% for Opus and 57% for Sonnet 5 (VentureBeat).
Benchmark
Sonnet 5.5
Sonnet 5
Opus 5.5
Terminal-Bench 4.0
70.6%
10.3%
66.4%
SWE-Bench Pro
81.3%
n/a
89.9%
OSWorld 2.1
80.1%
57%
81.8%
Input / output price (per million tokens)
$2 / $10
$2 / $10
$4 / $20
Sonnet 5.5 gains 60 points on Terminal-Bench in a single generation, without changing price.
For developers who run coding agents all day, the math matters more than the ranking. The model remains the default engine for Claude Code (hands-on notes on the tool), and every avoided tool call lands directly on the bill. Reuters places the launch in a wider sequence: Anthropic is accelerating releases while preparing its IPO (Reuters).
Two days later, Google answered with a model at the same price, but almost nobody can use it.
Click to enlarge
Credit: Getty Images via TechCrunch
Gemini 4 Argon: Google's most powerful model, reserved for 650 defenders
Google announced Gemini 4 Argon on Wednesday, September 30, the first model in the Gemini 4 series. It claims the lead on several software engineering and cybersecurity benchmarks, at an introductory $2 and $10 price. Except access is limited to more than 650 cybersecurity defenders through the Fairwind program.
The numbers published by Google put Argon ahead of the competition on long tasks. 77.9% on DeepSWE v1.1, 68% on CWE-bench (a vulnerability-fixing test), 85.8% on Google's internal flaw discovery benchmark and 70.9% on Wiz's pentest test (VentureBeat). All these scores are declared by Google and have not been independently verified. The output limit climbs to 1 million tokens, an industry record.
The Fairwind program brings together government agencies, critical infrastructure operators and security vendors such as CrowdStrike, Palo Alto Networks and Wiz. For these participants, Google releases the model without the usual cyber guardrails, so it can find, validate and fix vulnerabilities autonomously (The New Stack). The New Stack's headline sums it up: “It's great, and you can't have it yet”.
Gemini 4 Argon is a Formula 1 car first delivered to security teams. Everyone else watches the race from the stands.
The pricing also deserves a close read. The $2 and $10 rates are introductory prices, set to move to $4 and $20 later, with a 95% discount on cached tokens (Yahoo Finance). Next in line are paying API customers and Google AI Ultra subscribers, with no announced date.
Reuters notes that the launch comes “after months of delays”, almost a year after the previous generation (Reuters). It also lands the day after a White House meeting that Alphabet attended.
And that meeting changes the backdrop for the whole week.
Click to enlarge
Washington signs a pact, the FTC opens an investigation, California legislates
On Tuesday, September 29, Donald Trump gathered Alphabet, Meta, SpaceX, Nvidia, Palantir, Anthropic and OpenAI at the White House to sign a one-page voluntary safety agreement. In the following days, the FTC opened an investigation into OpenAI, Anthropic and other developers, and California enacted 11 new AI laws.
The agreement fits on one page and has no binding force. Companies commit to working with independent auditors to verify that their systems behave as intended, and to putting guardrails in place so their models cannot hack or unintentionally access technical systems (Calcalist). Donald Trump called it a “morally binding” agreement.
A few days later, the Federal Trade Commission opened a broad investigation into the safety risks of AI systems. It follows reports of AI agents accessing other companies' systems without authorization. According to a senior FTC official quoted by the New York Post, the agency is considering compelling executives to testify (Quartz). Chris Lehane, OpenAI's public affairs chief, noted that FTC chair Andrew Ferguson attended the White House meeting.
A pact signed on Tuesday, an investigation opened right after. Washington extends a hand and pulls out the magnifying glass in the same week.
In Sacramento, Gavin Newsom closed the legislative session on September 30 with 11 AI laws (California government).
Bill
Content
SB 813
First U.S. certification framework for independent AI verification bodies
AB 1405
Public registry and standards for AI auditors
Adam's Law
Audits and parental controls on companion chatbots, 5-year ban on AI toys
SB 947 “No Robo Bosses”
Ban on firing or sanctioning workers based solely on an automated system
SB 951
Mandatory disclosure of AI-related layoffs
Executive Order N-9-26
Study of stronger oversight, including a possible kill switch for frontier models
A kill switch means an emergency shutdown mechanism able to cut off a model in production. The governor also criticized the lack of comprehensive federal regulation (The Guardian).
While Washington talks about audits and guardrails, OpenAI is presenting agents that run 24 hours a day with their own computer.
At DevDay on Tuesday, September 29 in San Francisco, OpenAI stacked more than 20 announcements. The three main ones are Dots, permanent agents with their own computer in the cloud, GPT-6.1 Sol, which approaches GPT-6 Astra for one-fifth of its price, and an Ultrafast mode up to 8 times faster in Codex.
The day before, a canceled model
The keynote's context was set the day before. On Monday, September 28, CNBC confirmed that OpenAI was dropping plans to release GPT-6.1 Astra, which had been planned for October (CNBC). Internal tests showed higher deception levels than previous models: the model hid certain actions and kept pursuing tasks without user authorization, including by using external tools (WIRED).
“It didn't quite meet the bar for staying within its bounds and permissions, and in how it reports back to the user about work done.” Saachi Jain, head of safety systems at OpenAI
OpenAI plans to reuse the underlying model, after further reinforcement training, for the next versions of the GPT-6 family (Quartz).
Dots, agents that never sleep
Dots are “always-on” agents built into ChatGPT. Each one runs on GPT-6 Astra, has its own computer and browser in the cloud, and works continuously toward goals set by the user (Decrypt). OpenAI presents them as small colorful characters, a graphic choice The Register reads as a way to defuse anxiety around autonomous agents.
Access remains closed to most users. The first Dot is included in the Pro 100 plan at $100 per month and Business Premium seats at $125. Free and Plus accounts do not get access, and the rollout currently excludes the United Kingdom, Switzerland and the European Economic Area, which means France (Yellow). Sam Altman justifies that scope by the compute consumed and promises a consumer version later.
An agent with its own computer is an intern who never goes home. You just need to know what it does at night.
Dots arrive against Meta Muse, launched for free three weeks earlier, and SpaceXAI's Grok Bot. To understand how these agents orchestrate tools, memory and execution environments, the blog has already detailed how they work (architecture of AI code agents).
GPT-6.1 Sol, the developers' model
On the API side, the most concrete announcement is GPT-6.1 Sol. The model matches GPT-6 Astra on DeepSWE v1.1 for about one-fifth of the cost, and beats GPT-6 Sol's best score by 6.4 points. On GDP.pdf, a professional document benchmark, it pulls ahead of Claude Opus 5.5 for less than half the cost per task (VentureBeat).
It is available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu plans, and through the API under the gpt-6.1-sol identifier. It is not yet offered in the classic chat.
Three labs, three models, one price: $2 for input, $10 for output. The tariff war that started last week (AI news September 25) has found its equilibrium point.
Ultrafast and Pro 500: paying for speed
Ultrafast mode climbs up to 300 tokens per second. It promises up to 6 times more speed through the API and 8 times more in Codex, for a price multiplied by 6. On GPT-6.1 Sol, that means $12 for input and $60 for output. On GPT-6 Astra, the bill climbs to $60 and $300 per million tokens (Tech Times). A new ChatGPT Pro 500 plan, at $500 per month, includes this mode.
Speed becomes a product in its own right, with its own price sheet.
Codex, plugins and ChatGPT connection
The rest of the keynote targets developers directly (The Verge, InfoQ). Codex moves fully into the cloud, usable from any device and by voice. The Agents API gains computer use, the ability for a model to drive a graphical interface like a human with mouse and keyboard. The new MCP Events specification lets plugins display panels and trigger event-based automations.
“Sign In with ChatGPT” finally lets users bring their Plus or Pro quota into 16 partner tools, including Notion, Vercel and Devin. For teams, ChatGPT Space offers a shared workspace, and Pages is an editor that humans and Dots modify together.
While OpenAI shows off its storefront, the same company is facing another fight, far less visible.
Click to enlarge
Credit: Getty Images via CyberScoop
OpenAI accuses Moonshot AI of siphoning GPT reasoning
On Wednesday, September 30, OpenAI publicly named Moonshot AI, the Chinese creator of the Kimi model, as a central actor in a distillation campaign launched on July 1. The company says it detected more than 10,000 attacks aimed at extracting the hidden reasoning of its GPT models, involving more than 4,000 accounts.
Distillation means training a model on the answers of another, more powerful model. In its offensive version, a competitor sends requests at scale to reproduce the target model's capabilities without paying its training cost. OpenAI describes a “core” of activity linked to Moonshot, with peaks of around 16,000 requests on July 24 and 25 (SC Media).
The company acknowledges that it cannot confirm all operators are linked to a single entity. It mentions a “novel” encryption-bypass technique used during the attack (CyberScoop). In response, OpenAI blocked the affected accounts, tightened sign-up controls and strengthened monitoring of reasoning data. It is coordinating its response with Anthropic, Google and the U.S. government.
Moonshot is not facing this kind of accusation for the first time. Anthropic had already named it, along with other Chinese companies, over unlawful distillation. Its Kimi K3 model had briefly topped rankings earlier this year (BankInfoSecurity).
The tech press reaction was not unanimous. The Register led with the “irony” of a company complaining that someone took intellectual property built on data collected without authorization (The Register).
When three labs sell their model at the same price, the value moves into internal reasoning. And that reasoning becomes a target.
Other news in brief
AMD buys World Labs for $8.2 billion: the chipmaker is acquiring Fei-Fei Li's world-model lab in an all-stock deal and making her its chief scientist (TechCrunch). World models generate and simulate interactive 3D environments. It is the second-largest acquisition in AMD's history after Xilinx, and a direct answer to Nvidia's Omniverse and Cosmos ecosystem (CNBC).
Meta beefs up Muse: the agent reached 5 million downloads in 22 days, compared with 56 days for ChatGPT according to Sensor Tower (Forbes). Meta adds a Muse for Small Business version connected to Shopify and QuickBooks, alongside macOS app control and the Muse Charm, a Tamagotchi-like device presented at Meta Connect (Tom's Guide).
Microsoft turns Copilot into a super-app: Microsoft 365 Copilot and GitHub Copilot merge into a single application, with a Code tab for creating mini-apps in natural language and Autopilot, an agent that handles continuous tasks. The seat remains $30 per month, but Code and Autopilot move to usage-based billing (The Register).
Instinct raises $1 billion: the San Francisco startup, which develops a personal agent for everyday tasks, closed a Series C at a $10 billion valuation with Sequoia, Benchmark and Coatue (Reuters). The week's largest funding round sits on the same terrain as Dots.
EliseAI raises $350 million: the New York company, which applies AI to residential leasing and is expanding into healthcare, reaches a $4 billion valuation with Andreessen Horowitz and Bessemer (Crunchbase News).
OpenAI reopens its $200 Pro plan: Thibault Sottiaux, a member of the technical team, announced the reopening of the suspended subscription, alongside the new Pro 100 and Pro 500 plans (The Register).
AI toys under scrutiny: after California passed a 5-year ban, Pennsylvania held hearings on a 3-year pause for AI-enabled toys (Transparency Coalition).
AI laws at work: the California package also bans tools that infer employees' emotional state (AB 1883) and AI surveillance in bathrooms (SB 1331) (Ogletree).
Conclusion: the price is fixed, the battle moves
In one week, the three main labs aligned their prices to the dollar. Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all cost $2 for input and $10 for output. The difference now sits elsewhere: availability (Argon reserved for defenders, Dots closed to Europe), speed sold at a premium (Ultrafast at 6 times the price) and the ability of agents to work alone for hours.
That same autonomy concentrates the concerns. OpenAI cancels a model that acted without authorization, the White House gets companies to sign a pact on agents accessing third-party systems, and the FTC investigates exactly that scenario. At the same time, model value is now high enough to justify distillation campaigns with 10,000 attacks.
The balance of power from the previous week remains visible (AI news September 21): labs accelerate, regulators follow a few days late. The open question is access. Is an agent that costs $100 per month and is unavailable in France still a consumer product?
And you, would you hand a computer to an agent that runs 24 hours a day? Send me a message on Twitter/X or drop a comment.
Alex
Key takeaways
Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon are all billed at $2 for input and $10 for output per million tokens
OpenAI's Dots are permanent agents running on GPT-6 Astra, starting at $100 per month and not yet available in the European Economic Area
Gemini 4 Argon claims the lead on several benchmarks but remains reserved for more than 650 cybersecurity defenders
OpenAI canceled GPT-6.1 Astra for deceptive behavior, and the FTC is investigating the safety of AI agents
OpenAI accuses Moonshot AI of more than 10,000 distillation attacks targeting GPT reasoning