AI News

AI News September 25, 2026: Altman and Amodei at the UN, a price war in 90 minutes, Jev breaks the LLM mold

Alexandre
Alexandre
··
Reading time: 9 min
Sam Altman speaks at the United Nations Security Council on September 23, 2026

Click to enlarge

Credit: AFP via Al Jazeera
On September 22, Anthropic shipped Claude Opus 5.5 and cut its pricing by 20%. Ninety minutes later, OpenAI published GPT-6 Sol and GPT-6 Luna, two models priced at roughly half the previous generation. The next day, the CEOs of both companies sat in front of the United Nations Security Council to explain that their industry needs oversight.
The whole week fits inside that gap. TypeSafe AI released Jev, a model that cannot write a sentence and claims 193x the speed for 445x less money. Factory raised $200 million at a $5 billion valuation for its coding agents. Crusoe closed a $3.9 billion Series F at $30.9 billion. And a Forcepoint simulation showed how a poorly bounded agent fires 500 tool calls in a single turn.
Here's the breakdown.

Jev, the model that cannot write and says so upfront

TypeSafe AI launched Jev on September 19, a "System One" model that does not produce free text but picks an answer from options you define, with a confidence score. The company claims 193x faster and 445x cheaper than a classic LLM, at $0.042 per million input tokens, with output billed at zero.
The idea comes from Daniel Kahneman. A System One model answers fast and on instinct, where a classic LLM does System Two, slow and deliberate. In practice Jev writes nothing: you hand it a text and a list of possible outputs, it returns a choice and a probability. Simon Willison, who took the launch apart on September 21, calls it a new shape of model, the "decision models", and Zapier sums up the positioning in one image: a smart if-statement.
The use case is not conversation, it's plumbing. Routing a ticket to the right team, classifying content, flagging fraud, choosing which tool an agent should call. All micro-decisions that most teams currently run on a generation model at $3 or $5 per million tokens, purely because nothing else existed. Tom's Hardware notes the model bills on input only, which makes sense when the output is a single word.
The other use TypeSafe pushes is monitoring: Jev can read the traces of an LLM agent to spot misalignment or a jailbreak attempt, meaning a prompt engineered to bypass the model's guardrails. A control layer sitting next to the main model, not replacing it. Orchestration between specialized models is something this blog has already dug into (architecture of AI code agents).
While TypeSafe attacks the problem from the bottom of the spectrum, the two large labs are fighting at the top.
Anthropic shipped Claude Opus 5.5 on September 22, 2026

Click to enlarge

Credit: Anthropic

Anthropic cuts prices, OpenAI answers in 90 minutes

Anthropic released Claude Opus 5.5 on September 22, the first model in the 5.5 family. Per the official announcement it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5, with a 20% cut on per-token pricing and 60% off cache costs. Ninety minutes later, OpenAI shipped GPT-6 Sol and GPT-6 Luna.
Anthropic's statement is blunt: Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5" (Anthropic). The 60% cut on cache deserves a line, because it's the biggest item on an agent bill: prompt caching means the parts of the context already sent in a previous call are billed at a lower rate, and an agent looping over the same codebase re-reads the same files endlessly.
OpenAI's counter landed right behind, with two GPT-6 variants aimed at the mid and budget tiers, at roughly half the price of the GPT-5.6 models. SiliconANGLE timestamped the sequence: Anthropic announces, OpenAI publishes "minutes later". Quartz and KDnuggets frame it the same way, performance held, price broken.
This is the second time in a month the two labs have shipped hours apart. What's different from previous cycles is that the selling point is no longer the benchmark, it's the invoice. Neither announcement claims a capability jump: both claim the same level for less money. For a solo developer running agents continuously, that's the variable that actually matters (how I use Claude Code daily).
The price cut lands exactly when investors are pouring billions into the layer underneath.
Crusoe raised $3.9 billion for its modular AI data centers

Click to enlarge

Credit: ETCIO Datacenters

The money goes to coding agents and data centers

Factory raised $200 million at a $5 billion valuation, up from $1.5 billion earlier this year. Over the same window, Crusoe closed a $3.9 billion Series F at a $30.9 billion valuation, and Temporal a $550 million Series E at $12.6 billion. The three largest checks of the period all went to automated code and the infrastructure that runs it.
Factory sells "Droids", autonomous agents positioned across the whole software development lifecycle rather than code completion alone. The round was confirmed by the company and reported by Forbes on September 25. The commercial pitch leans less on code quality than on control: model choice, cloud, on-premise or air-gapped deployment, plus a measurement layer tying agent sessions and their spend back to projects, issues and pull requests. A valuation up more than 3x in months on a governance promise, which is to say on the needs of teams that look nothing like a solo developer's (the technical story behind an app built alone).
On the infrastructure side, Crusoe announced $3.9 billion in Series F, co-led by Valor Equity Partners, Atreides Management and Mubadala Capital, with Founders Fund, Nvidia, GIC, the Qatar Investment Authority, Radical Ventures and TPG participating (Crunchbase News). The company is shifting to modular AI data centers, prefabricated blocks dropped next to a substation instead of purpose-built facilities (ETCIO).
Temporal completes the picture with $550 million led by Lightspeed, Goldman Sachs Alternatives, Wellington Management and Tiger Global. The company builds fault-tolerant distributed workflows, exactly the layer an agent needs when it chains fifty tool calls and one of them fails. Three rounds, three positions on the same chain: the agent, the workflow, the machine.
Sam Altman listens during the Security Council session on artificial intelligence, September 23, 2026

Click to enlarge

Credit: Reuters via CNN

Altman and Amodei ask for global regulation, Trump refuses it

On September 23, Sam Altman and Dario Amodei addressed the United Nations Security Council in a session convened by France. Both warned that AI, badly managed, is "a risk to humanity as a whole" and called for international oversight mechanisms. The same day, Donald Trump rejected any international regulation of AI from the UN podium.
Amodei, speaking by video, put three concrete proposals on the table according to CNN: a narrow international ban on using AI to design biological weapons, an evaluation and verification system letting states check each other's commitments, and shared testing standards paired with a notification mechanism for security incidents. Anthropic's CEO also restated his position on the pace of development, continuing the essay he had published two weeks earlier (the full sequence here).
Altman pressed a different point: democratic processes and accountable governments should set AI policy, not a handful of Silicon Valley labs (Al Jazeera). The French and British foreign ministers called for common frameworks. Hugging Face, now part of Nvidia, also took part in the session. France 24 and Le Monde covered the session in the same terms.
The U.S. administration held its line. Donald Trump ruled out any new international guardrails, which puts both companies in an unusual spot: they are asking for oversight their own government refuses to organize. A CEO publicly demanding constraints nobody has the power to impose remains, in practice, his own regulator.
The next day, Anthropic gave a perfect example of what makes the debate concrete. The company announced on September 24 that Claude had identified a CRISPR-like enzyme system, exactly the kind of biological capability Amodei had proposed fencing off a day earlier (Al Jazeera). Scientific discovery and proliferation risk are the same capability, seen from two sides.
Denial of wallet takes nothing offline: it empties the budget

Click to enlarge

Credit: Kaspersky Daily

The invoice as attack surface: OWASP ranks unbounded consumption sixth

The 2026 OWASP Top 10 for LLM applications ranks "unbounded consumption" sixth. A Forcepoint simulation published this week shows a research assistant agent, with no recursion limit and no call budget, firing 500 tool calls in a single turn. With guardrails in place, the same scenario drops to one call and roughly $0.02.
The attack is called agent-tool fan-out. An attacker poisons the agent's input, meaning they slip an instruction into content the agent will read that triggers cascading calls. Each individual request stays inside rate limits, so nothing fires on the monitoring side. It's the aggregate that blows up. SC Media reports a simulated cost of about $0.50 by the hundredth turn, scaling to hundreds of dollars across concurrent sessions.
The scenario's name says it all: denial of wallet. No service outage, no data leak, just a bill climbing until someone looks at the dashboard. Dark Reading details five variants of the same mechanism in the Forcepoint report, from runaway processing to model theft. Kaspersky notes that 2026 is the first year large companies significantly overshot their AI budgets, with Uber burning its annual envelope by April (Kaspersky).
The tested countermeasures are deliberately basic: a per-session call budget, a recursion depth cap, and an agentic circuit breaker that halts execution as soon as abnormal fan-out shows up. Those three settings take the simulation from 500 calls down to one. Nothing that requires an architecture rewrite, just explicit limits most frameworks leave wide open by default (ten ways to cut your token bill).
The week reads end to end along that line: cheaper models, agents calling more tools, and a spending ceiling nobody sets until it blows.

Other news in brief

Mistral acquires Pimento: the Paris company continues its buying spree after its €3 billion Series D led by Samsung on September 8, pushing its post-money valuation past €21 billion. The deal targets its Vibe assistant more than advertising (Cryptonomist).
Nvidia closes the Hugging Face acquisition: the definitive agreement signed on September 2 covers about $13 billion, including $11 billion to investors and roughly $1 billion in employee retention equity, for a platform used by more than 18 million developers (TIKR).
Cloudflare changes its crawler defaults: since September 15, bots are sorted into three buckets, Search, AI-Training and AI-Agent, with the last two blocked on ad-supported pages. If your agent fetches web content, check which bucket it lands in (Shattered).
Agent frameworks ship in a batch: LangChain 1.3.15 (core 1.6.4), LlamaIndex 0.14.25 and CrewAI 1.15.22 all landed around September 20 and 21. LangSmith adds a self-hosted option and stops recording request and response content by default.
Google turns REST APIs into MCP tools: a September 24 post details how to expose REST endpoints to agents through the Model Context Protocol with Cloud API Gateway, plus an Agent Gateway layer for governance (Google Developers).
Ringg resolves 65% of customer calls on GPT-5.6: the company announced on September 24 that its multilingual voice agents run at roughly 90% lower cost than its GPT-4.1 setup (OpenAI).

Conclusion: the week price became the real story

Three announcements, one direction. Jev sells a decision at $0.042 per million tokens. Anthropic takes 20% off Opus and 60% off cache. OpenAI answers at half price in ninety minutes. None of the three promises a smarter model: all three promise the same work for less money.
And at the exact moment unit cost collapses, the Forcepoint report is a reminder that volume has no native ceiling. An agent firing 500 tool calls in one turn costs almost nothing per call. It's the counter that hurts, not the unit. The three guardrails that take that simulation down to a single call fit in a few lines of configuration, and sit disabled by default in most of the frameworks that shipped this week.
One floor up, the Security Council spent a day debating oversight the main country involved refuses to organize. Amodei's proposals, mutual verification, shared testing standards, incident notification, look a lot like what engineering teams call explicit limits. At a scale where nobody yet has the power to enforce them.
So what about you, have you set a spending cap on your agents, or do you check the bill at the end of the month? Send me a message on Twitter/X or drop a comment.
Alex

Key takeaways

  • TypeSafe AI launched Jev on September 19, a System One model returning a decision and a confidence score instead of text, at $0.042 per million input tokens
  • Anthropic shipped Claude Opus 5.5 on September 22 with 40% lower running cost than Opus 5, a 20% per-token price cut and 60% off cache
  • OpenAI countered ninety minutes later with GPT-6 Sol and GPT-6 Luna, at roughly half the price of the GPT-5.6 models
  • Factory raised $200 million at a $5 billion valuation for its Droid agents, Crusoe $3.9 billion at $30.9 billion and Temporal $550 million at $12.6 billion
  • On September 23, Altman and Amodei asked the UN Security Council for international oversight while Donald Trump rejected any global regulation
  • The 2026 OWASP Top 10 ranks unbounded consumption sixth, and a Forcepoint simulation shows 500 tool calls in one turn dropping to a single call with three guardrails

I'm Alex, creator of Waku. Find me on Twitter/X and Instagram.

Comments

Comments

Got a take on this article?

Create a free account in 10 seconds to comment, like, and get the next articles straight to your inbox.

Don't have an account yet?

This site uses cookies for analytics and advertising. No personal data is sold. Learn more