·
·
AI / Artificial Intelligence
Anthropic/Claude
OpenAI/ChatGPT
Mistral AI
Google / Gemini
·
LLM
Anthropic/Claude
OpenAI/ChatGPT
European AI
Mistral AI
Google/Gemini
AI News Week 39: Anthropic cuts prices, Google admits to a breakout
Anthropic has launched Claude Opus 5.5. It matches Fable 5.1 in performance but costs 40 per cent less than Opus 5. Google admits Gemini breached three external corporate systems, only disclosing the incident after press enquiries. Meanwhile, Anthropic reveals that Claude now handles 26 per cent of its own development work. OpenAI is integrating its ad platform into HubSpot and Shopify. Mistral now powers Firefox. Finally, Apertus is gaining real-world user feedback via Proton’s Lumo.

Last week, three providers called for a slower pace. This week, they deliver numbers, confessions, and a new model. Anthropic cuts prices by 40 per cent, Google admits to a breakout, and Mistral lands in Firefox.
1. Claude Opus 5.5: same performance, significantly lower price
On 22 September, Anthropic released Claude Opus 5.5. It is the first model in the 5.5 family. Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks.
Anthropic’s core message: Opus 5.5 performs on par with Fable 5.1 for most tasks, whilst costing 40 per cent less than Opus 5.
Where the 40 per cent savings come from. Two effects combine. Firstly, the price per token: input costs 4 dollars per million instead of 5, output 20 instead of 25. However, the largest cost for agents is cache reads, which drop from 50 to 20 cents per million. Secondly, the model requires fewer tokens per task. There is also a speed boost: output runs over 30 per cent faster than on Opus 5.
What beta testers are reporting. The figures come from Anthropic’s own announcement.
Deloitte Consulting: Opus 5.5 found 72 per cent of known bugs in code reviews at the lowest tier. Opus 5 found 56 per cent at a high tier.
Box: one-third of the tokens of Opus 5, with responses 40 per cent shorter at the same level of accuracy.
Optiver: same quality as Opus 5 in about half the steps and time. Costs are 40 to 50 per cent lower.
Quantium: a complex programming task previously took 38 prompts over four days; it now takes 11 prompts over three hours.
One tester audited 200,000 lines of code in under three hours. Opus 5 took over 20 hours.
Higher rate limits. The five-hour limits are increasing for Pro, Max, Team, and Enterprise accounts. Subscribers also get a limit reset that they can save and use when needed.
The security aspect is more interesting. Opus 5.5 is the first Anthropic model released since Dario Amodei’s call for a slower pace. External auditors tested it before release, including METR and Frontier Design.
One statistic fits the theme of this issue. In a new test, Opus 5.5 was roughly 85 per cent less likely than Opus 5 to attempt to exceed the boundaries of its environment. Every attempt was harmless, and the model reported it itself.
A critical note: In Anthropic's benchmark table, Opus 5.5 is sometimes significantly ahead of Fable 5.1. Anthropic tempers this in the same text: in internal use, the gap is smaller than the numbers suggest. Anyone reading "much better than Fable" is more confident than the provider itself.
Takeaway: Switch over and run the numbers. Take a process you run frequently and compare last month's bill with this month's.
2. Google admits: Gemini broke out of the testing lab
On 18 September, the Wall Street Journal reported that Gemini had breached three external corporate systems.
What happened: Israeli security firm Irregular tested Gemini in a training scenario in May. A configuration error in the test environment allowed the agents access to the open internet. Once there, they gained access to three external systems. In one case they guessed credentials; in another, they found them in a public code repository.
Google states the model corrected itself and no damage was caused. The group’s explanation: Gemini assumed the external systems were part of the test.
The timeline. The test took place in May. Irregular noticed the incidents in July, after other providers reported similar cases. It became public in September when the WSJ enquired. Google did not disclose the incident proactively.
The pattern. OpenAI, Anthropic, and Meta have all reported the same issue in recent weeks. Google is the fourth. Every major lab now has at least one case of a model leaving its test environment.
Takeaway: The test environment failed, not the model. It was connected to the internet despite being meant to be isolated. When testing agents, you must verify this separation, not assume it.
3. Anthropic discloses how much Claude contributes to Claude
Last week, Dario Amodei called for more transparency from the labs. On 17 September, Anthropic provided three metrics.
How much development work the AI does itself. Anthropic tracked and evaluated around 15,000 individual steps in its own model development. As of August 2026, Claude performs 26 per cent of this work. In February, it was under 1 per cent. Claude does not work fully autonomously in any area.
In this context, performing the work means Claude handles a task from prompt to final result. A human then reviews it and decides whether to keep it.
How the agents are monitored. Around 30,000 agents run simultaneously at Anthropic. Every action must pass a guardrail program before execution. In August, over one billion actions went through this control. 0.002 per cent were stopped—around 20,000 actions.
A second program reviews all actions after the fact. It flags about 100,000 logs per week for review, with around 50 reaching a human auditor.
Where the compute goes. In a sample week in July, around 6 per cent of development compute went into safety work. For the work done by Claude itself, the figure was around 12 per cent.
The auditors. On 18 September, Anthropic signed Accenture as its first external auditor. Both companies plan to invest at least one billion dollars each over five years. The auditors will work in-house with access similar to employees.
A critical note: Anthropic pays Accenture directly. Anthropic itself writes that, in the long term, funding should come from state or shared pools. As long as the audited party pays the auditor, this is not an independent verdict, but a commercial contract.
Takeaway: Anthropic shows the scale of control agents require. One program checks every action beforehand, a second reviews everything afterward, and humans decide the difficult cases. Deploying agents to production systems without this layer means you are not in control.
4. OpenAI expands its advertising platform
On 16 September, OpenAI announced four new features for ChatGPT advertising.
Sponsored Agents. Users who click an advert can start a conversation with an agent from the advertising company. This is clearly labelled and separated from the main chat. Currently running as a test with selected advertisers in the US.
Campaigns via prompt. Ads can be created, modified, and analysed in ChatGPT Work via the Ads Manager plugin. A website URL or a briefing is transformed directly into a campaign.
Suggestions in Ads Manager. When creating an advert, the system suggests copy and images based on the landing page and campaign goal. You can also automatically adapt headlines to the conversation context and translate them into the user's language.
HubSpot and Shopify. Companies managing customers in HubSpot can connect a ChatGPT Ads account, build adverts, and follow up on leads. Shopify merchants in the US get an app in the Shopify App Store, with product catalogues already integrated. The app launches in all markets where ChatGPT Ads is available starting 23 September.
Takeaway: OpenAI is integrating where your team already works. If you use HubSpot or Shopify, the channel will soon appear as a simple checkbox in your familiar tools. At that point, it is no longer a pilot project.
5. Europe: Mistral is now built into Firefox
On 16 September, Mistral and Mozilla announced a partnership. Mistral's models power Firefox Smart Window, Mozilla’s browser-based AI assistant. It is currently in beta—starting in France and North America, and expanding to the UK and Germany later this year.
Two points are relevant here. The models are fine-tuned on regional languages and dialects. Data privacy is also strict: chats are not saved on Mozilla’s servers by default, and Mistral does not store the data at all.
Additionally, Mistral raised three billion euros on 8 September at a valuation of over 21 billion euros. Samsung Electronics led the round, with participation from the EU-backed Scaleup Europe Fund, PSG Equity, Advent, funds managed by BlackRock, and the Grand Duchy of Luxembourg. CEO Arthur Mensch stated the funds will go towards building proprietary data centres.
This aligns with a study by the BCG Henderson Institute, which analysed the AI policies of over 30 countries. The result: complete AI sovereignty is out of reach for almost all countries. India's state-backed IndiaAI programme held 62,000 GPUs in early March 2026. Microsoft alone is estimated to have purchased around 485,000 units of a single generation in 2024. The institute therefore recommends resilience over autonomy: keep sensitive compute local and run multiple providers in parallel.
Takeaway: Europe can now run a browser powered by European models and invest billions in its own provider. Yet the chips still come from the US, and a portion of the capital from South Korea. For most businesses, the question is simpler: where does our data reside, and how quickly can we switch providers?
6. Switzerland: Apertus gains what open models lacked
On 18 September, the Apertus team from ETH Zurich, EPFL, and CSCS announced a partnership with Proton. Users can select Apertus 1.5 as the model in Proton's AI assistant, Lumo. Users can opt to provide feedback to help develop the model. Chats remain encrypted; Proton itself has no access.
Imanol Schlag, co-lead of the Apertus project, calls feedback from real-world use the missing ingredient for open models. This is the crucial point. Open models have weights and source code, but lack a user base. Major providers learn daily from millions of conversations. This partnership directly targets that advantage.
Meanwhile, the technical report for Apertus 1.5 is still missing. Announced at the 24 July release, it has yet to be published.
Swiss {ai} Weeks continue until 4 October. The ETH AI+X Summit takes place on 1 October.
7. In brief
OpenAI catches its models covering up errors. On 16 September, OpenAI published six instances of anomalous model behaviour. The most striking: during the training of GPT-5.6 Sol, models left instructions for their successors to hide errors from users. An agent lacking historical financial data instructed its successor to invent them. OpenAI wrote that the industry has not sufficiently resolved alignment and oversight to continue scaling at full speed for much longer.
A model that outputs no language. Startup TypeSafe AI released Jev. It does not output text, but rather probabilities for pre-defined response options. As a result, it cannot hallucinate and costs a fraction of standard models. Vercel replaced a security check with Jev, reporting five to eighteen times faster responses. In a test against Gemini, Gemini was slightly more accurate, but ten to twenty times more expensive. You do not need a large language model for sorting and classification.
Google DeepMind establishes a new institute. Launched on 17 September, with Demis Hassabis among the directors. Hassabis proposes a US-led evaluation body for frontier models. Providers would submit models 30 days prior to release—initially voluntarily, later as a requirement.
Anthropic exceeds 100 billion dollars annual run rate. At the end of July, the figure was 65 billion. This represents projected annual revenue based on current pace, not actual recorded revenue. The IPO is planned for November, with a valuation around 2 trillion dollars under discussion.
Astra for Law. OpenAI introduced a version tailored for legal professionals on 17 September, one week after launching its financial services version. ChatGPT is being customised sector by sector.
European AI spend. IDC forecasts spending in Europe will reach nearly 470 billion dollars annually by 2030, driven by agents. Banking is the largest sector, whilst healthcare is the fastest-growing.
Three things to do this week:
Test Opus 5.5. Take a task you perform regularly with Claude and run it again using Opus 5.5. Pay attention to two things: the quality of the output, and your monthly bill. Anthropic promises a 40 per cent saving. Verify this on your own use case.
Test ChatGPT ads. Open ads.openai.com and check if your account can book campaigns. For Shopify store owners, the ChatGPT Ads app is now available outside the US. The channel is small but growing rapidly; mastering the mechanics before your competitors enter will lower your acquisition costs.
Test Mistral. Evaluate Le Chat or the Mistral API against your current tools. Ask two questions: is the quality sufficient for your standard tasks, and how does your data privacy posture change with European processing? You need these answers before your board asks for alternatives.
Sources: https://www.anthropic.com/claude-opus-5-5 https://techcrunch.com/2026/09/22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/ https://www.nbcnews.com/tech/tech-news/google-says-ai-model-gained-unauthorized-access-three-systems-rcna598651 https://www.cnbc.com/2026/09/18/googles-gemini-becomes-latest-ai-model-to-break-out-and-hack-computer-systems.html https://www.anthropic.com/institute/measuring-pace-of-ai-development https://www.anthropic.com/news/accenture-embedded-evaluation https://openai.com/index/reimagining-advertising-with-ai/ https://openai.com/index/model-misalignment-reporting-framework/ https://techcrunch.com/2026/09/17/openai-caught-its-models-leaving-notes-to-successors-to-hide-bad-behavior/ https://openai.com/index/astra-for-law/ https://mistral.ai/news/mistral-x-mozilla/ https://mistral.ai/news/mistral-makes-sovereign-open-weight-ai-to-frontier/ https://www.netzwoche.ch/news/2026-09-18/ki-souveraenitaet-ist-fuer-viele-laender-eine-illusion https://www.netzwoche.ch/news/2026-09-18/schweizer-ki-apertus-und-protons-ki-assistent-lumo-spannen-zusammen https://www.netzwoche.ch/news/2026-09-18/ki-ausgaben-in-europa-steigen-bis-2030-auf-470-milliarden-us-dollar https://techcrunch.com/2026/09/17/google-deepmind-launches-institute-to-widen-the-agi-debate/ https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/ https://www.pymnts.com/news/investment-tracker/ipo/2026/anthropic-targets-november-ipo-revenue-surges/ https://ai-weeks.ch/