·
·
AI / Artificial Intelligence
Anthropic/Claude
OpenAI/ChatGPT
Mistral AI
Google / Gemini
·
LLM
Anthropic/Claude
OpenAI/ChatGPT
Google/Gemini
Mistral AI
European AI
Image generation
AI News Week 29 – The agent race reaches the office: Cowork mobile, OpenAI ChatGPT Work, Grok 4.5 and a German sovereign model
The AI agent race enters the workplace. Anthropic's Claude Cowork now runs on web and mobile. OpenAI counters with GPT-5.6 and ChatGPT Work. xAI launches Grok 4.5, a frontier-class model at half the price. Meanwhile, major platforms are challenging Anthropic’s MCP with their own standard, ARD. In Europe, Germany has launched Soofi S, an open 30B industrial model that is said to outperform Switzerland's Apertus. Finally, a new Microsoft study assesses AI productivity in Switzerland—though the findings require cautious interpretation.

1. Claude Cowork lands on Web and Mobile
Anthropic is bringing Cowork to web and mobile. Previously, the agent only ran in the desktop app. Now, sessions and files sync seamlessly across devices. The rollout has been underway since last week, starting with Max subscribers, and will continue over the next few weeks. Anthropic has extended the doubled usage limits until August 5.
Three things change in your daily routine:
Your work follows you. Start a task at your desk, review it on your phone whilst on the go, and access results from anywhere.
Tasks continue in the background. Close your laptop, and Claude keeps working – planned tasks run without requiring an active device.
Decisions come to you. Whenever a step requires your input, Claude asks for clarification; nothing is sent out without your explicit approval.
A figure from their blog puts things into perspective regarding what Cowork is used for: over 90 per cent of usage is not software development. The largest categories are Business Operations and Content Creation – together accounting for around half of the total.
Our take: This moves Cowork out of the desk-bound power-user niche. The clear edge over pure chatbots remains the human-in-the-loop principle: the agent works independently but requests your approval before any critical steps. This week highlights why this matters, as competitors attempt to follow suit with identical promises.
2. OpenAI counters with GPT-5.6 and ChatGPT Work
On July 9, OpenAI publicly released GPT-5.6 – claiming it as their most capable system yet. It comes in three tiers: Sol as the flagship model, Terra as the balanced mid-range option, and Luna as the fast, cost-effective version. OpenAI is now separating the naming convention: the number represents the generation, whilst Sol, Terra, and Luna signify permanent tier levels that will evolve independently.
Pricing per million tokens: Sol is $5 input and $30 output, Terra is $2.50 and $15, and Luna is $1 and $6. For context, Sol matches the pricing of Claude Opus 4.8 ($5 and $25), whilst Grok 4.5 undercuts both significantly ($2 and $6). The price war at the top continues.
More importantly for everyday business operations is ChatGPT Work. This agent is designed to complete entire tasks rather than just answering queries. It combines OpenAI’s coding tool, Codex, with ChatGPT and pulls context from your team's existing tools. GPT-5.6 will also become the default model in Microsoft 365 Copilot – within Word, Excel, PowerPoint, chat, and Copilot's agents. In tandem, OpenAI has acquired Northslope, whose engineers embed directly inside client organisations to build AI systems around their actual workflows – a clear signal that OpenAI is strengthening its enterprise arm around ChatGPT Work.
Our take: ChatGPT Work is OpenAI’s direct answer to Cowork and Claude Tag. Both providers are now selling the same proposition: an agent that completes work rather than merely providing answers. Anthropic maintains an edge in agentic knowledge work as Cowork, Claude Tag, and their Managed Agents have been running for weeks. A better benchmark does not automatically guarantee a better everyday experience – what matters is how reliably the agent completes real-world tasks. For organisations using Microsoft 365, the Copilot upgrade is the most practically relevant point: the default standard upgrades automatically.
3. Grok 4.5 – Opus-class performance at half the price
SpaceXAI unveiled Grok 4.5 on July 8. Elon Musk describes it as "an Opus-class model, but faster, more token-efficient, and cheaper". The model is built on their 1.5-trillion-parameter V9 foundation and trained using real developer sessions from Cursor – meaning it leverages data from actual work in real codebases over long sessions.
Grok 4.5 costs $2 per million input tokens and $6 per million output tokens. In comparison, Claude Opus 4.8 sits at $5 and $25. Grok 4.5 is available within Grok Build, on all Cursor plans, and via the SpaceXAI console. It is not yet approved in the EU, with availability expected in mid-July.
Our take: Three models in a single week – GPT-5.6, Grok 4.5, and Cowork mobile – all targeting the same objective: agentic work, cheaper and faster. Grok is betting heavily on developers, leveraging its new corporate proximity to Cursor (housed under the same roof since the SpaceX acquisition). Price pressure on premium models is intensifying.
4. The battle for the agent standard: MCP vs ARD
A strategic conflict is emerging. According to "The Information", a broad alliance is backing a new technical standard this week to connect AI agents with enterprise software: Agentic Resource Discovery (ARD). This standard is supported by Google, Microsoft, Salesforce, Snowflake, ServiceNow, Cisco, Databricks, GitHub, NVIDIA, and Hugging Face. Licensed under Apache 2.0, it is developed in open repositories and managed by the Linux Foundation.
This directly challenges Anthropic. Their Model Context Protocol (MCP) has quietly become the de facto standard over the past 18 months – with OpenAI relying on it as well. Crucially, MCP is itself an open standard that anyone can use freely, and the ARD supporters have been doing so for some time. Since ARD is also open (Apache 2.0, Linux Foundation), this is not a battle of open versus closed, but rather about which standard will dominate and dictate the direction of the industry.
Our take: For you as a user, nothing changes in the short term – MCP remains operational. However, this reveals where the real power struggle lies: not in who has the most elegant chat interface, but in which interface connects agents to your systems. Controlling the standard means controlling the ecosystem. The fact that the largest software giants have united to counter Anthropic’s MCP is a testament to its success.
5. New image models from Meta and Google
Two major tech giants have upgraded their image creation suites.
Meta Muse Image. Meta released its first proprietary image model on July 7, built by its Superintelligence Labs (internally codenamed "Mango"). It is available free of charge within the Meta AI app, Instagram Stories, and WhatsApp. Muse Image operates agentically: it browses the web, writes and executes code, and refines its own output. It can merge multiple photos into a single image and edit via prompts – such as placing a person in front of a landmark or removing background clutter. Controversy arose immediately: users can modify images from other Instagram accounts using AI, provided the profile is public.
Google Photos Video Remix. Powered by the Gemini Omni model, Google Photos is introducing cinematic effects to user videos starting July 8: including new lighting, background replacement, and artistic styles. The rollout is underway for eligible subscribers.
Our take: Image generation is turning agentic and moving directly into the everyday apps used by billions. For content teams, this lowers barriers to entry even further. However, Meta's feature allowing users to modify external public photos raises data privacy concerns – corporate Instagram users should quickly audit their photo visibility settings.
6. Europe and Switzerland: Specialisation over Chatbots
Whilst major US labs compete over general-purpose agents, Europe is focusing on a different strategy.
Soofi S – Germany's sovereign industrial model. A German consortium released Soofi S on July 13, an open-source 30B model engineered for industrial AI. The project is coordinated by the German AI Association (KI-Bundesverband) and funded by the Federal Ministry for Economic Affairs under the European IPCEI-CIS initiative. Partners include the Fraunhofer Institutes IAIS and IIS, DFKI, TU Darmstadt, University of Würzburg, alongside Ellamind and Merantix Momentum.
On a technical level, Soofi S is built for efficiency: only 3.2 billion of its 31.6 billion parameters are active per query – reducing computing demands and overall costs. A specialised architecture ensures the model remains fast and cost-effective even with very long texts – such as extensive technical documentation or contracts, where traditional models typically experience latency and cost spikes. It was trained primarily on German and English datasets in the Deutsche Telekom cloud in Munich. According to the pretraining report, Soofi S achieves the highest scores among fully open-source models on German and English benchmarks – surpassing the Swiss Apertus 70B.
The base weights are available on Hugging Face. However, a general version for direct use is not yet live; a testing phase is currently underway with industrial partners, and the consortium is actively looking for more companies to participate.
Our take: Soofi is not built for consumer chat, but for technical documentation, code, and enterprise agentic systems – hosted on sovereign European infrastructure. This is a significant step forward for the digital sovereignty debate: introducing the first fully open-source DACH model that is verifiably more capable than Apertus. Germany and Switzerland are pursuing the same vision with different priorities: Soofi targets industrial applications, whilst Apertus remains a public good.
Mistral Leanstral 1.5. Mistral released an open-source model on July 2 that mathematically proves software works correctly – rather than simply testing it. The difference is critical: a test merely shows that a program works under specified test cases; a mathematical proof guarantees it operates correctly under all scenarios. According to Mistral, the model has already identified previously unknown bugs in existing open-source codebases.
This is highly relevant for safety-critical software – in medicine, aviation, or financial systems – where "usually works" is not an option. Europe's competitive edge lies not in building the next chat assistant, but in delivering specialised, reliable AI systems like these.
Swiss businesses use AI productively – but without structure. Microsoft Switzerland published its "Work Trend Index". The study is based on anonymised productivity data from Microsoft 365 alongside a survey of 20,000 AI users across ten countries. 65 per cent of AI users in Switzerland state they now achieve results that were impossible a year ago. Globally, this figure stands at 58 per cent.
The catch: this advantage relies heavily on individuals, not strategic implementation. Only 24 per cent say executive leadership is aligned on AI strategy. Only 48 per cent feel secure enough to focus on current goals. Consequently, productivity gains are rarely used to redesign workflows. This matches BDO's analysis: the gap between large corporate enterprises and SMEs is widening – large organisations are integrating AI into their value chains, whilst many SMEs remain stuck at individual use cases.
A caveat on the study: it is published by Microsoft, the vendors of Copilot. In practice, Copilot is often rolled out poorly, and usage remains ad-hoc. Many employees are disappointed by the performance and bypass IT to use ChatGPT or Claude – a form of shadow IT that these numbers rarely capture.
In brief
"We must act now". Over 200 economists and AI researchers, including 16 Nobel laureates, signed a manifesto by the Stanford Digital Economy Lab on July 13. Organised by Erik Brynjolfsson and colleagues, it warns of a disruption "larger than the Industrial Revolution, but occurring in a much shorter timeframe" – risking widespread job displacement. Executives from Anthropic, Google, and OpenAI also signed. Their demand: policymakers, business leaders, and researchers must build institutions and regulations now to ensure AI augments rather than replaces human capabilities.
Apple sues OpenAI. On July 11, Apple filed a lawsuit alleging the theft of trade secrets. The core issue: over 400 former Apple employees have moved to OpenAI, many from silicon and on-device AI teams. The dispute escalated over the weekend into a public feud between Elon Musk and Sam Altman on X.
Claude Corps. Anthropic is launching a paid 12-month fellowship program that trains future AI specialists directly within non-profit organisations.
GPT-Live. OpenAI introduced new full-duplex voice models (GPT-Live-1 and mini). They listen, speak, and process information simultaneously rather than taking turns – featuring real-time translation, web search during conversation, and task delegation to other agents.
AI models remain vulnerable. A Cisco report warns that scoring well on basic safety benchmarks means little – multi-stage attacks continue to bypass AI model safeguards reliably.
Outlook. Leaked plans suggest Google will make Gemini 3.5 Pro generally available on July 17 – currently unconfirmed.
Three things to focus on this week
1. Try Cowork on mobile if you have a Max plan. Your work now transitions seamlessly to your mobile, and scheduled tasks run in the background without needing your laptop active. This is the shift from chatbot to executive agent – make sure you keep an eye on approvals for critical tasks.
2. Audit your Microsoft 365 environment. GPT-5.6 is becoming the default model in Copilot – across Word, Excel, PowerPoint, and chat. This transition happens automatically. Assess how this affects your team's workflow.
3. Close your strategic gaps. Switzerland leads in individual AI productivity, but lacks organisational structure. True ROI comes from redesigning business processes and aligning executive leadership, not from deploying yet another tool.