·

·

AI / Artificial Intelligence

Anthropic/Claude

Mistral AI

·

LLM

Anthropic/Claude

Mistral AI

European AI

Agnostic

AI for business: Choosing the right model, infrastructure, and the true cost

95 per cent of corporate AI pilots fail to deliver measurable value. This is not due to poor models. It is because of a lack of strategy, infrastructure and adoption. Success requires the right model for the right use case, an adaptable infrastructure, and costs that scale without destroying your margin.

Only 5 per cent of all corporate AI pilots create measurable value for revenue or profit. This is no rumour. It is the result of a 2025 MIT/NANDA study evaluating over a thousand companies. The other 95 per cent? Investment with no provable effect.

Not because the models are bad. But because strategy, infrastructure, and adoption are missing.

This article shows how to do it differently: what AI should actually achieve in a business, how to build a sensible infrastructure, which model fits which purpose—and why the subscription era is ending right now.

The problem: how AI lands in companies today

Most organisations are in an early stage: a few employees use ChatGPT on their own initiative. The IT department blocks it—or rolls out Copilot because it is integrated into M365 and sounds "safe". There is no clear strategy, no defined use cases, no training.

The result is predictable.

Unstructured use without strategy. Simply distributing an AI tool without defining its purpose creates chaos. Everyone prompts differently, no one exploits the full potential, and results are inconsistent.

Shadow AI as standard reality. 80 per cent of employees use AI tools that their company has not approved—among security professionals, it is nearly 90 per cent (UpGuard 2025). The "bring your own AI" phenomenon is no longer an exception: according to the Microsoft Work Trend Index 2024, 78 per cent of employees bring their own AI tools. Not out of malice—but because official solutions are missing or too restrictive. 82 per cent of business-relevant copy-paste operations come from unmanaged personal accounts, 22 per cent of which contain personal data, credentials, health information, financial data, or other confidential content (LayerX 2025).

No one can prompt properly—and that won't change. Expecting all employees to learn prompt engineering is unrealistic and unnecessary. If the knowledge of how to brief an AI well lies with only a few, only those few gain productivity. The rest work with half-baked results.

Costs spiral out of control. What starts as a CHF 20/month subscription can quickly become a significant budget item in real corporate use. At the same time, there is little transparency on who consumes how much and for what tasks.

Data protection remains unresolved. Which data goes to which model? Which jurisdiction applies to the provider? These questions remain open for many organisations—meaning sensitive content is inadvertently shared with cloud services outside their control.

Myths & Facts

Before looking at solutions: a quick fact-check on the most common beliefs about corporate AI.

MYTH: "95 per cent of Copilot pilots are successful." Opposite. According to analysis, most corporate AI projects fail—only 5 per cent ever get rolled out due to a lack of business cases, poor data quality, and weak results. Copilot's accuracy NPS sits at –24 (industry average: +30 to +40). Even Microsoft CEO Satya Nadella said bluntly in December 2025: "For the most part they don't really work, and are not smart." (The Information, Dec 2025)

MYTH: "Prompt engineering is the key skill everyone must learn." Modern language models understand user intent well—what they need is context: about the company, the industry, the specific use case. This context knowledge belongs in company-wide skills and plugins, not in the head of every single employee. The difference is context engineering instead of prompt engineering—and you distribute that via infrastructure, not training.

MYTH: "A more restrictive AI policy means less risk." False. 80 per cent of employees already use unapproved AI tools. A ban without an alternative drives usage into the shadows—with far less control. Structured, secure solutions create more safety than bans.

MYTH: "The best model wins." The "best" model changes almost monthly. Top models usually perform within each other's margin of error, and benchmarks are often biased because models were trained on those exact test questions. What wins in practice: adoption and the workflow around the tool—not the benchmark leader.

FACT: "AI lifts overall team performance." True. In a Procter & Gamble study with 776 professionals, a single person with AI support was as productive as an entire team without AI—and 12–16 per cent faster. AI primarily boosts weaker performers, not just the stars. Judgement and quality control remain human.

MYTH: "In-house development beats partnerships with external providers." No. According to MIT/NANDA, purchased solutions and provider partnerships succeed in ~67 per cent of cases, whilst internal developments succeed in only ~23 per cent. Buying is statistically about three times as successful as building it yourself.

MYTH: "Europe is lagging far behind in AI models." No longer true. The US and China lead at the absolute frontier level—but Europe is closer than often claimed: Mistral (FR), Black Forest Labs FLUX (DE), DeepL (DE), and Apertus 1.5 from Switzerland matches Google's Gemma model, is multimodal and capable of reasoning. In practice, around 70 per cent of everyday office tasks do not need a frontier model at all.

What AI should actually achieve in a business

Efficiency is the most frequently cited reason for AI investments—and it misses the point.

The larger opportunity lies elsewhere. In my seminars with several hundred participants from Swiss companies, the data shows: 43 per cent use the time saved for strategic tasks, 29 per cent for more creative work, 14 per cent work fewer hours. AI does not replace jobs—it offloads them. Potential flows back into work that creates real value.

What a good AI infrastructure should deliver:

Broad-based efficiency. Not just individual power users who prompt well should benefit—everyone should. This is only possible with well-developed skills and plugins that democratise context knowledge. A junior with a good company plugin can deliver senior-level output.

Better decisions. AI can read large amounts of data, deliver summaries, and structure options. This improves decision-making—if the infrastructure is right.

Innovation and hard problem solving. Once routine work is automated, the team has the capacity for the difficult questions that previously went unanswered.

Controllable costs. Growing usage must not lead to linear cost growth. A smart infrastructure separates expensive frontier models for complex tasks from cheaper or proprietary models for volume tasks.

AI is not a tool—it is a stack

The biggest misunderstanding in practice: buying and rolling out AI as a single product. Relying on one tool tethers your strategy to the roadmap and pricing of a single vendor.

A robust AI infrastructure is a stack of interchangeable layers—and keeping these layers clean keeps you free.

Layer 1 – Interface. What employees use daily. Crucial: a stable, uniform interface for everyone. Productivity and user experience are built here. If the underlying model changes, users ideally do not notice.

Layer 2 – Models (interchangeable). Frontier cloud (Anthropic, OpenAI), European (Mistral, Apertus), or run locally on your own servers—on-premises, meaning on your own hardware in-house or with a Swiss hoster like Infomaniak. Chosen per task, replaceable at any time. Claude for complex analysis today, Apertus for sensitive customer data tomorrow—without rebuilding the infrastructure.

Layer 3 – Connectors. Access to data, web, office suite, CRM, and other tools. This is where AI connects to what already exists in the company.

Layer 4 – Skills & Plugins. The decisive multiplier. This is where codified company knowledge lives: wording, processes, use cases, brand rules. A rollout without skills is a rollout without impact. Skills lift juniors to senior level—and the knowledge stays in-house, not in the heads of individuals.

Layer 5 – Knowledge base. The memory of the company. Structured and accessible via retrieval (RAG)—so the AI brings company knowledge, not just general knowledge.

Governance all around. Policies, access control, logging, reporting. What can which tool see? Who uses what, how much, and what costs are they generating? These questions must be answered before you scale.

The guiding principle: model-agnostic behind a stable interface. The interface remains—the model behind it is swapped when a better, cheaper, or more secure one arrives. Models are renamed, retired, or geoblocked. (In June 2026, Anthropic's Fable 5 / Mythos 5 were temporarily blocked for non-US users; at the same time, OpenAI withheld GPT-5.6 by order of the US government.) If you relied on a single model, you were blocked at that moment.

Comparing the models

There are currently five realistic categories for corporate use. No single model wins across all of them—each has its strengths.

Claude (Anthropic)

The strongest model portfolio for complex, multi-step tasks. Claude reads large documents fully, works in tasks (not just chats), conducts research, writes code, analyses data, and can handle routines fully automatically via scheduled tasks.

Via Claude Cowork, the model runs directly—without middleware. This is crucial: larger, multi-step tasks do not get lost in chunk retrieval or safety filters. Claude can work with local files, remember context permanently (CLAUDE.md), and be active in multiple tasks in parallel.

MCP and the plugin ecosystem. Anthropic developed the Model Context Protocol (MCP) as an open standard—now the de facto standard for connecting AI assistants to external tools. This means a growing ecosystem of connectors for CRM, ERP, calendar, email, project management, databases, and more. Claude benefits first. Building a connector for Claude means building it for the entire ecosystem.

Additionally, there are high-quality skills for Microsoft Office: creating PowerPoint decks, structuring Word documents, conducting Excel analyses—directly from Cowork. Recently, it also gained write access to M365, so it can write directly into existing documents, not just read them. Companies using Cowork with these skills report higher satisfaction than with standard Copilot solutions—because the task is solved in one go rather than fragmented by middleware.

Important note on data residency: claude.ai (the browser interface) runs on US servers. Since around 90 per cent of productive work in practice happens in Cowork—not in the browser chat—this is resolvable: in 3P mode (third-party provider mode), Cowork runs via AWS Bedrock or Google Vertex, Frankfurt region, with EU data residency.

Jurisdiction: US (AWS/Anthropic). Mitigated via 3P mode, but technically not fully removed from the US CLOUD Act.

Strength: Agentic, complex, large tasks. MCP ecosystem. Strongest Office integration via skills. The most powerful assistant for daily work.

Microsoft 365 Copilot

The most obvious choice for companies already operating in the M365 environment—and the most frequent disappointment.

Copilot is strong for small, context-specific queries directly within Office applications: "Summarise this email", "What is in this document?". Deep Office integration is a genuine advantage.

The problem lies in the architecture. Copilot works via a "semantic index": documents are split into fragments, and only the most similar fragments enter the model per query. What lies outside the top matches is lost—the so-called "lost in the middle" effect (Liu et al. 2023). For large documents or multi-step tasks, this significantly weakens quality—even with the same powerful model behind it.

Added to this is governance overhead: every query runs through identity checks, access controls, data protection labels, security filters, and audit logging—which takes time and resources. Great for compliance—but because users are impatient, the system optimises for quick single answers, not deep analysis.

Agentic use (Copilot Researcher, Cowork) is possible, but runs on paid credits in addition to the flat-rate licence.

Copilot can run with various models—including Claude by Anthropic or other frontier models. The limitation remains: even the strongest model delivers weaker results if the middleware in front of it cuts context.

Another detail rarely communicated: Claude models in Copilot are disabled by default in the EU and EFTA, running outside the EU Data Boundary. The new "Flex Routing" explicitly allows Copilot to route queries outside the EU boundary during peak loads.

Jurisdiction: US (Microsoft). EU Data Boundary present, but with limitations.

Strength: M365 integration, small context-specific queries in daily office work. Weak on large, multi-step tasks.

Mistral (Le Chat / Le Chat Enterprise)

The strongest European provider—and the only one among these tools not subject to the US CLOUD Act. Mistral is a French company; data remains under European jurisdiction.

Mistral Le Chat connects directly to the model—without middleware. It has native connectors for SharePoint, OneDrive, and Google Drive, and can read, edit, and save M365 files in the cloud or download them as Office files.

Mistral's open-weight models can be run on-premises—no cloud, no external data transfer.

Mistral is more than a single language model. The company offers a broad portfolio: codex-capable models for software development, image generation models, embedding models for knowledge retrieval—and is increasingly working on physical AI (AI for robotics and embedded systems). Building a European strategy around this yields more than just a chat assistant.

On the most difficult tasks, Mistral currently lags behind US frontier models. For the vast majority of office tasks—and particularly for data-sensitive use cases—this is not a dealbreaker.

Jurisdiction: France (EU). No US CLOUD Act.

Strength: Sovereignty, IP-sensitive content, cheapest cloud option, on-premises possible.

Apertus 1.5 (Switzerland)

The Swiss option—and one to keep on your radar. Apertus 1.5 is a Swiss open-weight model that matches Google's Gemma family, is multimodal, and supports reasoning.

Crucial to understand: Apertus 1.5 is a pure language model—not a finished product with its own interface. To use it, you need a client: Claude Cowork, Copilot, or tools like LM Studio for local use. And it requires infrastructure: your own in-house servers or rented capacity at a Swiss hoster like Infomaniak. This is one step more than cloud services—but all data remains entirely in Switzerland.

Particularly relevant for companies with strict data sovereignty requirements: Apertus can run on-premises or on Swiss servers, understands Swiss German, knows the Swiss context better than US models, and is subject exclusively to Swiss law.

For daily office tasks, internal knowledge retrieval, or as a local baseline AI behind a stable interface, Apertus is a serious option—especially for public administrations, healthcare facilities, or financial service providers with strict data protection requirements.

Jurisdiction: Switzerland. No US CLOUD Act, no EU law—pure Swiss jurisdiction.

Strength: Swiss data sovereignty, Swiss German, open-weight for on-premises.

Chinese Open-Weight Models

DeepSeek, Qwen, and GLM are technically among the most capable open-weight models globally and are available for free. If you run your own GPU infrastructure, you can execute these models locally—no cloud, zero marginal cost per query.

Kimi K3 (Moonshot AI, 2.8 trillion parameters) ranks 3rd globally in benchmarks, behind Fable 5 and GPT-5.6—and even leads all US models in frontend coding. DeepSeek V4 Pro is the most cost-effective all-round model with an MIT licence, leading in agentic coding among freely downloadable weights. Qwen3 (Alibaba, Apache 2.0) is the most cost-efficient option, whilst GLM-5.2 (Zhipu AI) is the sharpest coding engine. Overall, these models now match the level of US frontier models from a year ago—at a fraction of the operating cost.

The caveat is significant: these are Chinese models. For many companies, particularly in regulated industries, finance, or defence, this is a governance, procurement, and reputation issue that must be decided independently of technical quality.

For high-volume scenarios without IP sensitivity where your own infrastructure exists, they can offer major cost optimisation.

Jurisdiction: China (technically on-premises if self-hosted—but origin and supply chain risks remain).

Strength: Performance, zero marginal operating costs, on-premises possible.

Overview: what to choose for what?

Criterion

Claude Cowork

MS 365 Copilot

Mistral

Apertus 1.5

CN Open-Weight

Complex / large tasks

✓✓ strong

○ weaker

✓ strong

○ solid

○ varies

M365 integration

✓ local + 3P

✓✓ native

✓ SharePoint/OneDrive

EU data residency

✓ via 3P (DE)

✓ with limitations

✓✓ (FR)

– (CH)

✓ on-premises

No US CLOUD Act

✗ (Not via Bedrock, but with local models)

✓✓

✓✓

✓ on-premises

On-premises possible

✗ (Cloud)

✓✓

Swiss context

✓✓

Marginal cost at volume

high

Flat rate + credits

low

~0 on-premises

~0 on-premises

Governance overhead

low

high

low

low

low

Digital sovereignty—two axes often confused

"Sovereignty" is not one problem, it is two—and they require different answers.

Data residency (where data is stored): solvable. Claude via 3P in Frankfurt, Copilot in the EU Data Boundary, Mistral in France, Apertus and on-premises at your own site. For most companies, residency is manageable.

Jurisdiction (who can force access): the harder part. Microsoft, AWS, Google, and Anthropic are US companies subject to the US CLOUD Act—regardless of where data is physically stored. Microsoft publicly admitted in 2025 that it cannot guarantee EU data remains shielded from US authorities. Only a European provider (Mistral) or genuine on-premises setups avoid this.

The pragmatic recommendation: tier based on data sensitivity.

  • For around 70 per cent of everyday office tasks, capability and adoption are more important than residency.

  • For truly sensitive data—prototypes, R&D, finance, patient records, IP—use EU/CH hosting or on-premises with open-weight.

Sovereignty is a legitimate argument. But an infrastructure operating with interchangeable models and clear data sensitivity tiering solves the problem practically—without sacrificing frontier performance.

The token economy: why cheap subscriptions are ending

When AI becomes every employee's starting tool—the first thing they open in the morning, even before Outlook—token volume explodes. Every query, every analysis, every piece of generated text creates tokens. This scales.

Brief cost reality (output prices, as of July 2026):

Model

Output price / 1M tokens

At ~100M tokens / month

Claude Opus (Cloud)

~USD 25

~USD 2,500

Claude Sonnet (Cloud)

~USD 15

~USD 1,500

Mistral Large 3 (Cloud)

~USD 1.50

~USD 150

Open-weight, own GPUs

~USD 0 marginal

The subsidy war among providers is ending. Sam Altman stated publicly in January 2025 that OpenAI loses money on the USD 200/month Pro subscription—he had set the price himself expecting a profit. Cursor abruptly shifted to a credit-based model in mid-2025 and had to issue a public apology. Anthropic introduced rate limits without warning.

The direction is clear: basic subscription for access, usage-based billing past a threshold. Power tiers at USD 100–200/month are standard (ChatGPT Pro, Claude Max, Google AI Ultra). Microsoft 365 Copilot costs around USD 30/user/month flat rate—plus usage-based credits for agentic tasks.

Solving this solely with subscriptions will surprise you in real corporate use. The answer is a low-token-cost layer: a self-hosted open-weight model for volume tasks. Expensive frontier cloud for the hardest tasks—cheap or proprietary models for the masses. CapEx instead of OpEx, predictable costs regardless of usage.

How to proceed now

Do not start with the tool. Start with the problem.

1. Appoint project leads and set up the project. AI initiatives without clear ownership fail. Someone must own the project—with mandate, budget, and time. This is not an IT task or a marketing task. It is a leadership task.

2. Find, prioritise, and select use cases. Which tasks repeat daily? Where is time lost on routines that require no thought? Where do decisions stall due to poor or missing data? Two to three concrete use cases are worth more than ten vague ideas. Criteria: frequency, time spent, measurability of improvement.

3. Test use cases in different models and setups. Do not decide based on marketing claims—try them out. Claude for the complex use case, Mistral for the data-sensitive one, Copilot for native Office tasks. Reality in your own environment matters more than any benchmark.

4. Develop skills, name AI Champions, and train teams. This is the step where most fail. A rollout without company-specific skills is a rollout without impact. Skills codify the context knowledge everyone needs—and make good results repeatable, regardless of who is operating the device. AI Champions are the team members who dive deep and help others progress. Not everyone needs to know everything—but someone must truly master the tool.

A Copilot rollout without training and company-specific skills is an expensive licence nobody uses. That is not a criticism of the employees—it is a structural issue.

Conclusion

The AI technology is ready. The models are good enough. What is missing is the right infrastructure and a strategy that goes beyond "buy and distribute a tool".

The key principles:

  1. AI is a stack, not a product—interface, models, connectors, skills, knowledge base, governance. Keeping the layers separate keeps you free.

  2. Relying on a single model or vendor binds your strategy to their roadmap and pricing. Model availability is a geopolitical risk.

  3. The path to the model determines quality: direct interfaces beat middleware architectures for complex tasks—even with the same model behind them.

  4. Think of sovereignty on two axes: residency is solvable, jurisdiction is the harder part. Tiering by data sensitivity is more robust than a one-size-fits-all solution.

  5. The cheap subscription era is ending. Scaling requires a low-token-cost layer—self-hosted open-weight for volume, frontier cloud for complexity.

And the most critical factor remains: adoption. The best model does not win. The model people actually use wins—because the infrastructure is right, the skills exist, and the training has happened.

Ready to get serious about AI?

30-minute initial consultation – free and non-binding. We will review together where you stand and what the right first step is.

Ready to get serious about AI?

30-minute initial consultation – free and non-binding. We will review together where you stand and what the right first step is.

Welche Newsletter möchtest du abonnieren?
Bitte wähle mindestens einen Newsletter.
Your registration was successful.
Your sign-up could not be saved. Please try again.
Welche Newsletter möchtest du abonnieren?
Bitte wähle mindestens einen Newsletter.
Your registration was successful.
Your sign-up could not be saved. Please try again.