By Thorsten Meyer

I want Europe to have a sovereign frontier lab. I do not care whether it is Mistral.

That distinction matters, because everything that follows is critical, and it would be easy to mistake it for disappointment in a favourite team. It isn’t. I have argued for a year that European AI sovereignty is real, that a continent which only consumes intelligence made elsewhere has outsourced the layer that matters, and that Europe needs a frontier-grade lab of its own — a real alternative to two American frontier labs and one Chinese bloc, something governments and regulated enterprises can build on instead of routing their sovereignty through a foreign API. Mistral is the company currently cast in that role: the one carrying the European-champion narrative, the default name in every sovereignty deck.

So I went looking, on the independent numbers rather than the launch decks, for evidence that the champion Europe has anointed is actually at the frontier. And the honest finding is the one that should worry anyone who wants European sovereignty to be real regardless of whose logo delivers it: it isn’t — and the gap is widening, not closing.

AI DISPATCH · REALITY CHECK Mistral vs the frontier · 6 Aug 2026
The European champion, on the independent numbers
Europe’s Frontier Lab Isn’t at the Frontier

I want Europe to have a sovereign frontier lab. I don’t care whether it’s Mistral. So I went looking on the independent benchmarks for evidence the anointed champion is at the frontier. The honest finding should worry anyone who wants EU sovereignty to be real: it isn’t, and the gap is widening.

▲ Opinion · loyal to the goal, not the mascot
30
Mistral Medium 3.5 · their best · AA Index
56–61
The current frontier · ~2× Mistral’s best
= 30
Claude 4.5 Haiku · a rival’s cheapest tier
~€20B
Valuation · a geopolitical premium
01
The comparison that should not be possible

Artificial Analysis Intelligence Index (v4.1) — the independent composite of nine evals including agentic coding, tool use, and reasoning. Mistral’s strongest current model against the field.

Claude Opus 5
frontier
61
the frontier
GPT-5.6 Sol
frontier
59
the frontier
Claude 4.1 Opus
old, superseded
34*
*AA estimate
Mistral Medium 3.5
their current best
30
Europe’s flagship
Claude 4.5 Haiku
a rival’s cheapest
30
budget tier
Europe’s flagship frontier model is level with a competitor’s budget tier — the model you reach for when you explicitly do not need intelligence — and trails a rival’s year-old, already-superseded flagship. The measured comparison is the damning one.
02
The slope, not the score

A snapshot could be a bad quarter. The trajectory is the structural finding: on Artificial Analysis’s intelligence-over-time chart, Mistral’s line is the flattest of any major lab.

2023 2026 60 0 the field → 56–61 Mistral → 30
Everyone else climbed from single digits to the high fifties. Mistral crawled to about thirty. The gap isn’t constant — it’s growing, generation over generation. A lab a fixed distance behind can catch up. A lab whose gap widens is on a different curve, and different curves don’t converge on their own.
03
Not even the cheap option

The obvious defense — “not the smartest, but the efficient workhorse” — doesn’t survive the cost data. Cost per Intelligence Index task, at each model’s measured intelligence.

Mistral Medium 3.5
30
intelligence
~$0.46
per task
Claude 4.5 Haiku
30
same intelligence
~$0.22
half the price
DeepSeek V4 Flash
50
far smarter
~$0.03
~1/15 the price
Dominated on price by a cheaper model of equal intelligence; buried on capability by cheaper models of far greater intelligence. Neither the smartest nor the cheapest in its own price band — a strategically homeless position.
04
The honest case — and why I’m hard on them anyway

The Index measures intelligence. It doesn’t measure what Mistral actually sells. Both columns are true.

The genuine case for Mistral
  • Open weights the benchmark can’t see — run it in your own jurisdiction, a real product Anthropic and OpenAI structurally can’t match
  • Sovereignty is the spec for EU defense, institutions, regulated buyers — not the score
  • Real infrastructure: €4B data centers, France + Sweden, partly nuclear; ASML’s ~11% stake
  • On ~1/10 the capital of US rivals — remarkable for a 3-year-old
Why the curve is the wrong grade
  • Europe is concentrating its AI independence behind one lab, at a ~€20B geopolitical premium
  • If the anointed option ties a rival’s cheapest model, sovereignty is being narrated, not secured
  • Loyalty to the goal not the logo turns a flat line from tragedy into information: Europe needs more shots on goal
  • The actually pro-sovereignty move is to stare at the numbers — the goal matters more than the mascot
Europe deserves a real frontier lab. The company it anointed isn’t there yet —
which is an argument for more contenders and less loyalty to any one mascot. The goal is the point.

The comparison that should not be possible

Here is the single fact that reframes the whole story. On Artificial Analysis's Intelligence Index — the independent composite that runs nine evaluations including agentic coding, tool use, and reasoning — Mistral's strongest current model, Mistral Medium 3.5, scores 30.

Now place that against the field. The current frontier sits at 56 to 61: Claude Opus 5 at 61, Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57. Mistral's best is roughly half the frontier score. That alone would be a gap. It is what Mistral ties at 30 that turns a gap into an indictment: Claude 4.5 Haiku — Anthropic's cheapest, smallest model, the one you reach for when you explicitly do not need intelligence — also scores 30. Europe's flagship frontier model is level with a competitor's budget tier.

Push one step further, into the comparison a reader can build for themselves on the same tool, and it gets worse. Anthropic's older, superseded Opus models — Claude 4.1 Opus, a model already replaced by 4.5, then 4.8, then 5 — sit at an estimated 34, above Mistral's current best. An honest caveat: those older-Claude figures are Artificial Analysis estimates pending full evaluation, marked as such, so I lean on the measured comparison rather than the estimated one. But the measured comparison is the damning one. Mistral's current flagship does not merely lose to the frontier. It ties the cheapest current model from one rival and trails the year-old flagship of another.

AI DISPATCH · REALITY CHECK · COMPANION What "30" means · 6 Aug 2026
Reading the Intelligence Index in context
What a Score of 30 Actually Means in August 2026

Mistral's best model scores 30 on the Artificial Analysis Intelligence Index. Abstract number — until you ask two questions: what does the index measure, and what year does a 30 belong to?

01
The index measures the work that pays

Nine evaluations, weighted toward what people actually pay models to do in 2026: agentic knowledge work (GDPval), terminal & coding tasks, tool use, long-context reasoning, hallucination resistance.

A 55+ model
Carries the task
  • Plans, calls tools, recovers from errors, finishes
  • Holds a multi-step agentic workflow end to end
  • Gets work delegated, not just assisted
A 30 model
Needs its hand held
  • Stalls or loses the thread mid-task
  • Needs human supervision at every step
  • Below the threshold where 2026's work gets delegated at all
The gap between 30 and 55 isn't nicer prose. The scores separate exactly where the economic value now lives.
02
The time machine: 30 was the frontier — in early 2025

Same index, same chart, eighteen months apart. Mistral's flagship has arrived at last year's frontier, on time for a market that no longer exists.

early 2025 late 2025 Aug 2026 ~30 frontier then 40+ frontier 56–61 frontier now 30 Mistral today same level, 18 months later
A 30 would have been in the frontier conversation in early 2025. Since then the field crossed 40, then 50, and now sits at 56–61. The number didn't get worse — the year it belongs to did.
03
How fast the others moved

The pace-setters weren't only American. Two ecosystems climbed twenty points in the time Europe's champion climbed ten.

US labs
~40 → 61
Anthropic, OpenAI in twelve months. Claude Opus 5 at 61, GPT-5.6 Sol at 59 — and Meta went 43→54 in four months.
Chinese labs
+10 / release
DeepSeek jumped ten index points in a single release this summer. Kimi K3 (57), Qwen3.8 Max (56), GLM-5.2 (51), DeepSeek (50) — several open-weight, several radically cheaper.
Europe's champion
~20 → 30
Mistral over the same period — the flattest line of any major lab, now 20+ points behind four Chinese models alone.
The sovereignty question has a second front. Europe's champion isn't just behind the American frontier — it's twenty-plus points behind the open-weight Chinese models anyone can download and run.
04
The blunt version

Stated as the opinion it is — and the reason the gap can't be waved away with procurement mandates.

Opinion
Nobody chooses to work with a far less intelligent model. Not out of loyalty, not for the flag, not for long.
Intelligence is the product. The developer whose agent stalls where a rival's finishes, the enterprise whose workflows need double the supervision, the ministry whose sovereign deployment is eighteen months behind what its staff use privately at home — they drift to the capable option, or quietly wrap the sovereign model in a foreign one. Sovereignty can make a buyer tolerate a small gap. It cannot make anyone tolerate half the intelligence — and this year's market includes a 50-point model at three cents a task.
Amazon

European AI frontier research labs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The number that actually matters is the slope, not the score

A single snapshot could be a bad quarter. What makes this a structural finding rather than a bad-week story is the trajectory, and the trajectory is the part that genuinely worries me.

Artificial Analysis plots frontier intelligence over time, every major lab as a climbing line. Anthropic, OpenAI, Google, the Chinese labs — they all march up and to the right, from single digits in 2023 to the high fifties and low sixties now. Mistral's line is on the same chart, and it is the flattest of any major lab. It crawled from near zero to about thirty while the field climbed past fifty-five. The distance between Mistral and the frontier is not constant — it is growing, release over release, because everyone else is climbing faster than Mistral is.

That is the thing a valuation cannot paper over. A lab that is a fixed distance behind is a lab you can imagine catching up. A lab whose gap widens with every generation is on a different curve, and different curves do not converge on their own. The snapshot says Mistral is behind. The slope says Mistral is falling behind. Those are different diagnoses, and the second is the serious one.

Media Company in a Box: Build an Independent Media Business with AI: Answer Engine Optimization, Podcasting, and Nine Revenue Streams for the Sovereign Creator

Media Company in a Box: Build an Independent Media Business with AI: Answer Engine Optimization, Podcasting, and Nine Revenue Streams for the Sovereign Creator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What a 30 actually means in August 2026

Because "30 on an index" is abstract, let me make it concrete in the two ways that matter: what the number measures, and what year it belongs to.

The Intelligence Index is not trivia. Its nine evaluations lean heavily on the things people actually pay models to do in 2026 — agentic knowledge work, terminal and coding tasks, tool use, long-context reasoning, hallucination resistance. The gap between a 30-model and a 55-model is not that one writes slightly nicer prose. It is that the 55-model can carry a multi-step agentic task — plan, call tools, recover from errors, finish — where the 30-model stalls, loses the thread, or needs a human to hold its hand at every step. The scores separate exactly where the economic value now lives. A 30 in 2026 is not "a bit behind on benchmarks." It is below the threshold where the current generation of work gets delegated at all.

Now the time-machine view, which is the part the over-time chart makes brutally visual. A score of 30 was frontier-adjacent in early 2025 — roughly eighteen months ago. On that same chart, the leading models of that moment sat in the low-to-mid thirties; a 30 would have been in the conversation. Then the field moved: the frontier crossed 40 by late 2025 and stands at 56 to 61 today. Mistral's current flagship has, in effect, arrived at last year's frontier — on time for a market that no longer exists. And the pace-setters were not only American. The Chinese labs covered the same ground even faster: DeepSeek jumped ten full index points in a single release this summer, and today Kimi K3 (57), Qwen3.8 Max (56), GLM-5.2 (51), and DeepSeek V4 Flash (50) all sit twenty-plus points above Europe's best — several of them open-weight, several of them radically cheaper. Two ecosystems, American and Chinese, both climbed twenty points in the time Europe's champion climbed ten.

Which brings me to the bluntest version of my objection, and it is an opinion I will state as one: nobody chooses to work with a far less intelligent model. Not out of loyalty, not for the flag, not for long. Intelligence is the product. A developer whose agent stalls where a rival's finishes, an enterprise whose workflows need double the supervision, a ministry whose sovereign deployment is eighteen months behind what its own staff use privately at home — they all drift to the more capable option, or they quietly wrap the sovereign model in a foreign one and stop mentioning it in the slide decks. A model at last year's intelligence competes in this year's market, and this year's market includes a 50-point model at three cents a task. Sovereignty can make a buyer tolerate a small gap. It cannot make anyone tolerate working with half the intelligence — and pretending otherwise mistakes what procurement documents can mandate for what people will actually use.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It is not even the cheap option

The obvious defense — fine, it's not the smartest, but it's the efficient European workhorse — does not survive contact with the cost data either, and this surprised even me.

On cost per Intelligence Index task, Mistral Medium 3.5 comes in around $0.46. For that price you are buying intelligence of 30. But Claude 4.5 Haiku delivers the same 30 for about $0.22 — half the price for identical measured intelligence. And DeepSeek's V4 Flash delivers 50 — two-thirds more intelligence — for about $0.03, roughly a fifteenth of Mistral's cost per task. On the intelligence-versus-cost frontier that Artificial Analysis draws, Mistral Medium 3.5 does not sit on the efficient line at all; it sits below and to the right of it — the quadrant where you are paying more for less. It is dominated on price by a cheaper model of equal intelligence and buried on capability by cheaper models of far greater intelligence.

A model that is not the smartest can win on price. A model that is not the cheapest can win on capability. Mistral Medium 3.5, on the independent numbers, is neither the smartest nor the cheapest in its own price band. That is a strategically homeless position, and it is the hardest single fact in this whole piece.

Train It. Tame It. Teach It.: Build Your Personal AI Team and Get Every Model to Work Your Way (The AI Practitioner's Edge)

Train It. Tame It. Teach It.: Build Your Personal AI Team and Get Every Model to Work Your Way (The AI Practitioner's Edge)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The honest case for Mistral — because it exists

I promised myself I would not write a hit piece, so here is the genuine case for the defense, and none of it is charity.

The Intelligence Index measures intelligence. It does not measure the things Mistral actually sells. Open weights are a real product the benchmark cannot see: a model you can run entirely inside your own infrastructure, under your own jurisdiction, is worth more to a European hospital or ministry than four points on a leaderboard, and that is a category where Mistral genuinely leads and Anthropic and OpenAI structurally cannot follow. Sovereignty is a real wedge — French and German defense, EU institutions, and regulated enterprises increasingly require a sovereign option in procurement, and for those buyers Mistral's location and openness are the specification, not the intelligence score. The company is building real infrastructure, a €4 billion data-center plan across France and Sweden, partly nuclear-powered, and ASML's roughly eleven percent stake ties it to the actual heart of Europe's tech-industrial base. And on a fraction of its rivals' capital — around four billion dollars raised against OpenAI's and Anthropic's tens of billions — it produced a lineup that runs, ships, and earns a real and fast-growing revenue. Judged as a three-year-old company on a tenth of the budget, that is not failure. It is remarkable.

All of that is true. And none of it changes the slope.

Why I am hard on them anyway

Here is why I refuse to grade Mistral on the sovereignty curve, even though I care about sovereignty more than almost anything I write about — and precisely because I do not care whether Mistral specifically is the one that delivers it.

Europe is being asked to concentrate its AI independence behind this one company, at a valuation reported around €20 billion — a number everyone acknowledges is a geopolitical premium, a price paid for the idea of a European champion rather than the measured capability of one. Industrial policy, defense procurement, the InvestAI billions, a continent's hope of not depending on foreign intelligence — a great deal of it is being routed toward a single lab whose flagship currently ties a rival's cheapest model and whose trajectory is the flattest on the board. If that is true, then Europe's sovereignty is not being secured by anointing Mistral. It is being narrated, while the actual capability gap widens underneath the story.

This is exactly why anchoring the whole project to one company's fortunes is the mistake. If my loyalty were to Mistral, the flat trajectory would be a tragedy to explain away. Because my loyalty is to European sovereignty and not to any one logo, the flat trajectory is simply information — evidence that the continent may need more shots on goal than one anointed champion: more labs, open-weight efforts, national and academic efforts, whatever bends the curve. Narrating a gap you are not closing, behind a single name, is how you sleepwalk into permanent dependence. It is the most expensive kind of comfortable. The actually pro-sovereignty position is not to cheer the designated champion regardless of the numbers; it is to stare at the numbers precisely because the goal matters more than the mascot. Grading a lab on a curve because it is European — and because it is the European one everyone agreed to back — is the soft bigotry that lets a strategic vulnerability harden into a permanent one while everyone claps.

What would change my mind

This is not a verdict on Mistral forever; it is a reading of a slope, and slopes can bend. What I am watching for is not another model that ties Haiku at a slightly better price. It is a single release — from Mistral or from anyone else flying a European flag — that visibly changes the trajectory: that re-enters the frontier conversation, closes distance instead of holding it, tilts the line on that over-time chart up toward the field instead of away from it. One such release and I will write the opposite of this piece with more relief than I can describe, and I will not care in the slightest whose name is on it.

Until then, the honest read stands, and pretending otherwise would be the least sovereign thing I could do. Europe deserves a real frontier lab. Right now, on the independent numbers, the company it has designated for the role — well-funded, strategically positioned, genuinely open — is not at the frontier and is not currently closing the distance to it. That is not an argument against European sovereignty. It is an argument for pursuing it with clearer eyes, more contenders, and far less loyalty to any single mascot. The goal is the point. The logo is not.


Reality Check and opinion from a builder, founder, and post-labor economist running a local-first inference operation, who wants European AI sovereignty to succeed — through whichever lab delivers it, Mistral or otherwise. All intelligence and cost figures are from Artificial Analysis's independent Intelligence Index (v4.1) and cost-per-task analysis as displayed August 2026; older Claude Opus scores referenced from the same tool are Artificial Analysis estimates pending full evaluation and are flagged as such. Funding, valuation, and infrastructure details from contemporaneous coverage (Bloomberg, TNW, Sifted, and others, mid-2026). This is analysis of publicly benchmarked model performance, not a claim about the company's people or prospects, and not investment advice. Point-in-time as of 6 August 2026.

You May Also Like

From Copilots to Coordinators: Why 2026 Is the Year Agentic AI Hits the Operating Core

By Thorsten Meyer | ThorstenMeyerAI.com | February 2026 Executive Summary 40% of…

Meta’s AI Strategy Evolution: Balancing Open Source Ambitions and Proprietary Advances

Meta, historically a champion of open-source AI development, has undergone a significant…

Meta appoints Shengjia Zhao as chief scientist at Superintelligence Labs

Meta Platforms has made a major strategic move in artificial intelligence by…

Against Sovereignty: The Strongest Case for Just Using the Best Model

This publication has spent five weeks arguing one thing. Own the model,…