By Thorsten Meyer
I want Europe to have a sovereign frontier lab. I do not care whether it is Mistral.
That distinction matters, because everything that follows is critical, and it would be easy to mistake it for disappointment in a favourite team. It isn’t. I have argued for a year that European AI sovereignty is real, that a continent which only consumes intelligence made elsewhere has outsourced the layer that matters, and that Europe needs a frontier-grade lab of its own — a real alternative to two American frontier labs and one Chinese bloc, something governments and regulated enterprises can build on instead of routing their sovereignty through a foreign API. Mistral is the company currently cast in that role: the one carrying the European-champion narrative, the default name in every sovereignty deck.
So I went looking, on the independent numbers rather than the launch decks, for evidence that the champion Europe has anointed is actually at the frontier. And the honest finding is the one that should worry anyone who wants European sovereignty to be real regardless of whose logo delivers it: it isn’t — and the gap is widening, not closing.
I want Europe to have a sovereign frontier lab. I don’t care whether it’s Mistral. So I went looking on the independent benchmarks for evidence the anointed champion is at the frontier. The honest finding should worry anyone who wants EU sovereignty to be real: it isn’t, and the gap is widening.
▲ Opinion · loyal to the goal, not the mascotArtificial Analysis Intelligence Index (v4.1) — the independent composite of nine evals including agentic coding, tool use, and reasoning. Mistral’s strongest current model against the field.
frontier
frontier
old, superseded
their current best
a rival’s cheapest
A snapshot could be a bad quarter. The trajectory is the structural finding: on Artificial Analysis’s intelligence-over-time chart, Mistral’s line is the flattest of any major lab.
The obvious defense — “not the smartest, but the efficient workhorse” — doesn’t survive the cost data. Cost per Intelligence Index task, at each model’s measured intelligence.
The Index measures intelligence. It doesn’t measure what Mistral actually sells. Both columns are true.
- Open weights the benchmark can’t see — run it in your own jurisdiction, a real product Anthropic and OpenAI structurally can’t match
- Sovereignty is the spec for EU defense, institutions, regulated buyers — not the score
- Real infrastructure: €4B data centers, France + Sweden, partly nuclear; ASML’s ~11% stake
- On ~1/10 the capital of US rivals — remarkable for a 3-year-old
- Europe is concentrating its AI independence behind one lab, at a ~€20B geopolitical premium
- If the anointed option ties a rival’s cheapest model, sovereignty is being narrated, not secured
- Loyalty to the goal not the logo turns a flat line from tragedy into information: Europe needs more shots on goal
- The actually pro-sovereignty move is to stare at the numbers — the goal matters more than the mascot
which is an argument for more contenders and less loyalty to any one mascot. The goal is the point.
The comparison that should not be possible
Here is the single fact that reframes the whole story. On Artificial Analysis's Intelligence Index — the independent composite that runs nine evaluations including agentic coding, tool use, and reasoning — Mistral's strongest current model, Mistral Medium 3.5, scores 30.
Now place that against the field. The current frontier sits at 56 to 61: Claude Opus 5 at 61, Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57. Mistral's best is roughly half the frontier score. That alone would be a gap. It is what Mistral ties at 30 that turns a gap into an indictment: Claude 4.5 Haiku — Anthropic's cheapest, smallest model, the one you reach for when you explicitly do not need intelligence — also scores 30. Europe's flagship frontier model is level with a competitor's budget tier.
Push one step further, into the comparison a reader can build for themselves on the same tool, and it gets worse. Anthropic's older, superseded Opus models — Claude 4.1 Opus, a model already replaced by 4.5, then 4.8, then 5 — sit at an estimated 34, above Mistral's current best. An honest caveat: those older-Claude figures are Artificial Analysis estimates pending full evaluation, marked as such, so I lean on the measured comparison rather than the estimated one. But the measured comparison is the damning one. Mistral's current flagship does not merely lose to the frontier. It ties the cheapest current model from one rival and trails the year-old flagship of another.
Mistral's best model scores 30 on the Artificial Analysis Intelligence Index. Abstract number — until you ask two questions: what does the index measure, and what year does a 30 belong to?
Nine evaluations, weighted toward what people actually pay models to do in 2026: agentic knowledge work (GDPval), terminal & coding tasks, tool use, long-context reasoning, hallucination resistance.
- Plans, calls tools, recovers from errors, finishes
- Holds a multi-step agentic workflow end to end
- Gets work delegated, not just assisted
- Stalls or loses the thread mid-task
- Needs human supervision at every step
- Below the threshold where 2026's work gets delegated at all
Same index, same chart, eighteen months apart. Mistral's flagship has arrived at last year's frontier, on time for a market that no longer exists.
The pace-setters weren't only American. Two ecosystems climbed twenty points in the time Europe's champion climbed ten.
Stated as the opinion it is — and the reason the gap can't be waved away with procurement mandates.
European AI frontier research labs
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The number that actually matters is the slope, not the score
A single snapshot could be a bad quarter. What makes this a structural finding rather than a bad-week story is the trajectory, and the trajectory is the part that genuinely worries me.

Artificial Analysis plots frontier intelligence over time, every major lab as a climbing line. Anthropic, OpenAI, Google, the Chinese labs — they all march up and to the right, from single digits in 2023 to the high fifties and low sixties now. Mistral's line is on the same chart, and it is the flattest of any major lab. It crawled from near zero to about thirty while the field climbed past fifty-five. The distance between Mistral and the frontier is not constant — it is growing, release over release, because everyone else is climbing faster than Mistral is.
That is the thing a valuation cannot paper over. A lab that is a fixed distance behind is a lab you can imagine catching up. A lab whose gap widens with every generation is on a different curve, and different curves do not converge on their own. The snapshot says Mistral is behind. The slope says Mistral is falling behind. Those are different diagnoses, and the second is the serious one.

Media Company in a Box: Build an Independent Media Business with AI: Answer Engine Optimization, Podcasting, and Nine Revenue Streams for the Sovereign Creator
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What a 30 actually means in August 2026
Because "30 on an index" is abstract, let me make it concrete in the two ways that matter: what the number measures, and what year it belongs to.
The Intelligence Index is not trivia. Its nine evaluations lean heavily on the things people actually pay models to do in 2026 — agentic knowledge work, terminal and coding tasks, tool use, long-context reasoning, hallucination resistance. The gap between a 30-model and a 55-model is not that one writes slightly nicer prose. It is that the 55-model can carry a multi-step agentic task — plan, call tools, recover from errors, finish — where the 30-model stalls, loses the thread, or needs a human to hold its hand at every step. The scores separate exactly where the economic value now lives. A 30 in 2026 is not "a bit behind on benchmarks." It is below the threshold where the current generation of work gets delegated at all.
Now the time-machine view, which is the part the over-time chart makes brutally visual. A score of 30 was frontier-adjacent in early 2025 — roughly eighteen months ago. On that same chart, the leading models of that moment sat in the low-to-mid thirties; a 30 would have been in the conversation. Then the field moved: the frontier crossed 40 by late 2025 and stands at 56 to 61 today. Mistral's current flagship has, in effect, arrived at last year's frontier — on time for a market that no longer exists. And the pace-setters were not only American. The Chinese labs covered the same ground even faster: DeepSeek jumped ten full index points in a single release this summer, and today Kimi K3 (57), Qwen3.8 Max (56), GLM-5.2 (51), and DeepSeek V4 Flash (50) all sit twenty-plus points above Europe's best — several of them open-weight, several of them radically cheaper. Two ecosystems, American and Chinese, both climbed twenty points in the time Europe's champion climbed ten.
Which brings me to the bluntest version of my objection, and it is an opinion I will state as one: nobody chooses to work with a far less intelligent model. Not out of loyalty, not for the flag, not for long. Intelligence is the product. A developer whose agent stalls where a rival's finishes, an enterprise whose workflows need double the supervision, a ministry whose sovereign deployment is eighteen months behind what its own staff use privately at home — they all drift to the more capable option, or they quietly wrap the sovereign model in a foreign one and stop mentioning it in the slide decks. A model at last year's intelligence competes in this year's market, and this year's market includes a 50-point model at three cents a task. Sovereignty can make a buyer tolerate a small gap. It cannot make anyone tolerate working with half the intelligence — and pretending otherwise mistakes what procurement documents can mandate for what people will actually use.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
It is not even the cheap option
The obvious defense — fine, it's not the smartest, but it's the efficient European workhorse — does not survive contact with the cost data either, and this surprised even me.
On cost per Intelligence Index task, Mistral Medium 3.5 comes in around $0.46. For that price you are buying intelligence of 30. But Claude 4.5 Haiku delivers the same 30 for about $0.22 — half the price for identical measured intelligence. And DeepSeek's V4 Flash delivers 50 — two-thirds more intelligence — for about $0.03, roughly a fifteenth of Mistral's cost per task. On the intelligence-versus-cost frontier that Artificial Analysis draws, Mistral Medium 3.5 does not sit on the efficient line at all; it sits below and to the right of it — the quadrant where you are paying more for less. It is dominated on price by a cheaper model of equal intelligence and buried on capability by cheaper models of far greater intelligence.
A model that is not the smartest can win on price. A model that is not the cheapest can win on capability. Mistral Medium 3.5, on the independent numbers, is neither the smartest nor the cheapest in its own price band. That is a strategically homeless position, and it is the hardest single fact in this whole piece.

Train It. Tame It. Teach It.: Build Your Personal AI Team and Get Every Model to Work Your Way (The AI Practitioner's Edge)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The honest case for Mistral — because it exists
I promised myself I would not write a hit piece, so here is the genuine case for the defense, and none of it is charity.
The Intelligence Index measures intelligence. It does not measure the things Mistral actually sells. Open weights are a real product the benchmark cannot see: a model you can run entirely inside your own infrastructure, under your own jurisdiction, is worth more to a European hospital or ministry than four points on a leaderboard, and that is a category where Mistral genuinely leads and Anthropic and OpenAI structurally cannot follow. Sovereignty is a real wedge — French and German defense, EU institutions, and regulated enterprises increasingly require a sovereign option in procurement, and for those buyers Mistral's location and openness are the specification, not the intelligence score. The company is building real infrastructure, a €4 billion data-center plan across France and Sweden, partly nuclear-powered, and ASML's roughly eleven percent stake ties it to the actual heart of Europe's tech-industrial base. And on a fraction of its rivals' capital — around four billion dollars raised against OpenAI's and Anthropic's tens of billions — it produced a lineup that runs, ships, and earns a real and fast-growing revenue. Judged as a three-year-old company on a tenth of the budget, that is not failure. It is remarkable.
All of that is true. And none of it changes the slope.
Why I am hard on them anyway
Here is why I refuse to grade Mistral on the sovereignty curve, even though I care about sovereignty more than almost anything I write about — and precisely because I do not care whether Mistral specifically is the one that delivers it.
Europe is being asked to concentrate its AI independence behind this one company, at a valuation reported around €20 billion — a number everyone acknowledges is a geopolitical premium, a price paid for the idea of a European champion rather than the measured capability of one. Industrial policy, defense procurement, the InvestAI billions, a continent's hope of not depending on foreign intelligence — a great deal of it is being routed toward a single lab whose flagship currently ties a rival's cheapest model and whose trajectory is the flattest on the board. If that is true, then Europe's sovereignty is not being secured by anointing Mistral. It is being narrated, while the actual capability gap widens underneath the story.
This is exactly why anchoring the whole project to one company's fortunes is the mistake. If my loyalty were to Mistral, the flat trajectory would be a tragedy to explain away. Because my loyalty is to European sovereignty and not to any one logo, the flat trajectory is simply information — evidence that the continent may need more shots on goal than one anointed champion: more labs, open-weight efforts, national and academic efforts, whatever bends the curve. Narrating a gap you are not closing, behind a single name, is how you sleepwalk into permanent dependence. It is the most expensive kind of comfortable. The actually pro-sovereignty position is not to cheer the designated champion regardless of the numbers; it is to stare at the numbers precisely because the goal matters more than the mascot. Grading a lab on a curve because it is European — and because it is the European one everyone agreed to back — is the soft bigotry that lets a strategic vulnerability harden into a permanent one while everyone claps.
What would change my mind
This is not a verdict on Mistral forever; it is a reading of a slope, and slopes can bend. What I am watching for is not another model that ties Haiku at a slightly better price. It is a single release — from Mistral or from anyone else flying a European flag — that visibly changes the trajectory: that re-enters the frontier conversation, closes distance instead of holding it, tilts the line on that over-time chart up toward the field instead of away from it. One such release and I will write the opposite of this piece with more relief than I can describe, and I will not care in the slightest whose name is on it.
Until then, the honest read stands, and pretending otherwise would be the least sovereign thing I could do. Europe deserves a real frontier lab. Right now, on the independent numbers, the company it has designated for the role — well-funded, strategically positioned, genuinely open — is not at the frontier and is not currently closing the distance to it. That is not an argument against European sovereignty. It is an argument for pursuing it with clearer eyes, more contenders, and far less loyalty to any single mascot. The goal is the point. The logo is not.
Reality Check and opinion from a builder, founder, and post-labor economist running a local-first inference operation, who wants European AI sovereignty to succeed — through whichever lab delivers it, Mistral or otherwise. All intelligence and cost figures are from Artificial Analysis's independent Intelligence Index (v4.1) and cost-per-task analysis as displayed August 2026; older Claude Opus scores referenced from the same tool are Artificial Analysis estimates pending full evaluation and are flagged as such. Funding, valuation, and infrastructure details from contemporaneous coverage (Bloomberg, TNW, Sifted, and others, mid-2026). This is analysis of publicly benchmarked model performance, not a claim about the company's people or prospects, and not investment advice. Point-in-time as of 6 August 2026.