AIThis post was created with the assistance of artificial intelligence (AI).

Mistral released Mistral Large 4 yesterday, and the headline from Artificial Analysis is the one Paris wanted: France is home to the most intelligent model from outside the United States and China.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That is true. It is also the most generous possible framing of the result, and a buyer who stops reading there will make the wrong decision.

Here is the less flattering version, using the same independent data. Mistral Large 4 scores 38 on the Artificial Analysis Intelligence Index. Every current flagship from a major US or Chinese lab scores higher, the best by more than 19 points. It costs more than four times as much per task as Chinese open models that outscore it. It is two and a half times as verbose as the median model in its class. And in hands-on use I’ve seen something most of last month’s releases had stopped doing: confident hallucination.

Mistral wants to be a frontier lab. Large 4 shows it has closed a lot of ground — and that it is still not one.

Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

What was released

  • 1 trillion parameters, 49 billion active, natively multimodal (text and image in, text out), 512K context.
  • Research Public Preview on Mistral’s API; weights promised for the end of October. Until then it is a proprietary model, and the licence is unpublished.
  • Pricing: $1.36 / $4.18 per million input/output tokens, $0.14 cached input, with 50% off for the first two weeks.
  • Mistral says reinforcement learning is still running, so scores may move.

Credit where it’s due: this is an enormous jump. On the same Index version, Mistral Large 3 scored 9 and Medium 3.5 scored 14. Going from 9 to 38 in one release is the biggest step any European lab has taken. It also shows how far behind Mistral had fallen.

Amazon

AI language model comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where it actually ranks

All figures are from the Artificial Analysis Intelligence Index v4.3.2, the current version, so they compare like for like.

ModelLab · countryIndex
Claude Opus 5.5Anthropic · US57.6
Claude Sonnet 5.5Anthropic · US56.0
Claude Fable 5.1Anthropic · US53.4
GPT-6 AstraOpenAI · US52.7
Gemini 4 ArgonGoogle · US52.6
GPT-6.1 SolOpenAI · US51.8
GLM-5.3Z.ai · China44.8
Kimi K3Moonshot · China43.6
GLM-5.3-FlashZ.ai · China41.8
DeepSeek V4.1 FlashDeepSeek · China39.5
Mistral Large 4 (Preview)Mistral · France38.4
GPT-6 LunaOpenAI · US (small model)~38
DeepSeek V4 Pro 0813DeepSeek · China36.0
GLM-5.2Z.ai · China33.7

Read the table honestly and three things follow.

Against the US frontier, it is not in the conversation. Large 4 reaches about two-thirds of Opus 5.5’s score. The gap to the leader, 19 points, is larger than the entire spread among the six US frontier models. Large 4 sits level with GPT-6 Luna, OpenAI’s small model. Mistral’s flagship is matching an American lab’s budget tier.

Against China, it is eighth among open models. Once its weights ship, Large 4 would rank eighth among open-weights models — behind seven Chinese ones, including both GLM-5.3 variants, Kimi K3 and DeepSeek V4.1 Flash. It beats GLM-5.2 and DeepSeek V4 Pro, the comparisons Mistral chose for its launch. It loses to their successors.

Against Canada, the comparison barely applies. Cohere doesn’t compete at this tier. Its enterprise flagship is built for retrieval and tool use, not frontier reasoning. On AA-Omniscience it is reported to have a very low hallucination rate (around 14%) at very low accuracy (around 9%), because it declines most questions it can’t verify. Calibrated, but not capable enough to be the comparison. The “Western alternative outside the US” race is a field of one, and Large 4 wins it by default.

That last point is the uncomfortable truth behind the headline. “Most intelligent model outside the US and China” is accurate because almost nobody else is trying. The competitive set in that sentence is South Korea and the UAE, not the labs that matter.

Amazon

best AI models for enterprise use

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why I wouldn’t run agents on it

The Intelligence Index is now dominated by agentic work: AA-Briefcase (agentic knowledge work), GDPval-AA (real-world work tasks), AutomationBench (SaaS workflows) and Terminal-Bench 4.0 (agentic coding). A score of 38 on that Index is a direct statement about agentic capability, not a general-knowledge quiz.

For long-running or multi-step work, three things in the data compound:

The capability gap compounds over steps. A model 19 points behind the frontier doesn’t fail 19% more often on a ten-step task — errors multiply across steps. On short chat that’s tolerable. On a two-hour agent run it’s the difference between finished and abandoned.

It is extremely verbose. Large 4 generated 200 million output tokens to complete the Index, against a median of 81 million for comparable models. On agentic work, verbosity is cost and latency on every step, and the per-token price hides it.

Hallucination is back. This is my own observation, not an Artificial Analysis figure: in hands-on testing I saw Large 4 assert things confidently that weren’t true, a failure mode that had largely receded in last month’s US frontier releases. Gemini 4 Argon, for instance, posts a 15% AA-Omniscience hallucination rate, the lowest of any model above 45 on the Index. In fairness, Chinese open models are no better here — Kimi K3 sits at 51% and DeepSeek V4 Pro at 94%. But “no worse than the Chinese open models” is not a reason to put a model in charge of a workflow. In an agent, a confident fabrication isn’t one wrong answer; it’s a wrong premise every later step builds on.

Amazon

AI model hallucination detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The cost problem is worse than the intelligence problem

Here is the number I’d find hardest to defend in a procurement meeting.

At standard pricing, Large 4 costs $1.13 per Intelligence Index task. Compare:

  • GLM-5.3-Flash: $0.25 per task — and scores 41.8.
  • DeepSeek V4.1 Flash: $0.27 per task — and scores 39.5.
  • Gemini 4 Argon: ~$1.99 per task — and scores 52.6.

So for less than a quarter of the price, two Chinese open models outscore it. And for less than twice the price, a US frontier model delivers 14 more Index points. Even at the two-week launch discount ($0.57 per task), Large 4 costs more than twice what the stronger Chinese models cost.

Mistral’s per-token prices look competitive — $4.18 per million output tokens is well under the median. The verbosity cancels that out. Cheap tokens multiplied by two and a half times as many tokens is not a cheap model.

Amazon

large language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What it’s genuinely good at

Fairness requires the list, and it’s real:

  • Cyber defence. It scores 50 on the AA Cyber Index, level with GLM-5.3-Flash and ahead of Kimi K3 and DeepSeek V4.1 Flash. Its 82% on CyberGym-E2E beats GPT-6 Luna (78%). Once the weights ship it should be a top-three open model on cyber.
  • Documents and images. It scores 19% on GDP.pdf, an 18-point gain on Large 3, and the API now accepts 100 images per request (up from 8).
  • Speed. 116 tokens per second with a 1.46-second time to first token, both well above the median.
  • Jurisdiction. A French parent, EU hosting, and — if the licence matches Large 3’s — permissive open weights at the end of October.

None of that changes the verdict on agentic or long-horizon work. It does define the narrow cases where Large 4 is the right answer.

Who should actually use it

If you’re legally bound — defence, classified work, DORA-regulated finance, national health data — Large 4 is now the best European option by a wide margin. It’s a real upgrade over Large 3, and for you the alternative isn’t a better model, it’s no deployment. Wait for the weights, check the licence, and pilot it on cyber and document workloads, where it’s strongest.

If you’re not bound, don’t choose it for agentic or long-running work. For the same budget you can get a model that is meaningfully more capable (a US frontier model) or one that’s both more capable and four times cheaper (GLM-5.3-Flash). Large 4 is neither the best model nor the best value. It is the best European model, and that matters only if being European is a requirement rather than a preference.

The take

Mistral says it has “essentially closed the gap.” The data says it has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on.

That is genuine progress, and Europe should be glad of it: going from 9 to 38 in one release matters, and a strong cyber model under EU jurisdiction is useful. But a frontier lab is defined by being at the frontier, and on every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 isn’t. Nineteen points behind the leader, eighth among open models, four times the cost of stronger alternatives, and hallucinating in ways its peers have largely stopped.

Use it if you have to. Don’t use it because the headline said “most intelligent outside the US and China.” That sentence is true mainly because almost nobody else outside those two countries is competing.


Sources: Artificial Analysis, “Mistral has released Mistral Large 4, making France home to the most intelligent model outside the US and China” (6 October 2026) and the Mistral Large 4 Preview model and Mistral provider pages — Intelligence Index 38 (38.4), v4.3.2 composition, 1T/49B-active architecture, 512K context, Research Public Preview, end-of-October weights, $1.36/$4.18/$0.14 pricing and 50% two-week discount, $1.13 and $0.57 cost per task, 200M vs 81M median output tokens, 116 tokens/s and 1.46s TTFT, Cyber Index 50, CyberGym-E2E-AA 82%, GDP.pdf 19% (+18 vs Large 3), 100 images per request, and the comparison figures for GLM-5.3-Flash ($0.25), DeepSeek V4.1 Flash (39, $0.27), GPT-6 Luna (38) and MiMo-V2.6-Pro; Artificial Analysis comparison data for GLM-5.3 (44.8), Kimi K3 max (43.6) and low (34.5), GLM-5.3-Flash (41.8), DeepSeek V4 Pro 0813 (36.0); Trending Topics’ independent-ranking analysis (eighth among open models behind seven Chinese, GLM-5.2 33.7, Large 3 at 9, Medium 3.5 at 14, ~two-thirds of Opus 5.5, RL still underway); v4.3.2 frontier scores (Opus 5.5 57.6, Sonnet 5.5 56.0, Fable 5.1 53.4, Astra 52.7, Gemini 4 Argon 52.6, GPT-6.1 Sol 51.8) via Artificial Analysis-derived reporting; Gemini 4 Argon’s 15% hallucination rate and ~$1.99 cost per task via OfficeChai and FelloAI; Kimi K3’s 51% and DeepSeek V4 Pro’s 94% AA-Omniscience hallucination rates via Artificial Analysis-derived reporting; Cohere’s AA-Omniscience profile as reported by Suprmind. Mistral’s own statements per its launch post. Mistral Large 4’s AA-Omniscience result is not published in text; the hallucination observation is the author’s own hands-on testing, not an Artificial Analysis measurement. Preview scores may change as training continues. Not investment advice. Analysis and framing are the author’s.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Internet Is Forking: How Coinbase, Cloudflare, Stripe, and OpenAI Are Building a Parallel Web for AI Agents

AIThis post was created with the assistance of artificial intelligence (AI).The $285…

Limits, Levers, and a Roadmap: What It Will Take for Video Models to Become Vision Foundation Models

AIThis post was created with the assistance of artificial intelligence (AI).Claim under…

ASI-ARCH: A New Era of Autonomous AI Research

AIThis post was created with the assistance of artificial intelligence (AI).The AI…

The Free-Download Question: When Running Your Own Model Actually Beats Paying

AIThis post was created with the assistance of artificial intelligence (AI).The Mistral…