A Guide To Mistral Large 4’S Strengths Beyond The US And China
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Guide To Mistral Large 4’S Strengths Beyond The US And China on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get privacy and security gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index, a sharp rise from earlier Mistral models but below leading US and Chinese systems. The source material says it is available as a proprietary research preview, with weights promised for the end of October; its price per benchmark task is also higher than that of two Chinese models that scored better.

Mistral released Large 4 as a research preview, and it scored 38.4 on the Artificial Analysis Intelligence Index, making it the highest-ranked model in that index from outside the United States and China, according to the source material. The result marks a large improvement over Mistral’s earlier models, but the score remains below current leading US and Chinese systems, and the source’s cost comparison raises questions for buyers considering it for agentic workloads.

Artificial Analysis Index version 4.3.2 gives Large 4 a score of 38.4. The source material lists US models at the top, led by Claude Opus 5.5 at 57.6, and Chinese models including GLM-5.3 at 44.8, Kimi K3 at 43.6 and GLM-5.3-Flash at 41.8. Large 4 also scores below DeepSeek V4.1 Flash at 39.5. Those comparisons are based on the figures supplied by the source and refer to the index version it names.

The jump from Mistral Large 3’s score of 9 and Medium 3.5’s score of 14 is substantial. Large 4 has one trillion parameters, with 49 billion active, accepts text and images, produces text, and supports a 512,000-token context window. It is available through Mistral’s API as a research preview. The source says Mistral promised to release the model weights by the end of October; until then, the model is proprietary and its licence has not been published.

The listed API price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. The source says Mistral offered a 50% discount for the first two weeks and that reinforcement learning was still underway, so benchmark scores could change. It reports a benchmark-task cost of $1.13 for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, both of which scored higher on the index.

At a glance
reportWhen: Released the day before the source arti…
The developmentMistral released Large 4 as a research preview, with Artificial Analysis ranking it as the highest-scoring model from outside the United States and China while still placing it behind leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost Of Agentic Performance

The result matters because the index includes tests of agentic work and multi-step tasks, rather than measuring only conversational quality. The source identifies AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0 among the components. A lower score does not establish how a model will perform in every company workflow, but it is relevant evidence for teams comparing systems intended to plan, use tools or complete extended tasks.

For buyers, the combination of a lower index score and higher reported task cost deserves attention. The source’s figures put two Chinese models ahead of Large 4 while costing less per index task. It also reports that Large 4 generated 200 million output tokens during the index evaluation, compared with a median of 81 million for comparable models. That is a benchmark observation, not a universal measure of how much a customer’s prompts will cost; actual spending depends on usage and task design.

The source author also describes seeing confident false statements during hands-on testing. That is an attributed personal observation, not an Artificial Analysis score or a reported independent hallucination rate for Large 4. It nonetheless points to a practical concern for deployments: when an agent passes an unsupported claim into later steps, the error can affect the whole workflow. Teams should test factual reliability and cost in their own use cases rather than treating the model’s regional ranking as a deployment recommendation.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How The Ranking Is Framed

The source frames Large 4 as the most intelligent model outside the US and China on Artificial Analysis’s index. That distinction is geographically specific: it does not mean the model leads the overall table. The same supplied results show several US and Chinese models scoring higher. The source also notes that Canada’s Cohere does not compete at the same frontier-reasoning tier in this comparison, describing its enterprise model as focused on retrieval and tool use.

Mistral’s score gain is notable relative to its previous entries: Large 3 scored 9 and Medium 3.5 scored 14 on the same stated index version. But the source cautions that the rise also highlights how far Mistral’s earlier scores sat below the new result. It compares Large 4 with GLM-5.2 and DeepSeek V4 Pro, which it says Mistral selected for its launch, while noting that newer Chinese models in the supplied table score higher.

Large 4’s current access terms also shape how the result should be read. It is a proprietary API preview, not yet a publicly released open-weights model. The source says Mistral plans to publish its weights at the end of October, but provides no licence terms. Until those are available, customers cannot assess the model’s eventual reuse rights from the supplied information.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— Source article author

Amazon

large language model with image support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Scores And Licence Pending

Large 4’s result may change: the source says reinforcement learning is ongoing, but does not provide a timetable for updated scores or later evaluations. The supplied material also does not give a quantified hallucination rate for Large 4. The author’s testing observation should not be treated as a representative measurement of all prompts or deployments.

The planned release of weights is another open point. The source gives the end of October as Mistral’s target, but the material does not establish whether that release occurred or specify the licence. It also does not provide independent verification of the pricing or benchmark-task cost figures beyond attributing the index data to Artificial Analysis and reporting Mistral’s API terms.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch For Weights And Retests

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. Their publication and licence terms will determine whether customers can run or adapt the model outside Mistral’s API. Because the source says training is ongoing, later index results could also clarify whether the preview score changes.

For organizations evaluating the model now, the practical next step is to compare it against alternatives on their own tasks, including accuracy, output volume, latency and total cost. Those tests can show whether Large 4’s capabilities suit a particular workload; the index and the source author’s limited hands-on observations cannot settle that question on their own.

Amazon

AI model cost comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index version 4.3.2, according to the source material.

Does Large 4 lead the overall rankings?

No. The supplied table places several US and Chinese models above it. The source describes it as the highest-scoring model from outside the US and China in that index.

Can customers download its weights now?

The source describes Large 4 as a proprietary API research preview and says Mistral promised weights by the end of October. It does not confirm that the release has happened or provide licence terms.

How does its reported task cost compare with rivals?

The source reports $1.13 per Intelligence Index task for Large 4, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models also scored higher in the supplied table.

Is Large 4’s hallucination rate known?

The source does not provide a quantified hallucination rate for Large 4. It reports the author’s personal observation of confident false statements, which is not an independent benchmark measurement.

Source: ThorstenMeyerAI.com

COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
SenseTime-W’s Profitable Results Highlight AI Industry Momentum

SenseTime-W’s Profitable Results Highlight AI Industry Momentum

SenseTime-W posts interim profit of RMB 607M and 28.2% growth in generative AI revenue, signaling a strategic shift and sector momentum amid competitive pressures.
One Markdown File, Publish-ready For Every Platform

One Markdown File, Publish-ready For Every Platform

A web tool now enables creators to convert a single markdown file into platform-ready formats, streamlining content distribution for independent creators.
The City That Watches Itself: The Living Digital Twin, And The God’s-Eye View We’re Building

The City That Watches Itself: The Living Digital Twin, And The God’s-Eye View We’re Building

Cities are developing dynamic digital twins that combine real-time data and AI to improve planning and monitoring, raising both opportunities and privacy concerns.
When a Content Network Starts Publishing to Itself

When a Content Network Starts Publishing to Itself

A growing trend where content networks start publishing to their own properties, shifting from external distribution to internal ecosystem building—impacting control, engagement, and revenue.