Mistral Large 4 May Lead Outside The US And China, But Agents Need More
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 May Lead Outside The US And China, But Agents Need More on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4, released as a research preview, scores 38.4 on the Artificial Analysis Intelligence Index, a large gain over Mistral’s earlier models. It leads the cited field of models from outside the US and China, but trails current US and Chinese leaders, while its reported cost per benchmark task and hands-on hallucination concerns complicate its use for agents.

Mistral AI released Mistral Large 4 as a research public preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index, according to the cited benchmark data. That makes it the highest-scoring model in the report’s comparison set from outside the United States and China, but it remains behind leading models from both countries, a gap that matters for buyers considering it for complex agent workflows.

The model has 1 trillion total parameters, with 49 billion active, and accepts text and images while producing text. Mistral lists a 512,000-token context window. The company has made it available through its API as a research preview; it says model weights are planned for release at the end of October. Until then, the model is proprietary, and the source report says its licence has not been published.

Artificial Analysis Index version 4.3.2 gives Large 4 a score of 38.4. The same report lists Anthropic’s Claude Opus 5.5 at 57.6 and Google’s Gemini 4 Argon at 52.6, while Chinese models GLM-5.3 and Kimi K3 score 44.8 and 43.6. Mistral Large 3 scored 9 and Medium 3.5 scored 14 on the same index version, making Large 4 a marked improvement over the company’s earlier models.

Mistral’s listed API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. The source report says a 50% discount applies for the first two weeks. Using its benchmark-task estimates, the report puts Large 4 at $1.13 per task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models also score higher on the cited index, at 41.8 and 39.5 respectively.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, with independent benchmark data showing a substantial improvement but a continuing gap from leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Trade-Off for Agent Work

The result gives companies outside the US and China another model option, and the jump from Mistral’s earlier scores suggests the French company has improved its benchmark performance substantially. But the index is built around tasks that include agentic knowledge work, software workflows and coding, so its score is relevant to customers evaluating models for multi-step work rather than only short exchanges.

The source report argues that a performance gap can become more consequential over a long sequence of agent actions, where an error early in a task may affect later steps. It also reports that Large 4 used 200 million output tokens across the index, versus a median of 81 million for comparable models. That is a benchmark-specific observation, not a universal measure of how much a customer’s workloads will cost. Still, high output use can raise both latency and expense when a system makes repeated calls.

The report’s author also describes seeing confident false statements in hands-on testing. That is an attributed observation, not a finding from Artificial Analysis. In agent systems, an incorrect answer may be passed into later actions as if it were reliable, making evaluation and human oversight important before deployment. The available material does not establish how frequently that behavior occurs across users or tasks.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Benchmark Frames the Race

The claim that Large 4 is the most intelligent model outside the US and China rests on the comparison in the source report, which draws on the Artificial Analysis Intelligence Index v4.3.2. The phrase describes a geographic comparison, not a ranking above the leading US and Chinese systems. The report notes that the cited models from China score above Large 4, including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash.

The index combines evaluations including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0, which the source characterizes as agentic or real-world work tests. Its scores therefore speak to performance on that particular mix of tasks, rather than every possible use of a language model. Results can also shift: Mistral says reinforcement learning is still underway and that scores may change.

Large 4’s benchmark result arrives alongside a notable change in access planned by the company. At preview launch, customers can use the API, but the weights and licence details are not yet available in the source material. The promised weight release could affect how developers assess deployment options, though the final availability and terms will need confirmation.

“Reinforcement learning is still running, so scores may move.”

— Mistral AI

Amazon

large language model for agents

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Scores and Release Terms

Large 4 is still a preview, and Mistral says reinforcement learning is ongoing. It is not clear from the source material when the benchmark will be rerun or how much the score could change. The report’s cost-per-task figures are estimates tied to its benchmark runs; actual costs will depend on each customer’s prompts, output lengths, caching and usage patterns.

The source material also does not provide the eventual weights release licence, or confirm that the planned end-of-October release will occur on that schedule. The frequency and scope of the reported hallucinations are likewise not established by the hands-on observations described in the report. Buyers will need task-specific testing rather than treating one index score or anecdotal test as a complete assessment.

Amazon

AI text and image processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Further Testing

The next expected milestone is Mistral’s planned release of Large 4’s weights at the end of October, along with the licence terms needed to understand how they may be used. Mistral’s API preview remains the current access route described in the source material. The company may also update benchmark results as its reinforcement-learning work continues.

For prospective users, the practical next step is to test the preview on their own workloads, tracking task completion, factual errors, output volume, latency and total cost. Those results can show whether Large 4’s performance is suitable for a particular workflow and whether its price compares favorably with other models. The broader competitive picture may become clearer once the weights, licence and updated evaluations are available.

Amazon

AI benchmark analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral release?

Mistral released Large 4 as a research public preview through its API. The model takes text and images as input, produces text, and has a listed context window of 512,000 tokens.

How did Mistral Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source report. The same report lists several US and Chinese models above that score.

Is Large 4 available with open weights?

Not yet, according to the source material. Mistral plans to release the weights at the end of October, but the licence is not provided in the report and the planned release remains ahead.

What does the source report say about cost?

The listed API rates are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The report estimates $1.13 per benchmark task, but real-world costs will vary by workload and usage.

Is Large 4 suitable for agent workflows?

The index includes agent-oriented evaluations, but its score is below several cited US and Chinese models. The report’s author also raises concerns about output volume and describes observed hallucinations; those observations do not establish how the model will perform on every customer’s tasks. Testing with oversight is needed before relying on it for consequential workflows.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Editing Matters More Than Generation in AI-Driven Work

Never underestimate the importance of editing in AI-driven work, as it ensures accuracy and integrity—yet, the true value lies in…

Ethics of AI Art: Copyright and Originality Debates

Understanding the evolving ethics of AI art raises complex questions about copyright and originality that you won’t want to miss.

AmenGate: The Moment Before The Scroll

AmenGate introduces a faith-based prayer lock for iPhone, aiming to transform phone interruptions into meaningful moments of prayer, built on system-level security.

A Look At Claude Opus 5.5 From Anthropic

Anthropic has announced Claude Opus 5.5, but model performance, pricing, features and availability could not be verified from the material available.