🔍 Read the full analysis: Mistral Large 4’S Race To Keep Up With The AI Frontier on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral launched a public API preview of Large 4 on October 6, 2026, describing it as a one-trillion-parameter mixture-of-experts model. Artificial Analysis gave the preview an Intelligence Index score of 38, below leading US models and several Chinese competitors; the weights are scheduled for release later in October. A reviewer advised against relying on it for demanding agentic work, but that assessment reflects personal experience, not a controlled reliability study.
Mistral launched a public API preview of Mistral Large 4 on October 6, introducing its largest model to date as the French company competes for a place among leading AI developers. In results available the next day, Artificial Analysis scored the preview 38 on its Intelligence Index, below several US and Chinese models; Mistral says the model’s weights are scheduled for release later in October.
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. It accepts text and images. The company said it trained the model on its own infrastructure in Europe and continues to improve it. As of October 7, access is through a preview API: the weights are not yet publicly downloadable.
Artificial Analysis’ dated comparison places Large 4 Preview at 38 on its Intelligence Index. The same snapshot gives Claude Opus 5.5, at maximum reasoning effort with default fallback, 58; Google Gemini 4 Argon, at high effort, 53; and OpenAI GPT-6.1 Sol, at maximum effort, 52. Chinese models GLM-5.3 and Kimi K3 score 45 and 44, respectively, while DeepSeek V4.1 Flash scores 39 at maximum effort. GPT-6 Luna also scores 38 at maximum effort.
These are benchmark index points, not percentages or direct predictions of success on a specific task. The comparison uses different named reasoning settings, not identical compute budgets. Artificial Analysis reports a context capacity of roughly 512,000 tokens, but a large context window measures how much input a model can accept, not how reliably it reasons over that material.
Mistral Large 4’s Race to Keep Up With the AI Frontier
Mistral’s one-trillion-parameter model is now in public API preview. Its first benchmark snapshot trails several leading US and Chinese systems, while its promised weight release remains ahead.
What arrived in preview
API access is available; the publicly downloadable weights are not. Mistral says training took place on its own infrastructure in Europe.
Mixture of experts
Mistral describes Large 4 as a one-trillion-parameter system, with 49 billion parameters active for a given input.
Text and images
The model accepts text and image inputs. Its reported context capacity is roughly 512,000 tokens.
Preview first, weights later
As of October 7, access is through a public API preview. Mistral scheduled the weight release for later in October.
The benchmark gap
Artificial Analysis Intelligence Index scores available October 7, 2026. Bars show index points, not percentages.
The comparison uses different named reasoning settings, not identical compute budgets. Large 4 is 20 points behind Claude Opus 5.5, 15 behind Gemini 4 Argon, and 14 behind GPT-6.1 Sol. It is one point below DeepSeek V4.1 Flash and above Cohere Command A+.
Scores inform; task tests decide
A context window measures input capacity. An aggregate score cannot establish reliability on a specific workflow.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
Thorsten Meyer, ThorstenMeyerAI.com. This is a reviewer’s personal experience, not a controlled reliability study or a measured hallucination rate.
One snapshot
Versions and scores can change. The comparison does not use identical reasoning settings or compute budgets.
Errors can compound
Planning, tool use, and checking across multiple steps need evaluation on the team’s actual tasks.
Release terms unknown
No exact date or terms were specified in the source material. The weights remained pending on October 7.
No figures provided
The available information does not establish how Large 4’s cost compares with alternatives.
From API preview to workload tests
Test real tasks
Compare the API on representative coding, research, and professional workflows.
Check the process
Track constraint following, evidence quality, tool error recovery, and human review needs.
Review the weights
Assess the announced release when available, including its terms and practical deployment fit.
Update the picture
New versions and benchmark results may shift the current comparison.
The Benchmark Gap for Developers
The preview’s score gives developers an early, comparable data point, but it does not settle which model will work best in a particular product. For teams choosing a system for long-running agentic tasks, the reviewer’s concern is that planning, tool use and checking results can compound errors over multiple steps. A model’s aggregate benchmark score can inform that decision, but it cannot establish how often it will make unsupported claims or complete a specific workflow correctly.
The gap is relevant to Mistral’s competitive position. In this snapshot, Large 4 is 20 index points below Claude Opus 5.5, 15 below Gemini 4 Argon and 14 below GPT-6.1 Sol. It is also behind the listed GLM and Kimi models. Those figures describe differences on the index, not percentages of capability. DeepSeek V4.1 Flash is one point higher, while Mistral scores above Cohere Command A+, which has 13. That counterexample means it would be inaccurate to say every competing major developer is ahead.
Mistral’s release also has significance beyond a single ranking: the company says it trained Large 4 on European infrastructure. That is a development in Europe’s AI capacity, even if it does not show that the preview matches the strongest available models. Developers will need to weigh that consideration against measured task performance, cost, access arrangements and the results of their own tests.
As an affiliate, we earn on qualifying purchases.
From API Preview to Weights
The October 6 announcement is a preview launch, not a completed open-weight release. Mistral has made the model available through an API and says its weights will follow later in October. Until that release, developers cannot assess the publicly downloadable weights described in the announcement, and the available benchmark results apply to the preview evaluated at that time.
The score comparison is a dated snapshot from October 7, 2026, based on Artificial Analysis’ Intelligence Index. The developer locations in the comparison identify the organizations, not where any individual API request is processed. Scores and model versions can change, and results from different reasoning settings should not be treated as a controlled test using equal compute.
The source reviewer also reported encountering hallucinations while using the preview and said this reduced confidence in assigning it longer tasks. That is personal experience, not a controlled comparative study; it does not establish how often Large 4 hallucinates or prove that other models do not. The reviewer also noted that Mistral advertises strengths in agentic coding and professional tasks, claims that would need testing against the relevant workloads.
“The model was trained on Mistral’s own infrastructure in Europe, and the company says it continues to improve it.”
— Mistral
As an affiliate, we earn on qualifying purchases.
Preview Limits and Open Questions
Several points remain unsettled. Mistral has not yet made the model weights available, so the announced open-weight release remains pending as of October 7. The company has said the weights are scheduled for later in the month, but the source material does not give a specific release date or confirm what terms will apply.
The current benchmark comparison cannot show how Large 4 performs across every coding, research or professional workflow. It also does not provide an apples-to-apples comparison under identical reasoning and compute settings. The reviewer’s hallucination observations are anecdotal; no controlled comparative rate is provided. The source excerpt does not give cost figures, so claims about whether Large 4 is less expensive or more expensive than alternatives cannot be verified here.
As an affiliate, we earn on qualifying purchases.
The Weights and Workload Tests
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October. Until then, developers evaluating the preview can test the API against their own tasks rather than infer reliability from the Intelligence Index alone. A useful evaluation would track whether the model follows constraints, supports conclusions with evidence, recovers from tool errors and completes multi-step work without excessive human checking.
Further benchmark results, updated model versions and the weight release could change the current picture. For now, the confirmed development is an API preview with a dated benchmark score—not a final assessment of the model’s eventual performance or a demonstrated replacement for higher-scoring systems.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral announced a public API preview of Mistral Large 4 on October 6, 2026. It describes the model as a one-trillion-parameter mixture-of-experts system with 49 billion active parameters that accepts text and images.
Are Mistral Large 4’s weights available now?
No. As of October 7, 2026, the preview is available through an API. Mistral says the weights are scheduled for release later in October, but the source does not specify an exact date or release terms.
How did the preview score against competitors?
Artificial Analysis gave it an Intelligence Index score of 38 in results available October 7. That is below the listed scores for Claude Opus 5.5, Gemini 4 Argon, GPT-6.1 Sol, GLM-5.3 and Kimi K3; it matches GPT-6 Luna’s score and is one point below DeepSeek V4.1 Flash. The models were not all evaluated at identical reasoning settings.
Does the score prove Large 4 is unreliable?
No. An aggregate benchmark score is not a direct reliability rate and does not predict whether the model will fail a particular task. The reviewer advised caution for long agentic work based on the score and personal experience, but did not report a controlled comparative reliability study.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
