🔍 Read the full analysis: Can Mistral Large 4 Run Your Agents? Its Strengths And Limits on ThorstenMeyerAI.com
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index, a sharp improvement over the company’s previous models but below leading US and Chinese systems. Its long context and multimodal input may suit some tasks, but benchmark performance, reported verbosity, pricing and a hands-on account of hallucinations leave its value for autonomous agents uncertain.
Mistral AI has released Mistral Large 4 as a research preview, presenting a major performance jump for the French company while independent benchmark data places it behind leading US models and several Chinese systems. The result matters to teams considering the model for multi-step agent workflows, where capability, reliability, speed and the cost of repeated model calls all affect whether a task completes successfully.
On Artificial Analysis Intelligence Index v4.3.2, Large 4 scored 38.4. The source report says the index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. US frontier models at the top of the table scored above 52, while Chinese models GLM-5.3, Kimi K3, GLM-5.3-Flash and DeepSeek V4.1 Flash also exceeded Large 4. These scores measure performance across the index’s tasks; they do not establish how the model will perform in every customer’s agent system.
The release is a substantial improvement over Mistral’s earlier results in the same index: the report gives Large 3 a score of 9 and Medium 3.5 a score of 14. Large 4 has one trillion total parameters, 49 billion active parameters, text and image input, text output, and a 512,000-token context window. Mistral says reinforcement learning is still underway, so its scores may change. The model is available through the company’s API as a research preview; the report says Mistral has promised model weights for the end of October.
Pricing listed in the source is $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. A 50% introductory discount applies for the first two weeks, according to the report. The source calculates a cost of $1.13 per Artificial Analysis Index task at standard pricing. It gives lower task-cost figures for GLM-5.3-Flash ($0.25) and DeepSeek V4.1 Flash ($0.27), both of which scored higher in the index. These are benchmark task-cost comparisons, not a guarantee of costs for a particular production workflow.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
Agent Workloads Face Three Trade-offs
The benchmark matters because the index includes multi-step work and tool-oriented evaluations, rather than measuring only short-form conversation. In an agent workflow, an early error can shape later actions: a false assumption may affect subsequent research, code changes or decisions. The source author argues that capability gaps can compound across long runs. The benchmark score alone, however, does not quantify the probability that a given agent task will fail.
Cost is another practical concern. If a model produces more output tokens or needs repeated calls to finish work, token prices can understate the total expense. The source report says Large 4 generated 200 million output tokens across the index, compared with a median of 81 million for comparable models. That is a raw benchmark observation; the report does not provide enough detail here to treat it as a direct forecast of token use or latency in every deployment.
There is also a reliability concern, but the evidence needs careful separation. The source author describes seeing confident hallucinations in hands-on use; that is an attributed personal observation, not an Artificial Analysis finding included in the score. The report cites other models’ AA-Omniscience results as evidence that hallucination behavior can differ, but no Large 4 result for that measure is provided. Buyers should test their own tasks and verification controls before relying on the model to act without review.
AI agent workflow automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Model Preview to Deployment
Large 4’s result is both a gain for Mistral and a reminder of the distance to the highest-scoring systems. The source characterizes the move from Large 3’s score of 9 to 38.4 as an unusually large step for a European lab. Yet in the index table supplied, US models score above 52, and multiple Chinese models score above Large 4. The report also notes that Mistral’s release comparisons included earlier Chinese models that Large 4 beats, while newer versions in the table score higher.
Availability remains part of the evaluation. At the time described, Large 4 was a proprietary API preview, with weights promised for the end of October and the licence unpublished. Until weights and licence terms are available, developers cannot assess the release as an open-weight option on those terms. Mistral’s claim to be the most intelligent model outside the US and China is framed in the source through the Artificial Analysis results; the source argues that the comparison set is narrow. That geographic framing is not the same as ranking among the world’s leading models.
Preview Leaves Key Questions Open
The results are a snapshot of a research-preview model, and Mistral says reinforcement learning is continuing. The source does not give a date for when that process will finish or provide updated scores. It also does not identify a published licence for the promised weights, which leaves the future terms of local or modified use unclear.
The source offers no Large 4 result for AA-Omniscience, so its author’s report of confident hallucinations cannot be compared with that benchmark measure here. Nor does the index establish performance on a reader’s specific agent, with its own tools, prompts, safeguards and retry logic. Actual costs, latency and reliability will depend on those implementation details and the volume of input and generated output.
Watch for Weights and Retesting
The next stated milestone is Mistral’s promised release of model weights by the end of October; the source does not specify a year in the supplied material. Publication of the licence will be necessary for developers to judge what uses and distribution it permits. Further Artificial Analysis results could also change as Mistral continues reinforcement learning.
For teams evaluating agents, the practical next step is a controlled trial using representative tasks, with checks for factual errors, tool-use mistakes, completion rates, output volume and total cost. The available evidence supports a clear conclusion for now: Large 4 is a major step forward for Mistral, but the preview benchmark does not by itself establish it as a cost-effective or dependable choice for autonomous workflows.
Key Questions
What is Mistral Large 4?
Mistral Large 4 is Mistral AI’s one-trillion-parameter model, with 49 billion active parameters, text and image input, text output and a 512,000-token context window. The source describes it as a research preview available through Mistral’s API.
How did Large 4 score on the benchmark?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. In the source’s table, leading US models scored above 52 and several Chinese models also scored higher. The score is an index result, not a prediction for every application.
Can it run autonomous agents?
The supplied data does not establish that it is suitable for every agent workload. Its index includes agent-related tasks, but the score, reported output volume and an author’s account of hallucinations are reasons to test it on the intended workflow before deployment.
How much does it cost?
The source lists standard API prices of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a two-week 50% discount and a benchmark task cost of $1.13; actual costs depend on usage and task design.
When will its weights be available?
Mistral has promised the weights for the end of October, according to the source. The supplied material does not specify a year, and it says the licence had not been published at the time of reporting.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
