Mistral Large 4: Strong Outside The US And China, Less Suited To Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: Strong Outside The US And China, Less Suited To Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a substantial improvement over Mistral’s earlier models but below current US and several Chinese models. The source reports high token use and a $1.13 cost per benchmark task, while hands-on testing identified confident errors; these factors may limit its suitability for agents.

Mistral released Large 4 as a research preview, and the model scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the benchmark data cited in a report by ThorstenMeyerAI.com. The score marks a sharp improvement over Mistral’s previous models, but remains below current US flagships and several Chinese models, raising questions about its fit for multi-step agent work.

The model is described as having 1 trillion total parameters, with 49 billion active, and supporting text and image inputs with text output. It has a 512,000-token context window. Mistral made it available through its API as a Research Public Preview; the source says the company plans to release model weights at the end of October. The release date, year, and weights licence are not specified in the supplied material.

On the same Artificial Analysis index version, Large 4 scored 38.4, compared with 9 for Mistral Large 3 and 14 for Medium 3.5. The source calls this a major improvement, while noting that the score remains below listed US models such as Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7. Among the Chinese models listed, GLM-5.3 scored 44.8, Kimi K3 scored 43.6, and GLM-5.3-Flash scored 41.8.

The source lists API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. It says Mistral is offering a 50% discount for the first two weeks. Based on Artificial Analysis task-cost figures reported by the source, Large 4 costs $1.13 per benchmark task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models also scored higher on the cited index.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, with benchmark results showing a major step forward for the French lab but a continued gap to leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A Bigger Model Gap for Agents

For companies considering an AI model for automated, multi-step work, a benchmark gap can matter beyond a single answer. If an agent makes an error early in a sequence, later steps may rely on that mistake. The source argues that Large 4’s lower index score, along with its reported token use, could make it less reliable and more expensive for long-running tasks than its headline capability suggests.

Artificial Analysis Index v4.3.2 includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0, according to the source. The index is not a direct guarantee of performance on every company workflow, but its task mix makes the result relevant to buyers evaluating agents rather than ordinary chat. The source reports that Large 4 used 200 million output tokens across the index tasks, against a median of 81 million for comparable models. That is a benchmark observation, not a universal measure of token use in every deployment.

The practical question is not only whether the model is capable, but whether its performance justifies its cost. The reported task-cost comparison favors two lower-priced Chinese models that scored higher. Buyers will need to test models against their own workloads, including accuracy, latency, tool use and total cost, before drawing procurement conclusions.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Large 4 Compares

The launch framing cited in the source described Large 4 as the most intelligent model outside the US and China. The report treats that as a narrow comparison rather than evidence that the model matches the leading systems overall. Its cited table places Large 4 below the US and Chinese models listed, while its earlier Mistral comparison shows how much the company’s benchmark score has improved.

The source says Mistral has not yet released the model weights and that the licence is unpublished. Until the weights are available, Large 4 is a proprietary API model, rather than an open-weights release that developers can independently host. That distinction matters to organisations weighing control and deployment choices alongside model quality and price.

The release is also not presented as a settled benchmark result: the source says Mistral reported that reinforcement learning is still running, so scores may change. The reported figures should therefore be read as a snapshot of a preview model, not necessarily its final performance.

“Large 4 shows it has closed a lot of ground — and that it is still not one.”

— ThorstenMeyerAI.com report author

Amazon

large language model API key

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results Still May Shift

Several details remain unsettled. The source does not specify the calendar date or year for the planned end-of-October weights release, and it says the weights licence has not been published. It is also unclear whether the benchmark score will change after Mistral completes reinforcement learning or how the model will perform on particular customer workflows.

The source author reports seeing confident false assertions in hands-on testing. That is an attributed observation, not an Artificial Analysis benchmark result, and the supplied material does not give a test protocol, sample size or error rate for Large 4. The report also cites hallucination figures for other models, but those figures do not establish how often Large 4 will make errors in real deployments.

Pricing comparisons may also shift with the stated introductory discount, which applies for the first two weeks according to the source. The reported per-task benchmark costs are not the same as a customer’s bill: actual costs depend on prompts, output length, caching and the number of agent steps.

Amazon

AI model token counter tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Testing Ahead

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. Developers will then be able to assess the model’s hosting options and review the licence, if the release proceeds as described. The source provides no confirmed licensing terms or exact release date.

Mistral’s ongoing reinforcement learning may also change the model’s benchmark performance. Buyers evaluating it for agent use can compare updated results with their own task tests, paying attention to error propagation, output-token consumption, latency and total cost rather than index rank alone. Until those details emerge, Large 4’s score and the source author’s hands-on observations offer evidence to consider, but not a definitive verdict on every use case.

Amazon

AI model cost calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is a Mistral model offered as a Research Public Preview through the company’s API. The source describes it as a 1-trillion-parameter multimodal model with 49 billion active parameters and a 512,000-token context window.

How did Large 4 perform on the cited benchmark?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. That is higher than the cited scores for Mistral Large 3 and Medium 3.5, but below the US and several Chinese models listed in the report.

Why does the report question its use for agents?

The report points to Large 4’s lower benchmark score relative to leading models, its reported high output-token use, and the author’s observation of confident errors in hands-on testing. These concerns may matter more in multi-step workflows, where later actions can build on an earlier mistake. The source does not provide a controlled error-rate study for Large 4.

How much does Mistral Large 4 cost?

The source lists API pricing of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It reports a 50% discount for the first two weeks; customers’ actual costs will depend on usage.

When will the model weights be available?

Mistral is reported to have planned a weights release for the end of October, but the source does not give the year or an exact date. It also says the licence has not been published.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

2026 Guide: 6 AI Tools for Better Student Organization Management

Discover the top six AI-powered tools in 2026 that help students manage their organization and workflow efficiently, with insights on features, costs, and integration.

The Future Of Note-Taking: 11 AI Tools You Need In 2026

Discover the 11 leading AI-powered note-taking tools in 2026, their features, and what makes them essential for users across various needs.

Analyzing AI Compression Workflows: Focus On Local LLMs In 2026

Exploring how native quantization-aware training shapes local large language model deployment in 2026, with focus on new low-precision formats and workflows.

Claude’s AI Marketplace: A Look At Anthropic’s 2,000+ Plugins And Connectors

A BleepingComputer report says Anthropic is turning Claude into an AI marketplace with more than 2,000 plugins and connectors. Details remain unconfirmed.