Mistral Large 4 Preview Trains a Trillion Parameters in Europe
Mistral launched the Large 4 public preview on 6 October: a 1-trillion-parameter multimodal MoE with 49 billion active parameters, trained from scratch on 3,800 Grace Blackwell GPUs in the company's own European datacenters. Mistral claims leading cyber and coding results among open models, and open weights are due at the end of October after red-teaming.
On this page
A trillion parameters forged in Europe
Mistral opened the public preview of Mistral Large 4 on 6 October, the company's first trillion-parameter model: a natively multimodal Mixture-of-Experts with 49 billion active parameters, hybrid instruct-and-reasoning behavior, and training data covering 160-plus languages. Two details distinguish it from the usual frontier preview. It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, with the preview served from that same infrastructure. And the plan states plainly that open weights arrive at the end of October, after red-teaming with cybersecurity firms, partners, and state authorities.
The company calls the model an attempt to push the frontier of open-weight performance, notes that its reinforcement-learning run is still in flight with substantial headroom, and ties the roadmap to a 3 billion euro Series D that it describes as the largest equity round ever raised by a European technology company.
Where Mistral claims the lead
The benchmark claims concentrate on cyber and agents. Mistral reports a top-five global rank on the AA Cyber Index, an 82 percent score on a reproduce-and-patch test it says closed rivals like Claude Opus 5.5 and GPT-6 Astra decline to attempt, 93 percent on Cybench, and 93.3 percent attack resistance on Lakera's B3 benchmark. On coding, the model posts 61.7 percent on DeepSWE v1.1 and ranks second only to Claude Opus 5 in Surge AI's blind human evaluation.
Third-party numbers already support parts of the story: coverage of the launch notes a 1 million-token context window and preview pricing of about 1.36 dollars per million input tokens and 4.18 dollars per million output tokens, with independent trackers listing it as the largest model release of the month. The preview's own comparison tables put it ahead of DeepSeek V4 Pro and Qwen3.8 Max on the company's coding-agent index and ahead of GPT-6 Astra on a vision benchmark, claims that deserve the usual discount for vendor-selected tests but are unusually specific.
The open-weights promise, and its gate
For the local and self-hosting community, everything hinges on the end-of-month release. A 49B-active trillion-parameter MoE lands in the same deployment class as other frontier-adjacent open models: multi-GPU servers and expert-caching engines, not laptops. The unresolved question is the license, which the announcement does not name, and the red-teaming gate, which means the shipped weights will have passed through a moderation process the preview has not.
Mistral's argument is that sovereignty requires more than licenses; it requires the training run to happen on European soil, on European terms, with weights that European institutions can eventually hold. Kolibri made that argument for a 740M embedder and a national government. A trillion-parameter model trained in the company's own datacenters is the same claim at frontier scale, and the October release will show whether the weights arrive as promised.