最新模型
Mistral Large 4 Opens API Preview; Trillion-Parameter MoE Model Scores 81.7% on Vulnerability Reproduction and Patching Evaluation
Mistral launched a public preview of Large 4 on October 6, featuring a native multimodal MoE architecture. Model weights are expected by the end of the month. An independent cybersecurity evaluation found strong vulnerability-handling performance, but passing the test does not mean the original vulnerability was fully fixed.

Mistral launched a public preview of Large 4 on October 6, which developers can try through the Mistral Studio API. The company says the model has about one trillion parameters, with roughly 49 billion activated per token, and supports native multimodal understanding. Weights are expected by the end of the month, and red-team testing in real-world scenarios is still underway. For now, the release offers a hosted trial; self-hosting will have to wait for the weights and full technical documentation. Official announcement
The technical significance of MoE is that each generation uses only a subset of expert parameters, allowing total capacity and per-token computation to scale separately. However, this architecture means that 49 billion active parameters should not be treated as the memory requirement for the entire model: deployment still needs to accommodate or schedule a large pool of expert weights, in addition to the context cache and inference-system overhead. Engineering teams should wait for details on the actual weight format, quantization support, and throughput benchmarks before estimating hardware needs.
More concrete evidence of capability comes from Artificial Analysis’s CyberGym-E2E-AA evaluation: Large 4 Preview achieved an 81.7% success rate on a single attempt, ranking first among the models listed on the page. The evaluation covers 131 C/C++ memory-safety tasks, with one task selected per project. In an isolated sandbox, agents must find a vulnerability, generate a proof-of-concept input that triggers a crash, and modify the code; each task has a 90-minute time limit. This is closer to an engineering workflow involving source-code analysis, tool use, and result verification than simply answering cybersecurity questions. Independent evaluation
But this score has clear limits. Success requires making the unpatched version crash, having the patch eliminate that crash, and keeping the original functional tests passing. Whether the patch fixes the underlying vulnerability identified in the dataset is recorded only as diagnostic information and does not count toward the main score. Therefore, 81.7% should not be interpreted as the full vulnerability-fix rate, nor directly compared with results from the original Berkeley evaluation. Scoring methodology
Teams preparing to adopt cybersecurity or code agents should next validate patch correctness on their own projects, add further tests and human review, and track total task costs, failure modes, and tool-permission requirements. They should also follow up at the end of the month on licensing, architectural details, and deployment support. Mistral says reinforcement learning behind the preview is still ongoing, so today’s evaluation should be treated as a snapshot of the current version. Release updates