Mistral AI
Mistral AI is a Paris, France AI company (founded 2023), focused on open-source LLMs and inference optimization. Europe's strongest, world's #2 open-source LLM provider (behind Meta's Llama).
Key products
| Model | Params | Type | Highlights |
|---|---|---|---|
| Mistral 7B | 7B | Dense | Early open-source benchmark |
| Mixtral 8x7B | 47B | MoE | Early MoE benchmark |
| Mistral Small | 22B | Dense | Commercial mainstream |
| Mistral Medium | ~70B | Dense | Commercial high-end |
| Mistral Large | ~123B | Dense | Commercial flagship |
| Codestral | 22B | Dense | Code specialized |
| Mistral NeMo | 12B | Dense | NVIDIA collaboration |
Key technical highlights
- Sliding Window Attention (inference): each layer only attends to most recent N tokens, KV cache memory drops dramatically.
- Early MoE practice: Mixtral 8x7B showed industry MoE can be trained well.
- Extreme inference optimization: proprietary vLLM / TensorRT-LLM kernels, runs faster than Llama.
Business model
- Open + closed dual track: base models open-source (Apache 2.0), flagship models closed-source (La Plateforme API).
- EU compliance: GDPR / EU AI Act friendly, European government / financial customer first choice.
- Price advantage: 5-10x cheaper than GPT-4 / Claude, comparable quality.
Compared with Llama
- Ecosystem: Llama >> Mistral (community fine-tunes, HuggingFace integration).
- Multilingual: Mistral slightly better (French native + European multilingual training).
- Inference speed: Mistral usually 10-20% faster than same-size Llama.
- Documentation / papers: Mistral papers are high quality (technical details well disclosed).
Use for
- European enterprises (compliance requirements).
- Cost-sensitive small teams.
- Real-time applications needing inference speed.
Limitations
- Flagship models are closed-source, can't deploy Large version locally.
- Chinese capability weaker than Qwen / DeepSeek.
- Community scale an order of magnitude smaller than Llama.