Groq
Groq is an American AI inference chip + cloud service company. Its self-developed LPU (Language Processing Unit) inference chip is known for "extreme low latency". In the LLM API market, it's the third option beyond OpenAI / Anthropic.
LPU vs GPU
- GPU (NVIDIA): general-purpose parallel computing, HBM high bandwidth, but kernel launch overhead unfriendly for single batch.
- LPU (Groq): ASIC designed for LLM inference, deterministic latency, no kernel launch overhead. Single-token latency 5-10x lower than GPU.
Groq Cloud models
Groq doesn't develop its own models, provides open-source model inference services:
- Llama 3.1 8B / 70B / 405B
- Mixtral 8x7B
- Gemma 2 9B
- Whisper (voice transcription)
Price
Extremely cheap: Llama 70B inference $0.59/M tokens (input + output average), 20x cheaper than GPT-4o.
Real-world speed
Llama 3.1 70B:
- Groq LPU: ~280 tokens/s
- vLLM on H100: ~120 tokens/s
- OpenAI API (GPT-4o): ~80 tokens/s
Use cases
- Real-time chat / streaming output: low latency user experience much better than GPT-4.
- Code completion: instant response is productivity key.
- Batch API calls: cheap + fast, great for evaluation.
- Audio transcription: Whisper runs blazingly fast on LPU.
Limitations
- Open-source models only: can't use Claude / GPT-5 etc closed-source models.
- Limited capacity: peak hours need queueing or rate limit.
- No fine-tuning: prompt tuning only.
Compared with other inference services
| Service | Speed | Price | Model selection |
|---|---|---|---|
| Groq LPU | ★★★★★ | ★★★★★ | Open source |
| Together.ai | ★★★★ | ★★★★ | Open + closed |
| OpenRouter | ★★★★ | ★★★★ | Aggregates providers |
| Fireworks | ★★★★ | ★★★★ | Open source |
| Replicate | ★★★ | ★★★ | Open source (CPU/GPU) |
| OpenAI API | ★★★ | ★★ | Closed-source flagship |