Skip to main content
ZICQ

Wiki Models & Products

Groq

Models & Products
Aliases: Groq Groq LPU LPU Inference Engine ·2026-09-14

Groq

Groq is an American AI inference chip + cloud service company. Its self-developed LPU (Language Processing Unit) inference chip is known for "extreme low latency". In the LLM API market, it's the third option beyond OpenAI / Anthropic.

LPU vs GPU

  • GPU (NVIDIA): general-purpose parallel computing, HBM high bandwidth, but kernel launch overhead unfriendly for single batch.
  • LPU (Groq): ASIC designed for LLM inference, deterministic latency, no kernel launch overhead. Single-token latency 5-10x lower than GPU.

Groq Cloud models

Groq doesn't develop its own models, provides open-source model inference services:

  • Llama 3.1 8B / 70B / 405B
  • Mixtral 8x7B
  • Gemma 2 9B
  • Whisper (voice transcription)

Price

Extremely cheap: Llama 70B inference $0.59/M tokens (input + output average), 20x cheaper than GPT-4o.

Real-world speed

Llama 3.1 70B:

  • Groq LPU: ~280 tokens/s
  • vLLM on H100: ~120 tokens/s
  • OpenAI API (GPT-4o): ~80 tokens/s

Use cases

  • Real-time chat / streaming output: low latency user experience much better than GPT-4.
  • Code completion: instant response is productivity key.
  • Batch API calls: cheap + fast, great for evaluation.
  • Audio transcription: Whisper runs blazingly fast on LPU.

Limitations

  • Open-source models only: can't use Claude / GPT-5 etc closed-source models.
  • Limited capacity: peak hours need queueing or rate limit.
  • No fine-tuning: prompt tuning only.

Compared with other inference services

Service Speed Price Model selection
Groq LPU ★★★★★ ★★★★★ Open source
Together.ai ★★★★ ★★★★ Open + closed
OpenRouter ★★★★ ★★★★ Aggregates providers
Fireworks ★★★★ ★★★★ Open source
Replicate ★★★ ★★★ Open source (CPU/GPU)
OpenAI API ★★★ ★★ Closed-source flagship