Context Window
The context window is the maximum number of tokens an LLM can "see" in one go. Content beyond the window is truncated and the model cannot reference it directly.
Common specs
| Model | Context window |
|---|---|
| GPT-3.5 | 4k → 16k |
| GPT-4 | 8k → 128k |
| Claude 3.5 | 200k |
| Gemini 1.5 Pro | 1M → 2M |
| Llama 3 | 8k |
| Qwen2.5 | 128k |
| DeepSeek V3 | 64k |
Long-context pitfalls
- Effective length < official claim: models show noticeable recall drop in the "middle" of long contexts (the "Lost in the Middle" effect).
- Cost grows linearly: more input tokens = higher API cost + longer latency.
- Attention dilution: relevant content gets drowned in noise.
Useful tools
- Needle in a Haystack test: hide a "needle" in long text, check retrieval.
- RULER / LongBench benchmarks: stricter long-context evaluation.
- RAG first: if RAG can solve it, don't stuff it into context.