Token
A token is the smallest unit of text that a large language model processes. One token roughly corresponds to ~0.75 English words or 1-2 Chinese characters. Models don't see characters — they see token IDs.
Why it matters
- API billing: nearly all LLM services charge per token (input + output billed separately).
- Context window: a model's context window is measured in tokens (4k / 32k / 200k / 1M).
- Inference speed: tokens/s is the core throughput metric.
Rough estimates
| Text | Approx. tokens |
|---|---|
| 300-page book | ~100k |
| News article | ~1k |
| Line of code | ~10-30 |
Note: Chinese tokenization is typically 2-3x less efficient than English — one Chinese character may split into 1-3 tokens depending on the tokenizer.