Chain-of-Thought (CoT)
Chain-of-Thought (CoT) is the prompting technique that makes an LLM explicitly output intermediate reasoning steps before giving a final answer. Significantly improves accuracy on complex tasks (math, logic, planning).
Classic example
Ask: "Roger has 5 tennis balls. He buys 2 cans of 3 each. How many?"
- Without CoT: may answer wrong (5 + 2 × 3 = 11 instead of 5 + 2 + 3 = 10).
- With CoT: "Let's think step by step. Roger started with 5. 2 cans of 3 = 6. Total 5 + 6 = 11."
Why it works
- Force serialization: makes the model chain multi-step reasoning in a single forward pass, avoiding skipping steps.
- Debuggable: humans can see where the model goes wrong, and optimize the prompt accordingly.
- Gains scale with model size: small models may not learn CoT format (output nonsense); large models gain the most.
Advanced variants
- Self-Consistency: sample multiple CoT paths, take majority answer.
- Zero-Shot CoT: just add "Let's think step by step" to the prompt, no examples needed.
- ReAct: CoT + tool calls interleaved (Reasoning + Acting).
- Tree of Thoughts (ToT): expand multiple CoT paths, BFS/DFS to pick best.
- Program-Aided Language Models (PAL): convert CoT to code, run in Python interpreter.
Limitations
- Cost: explicit reasoning consumes more tokens (reasoning models 5-10x latency).
- Not always correct: "looks like reasoning" doesn't mean true reasoning, may be post-hoc rationalization.
- Small models regress: sub-1B models output poor CoT, direct answer is more accurate.
Practical experience
- CoT is the #1 cost-effective prompt technique.
- In Chinese: "请一步步思考" ≈ "Let's think step by step".
- Add "please put reasoning process in tags" — frontends can capture and display.
- Pairs best with reasoning models (o1/R1); with regular models gain is modest.