Code Generation
Code Generation is the capability of an LLM to write runnable code from natural language descriptions. One of the most mature and commercially valuable application areas of current LLMs.
Task spectrum
| Task | Example |
|---|---|
| Completion | Complete function body from signature |
| Rewrite | Convert Python code to Rust |
| Fix | Generate patch from error stacktrace |
| Refactor | Split 50-line function into 5 small ones |
| Generate tests | Write unit tests for a function |
| Understand | Explain what this code does |
| Whole-repo PR | Agent fixing bugs in multi-file codebase (swe-bench) |
Evaluation dimensions
- Correctness: unit test pass rate (HumanEval, SWE-bench).
- Consistency: matches existing code style.
- Safety: doesn't introduce SQL injection, XSS, etc.
- Maintainability: naming, abstraction, comment quality.
Representative products
- Editors: cursor, Windsurf, Cody, Continue, Zed
- CLI: Claude Code, Aider, OpenHands
- Cloud agents: Devin, Manus, Replit Agent, v0, Bolt
- Embedded: GitHub Copilot, JetBrains AI Assistant
Mainstream models
- Closed: claude (long-time SWE-bench leader), gpt-5
- Open: qwen-Coder, DeepSeek-Coder, Code Llama, StarCoder
Practical experience
- Context > Model: providing rich context (project structure, dependencies, style examples) beats upgrading to a more expensive model.
- Test-driven: have model write tests first, then implementation, then run tests — 30%+ accuracy gain over "write it all at once".
- Small changes first: break large features into multiple small PRs, each independently reviewed.
- Human review non-negotiable: models introduce subtle bugs (boundary conditions, null pointers); humans must review.