Rerank
Rerank is the second step of the RAG retrieval flow: use a stronger model to refine first-stage results, pushing the truly most relevant chunks to the top.
Why rerank is needed
First-stage (vector / BM25) uses bi-encoders: query and doc are independently encoded into vectors, then similarity is computed. Pros: fast, batchable. Cons: query-doc interaction info is lost.
Rerank uses cross-encoders: query and doc go into the model together, each pair gets a relevance score separately. Much higher precision, but N× slower — so can't be used on the full corpus.
Classic flow
1. First-stage (top-100): vector + BM25 recall 100 candidates
2. Rerank (top-5): use cross-encoder to refine 100 candidates down to 5
3. LLM generation: feed 5 chunks to LLM, get answer
First-stage ensures recall, rerank ensures precision, LLM gets high-quality context.
Mainstream rerank models
| Model | Type | Performance / price |
|---|---|---|
| Cohere Rerank 3 | SaaS | Industry SOTA, cheap |
| bge-reranker-v2-m3 | Open-source | Multilingual, good value |
| Jina Reranker | SaaS | Lightweight |
| gte-rerank | Open-source | Alibaba DAMO |
| ms-marco-MiniLM | Open-source | Classic English baseline |
Practical experience
- cohere / bge-reranker-v2-m3 is the default starting point.
- First-stage should be wide: top-100 reranks much better than top-20.
- Multilingual rerank is required: Chinese → bge-reranker-v2-m3, English → bge-reranker-large.
- Local deployment: bge-reranker on 4090 at 1ms/pair, single card handles 10k QPS.