Skip to main content
ZICQ

Wiki Concepts

Rerank

Concepts
Aliases: reranking cross-encoder ·2026-09-14

Rerank

Rerank is the second step of the RAG retrieval flow: use a stronger model to refine first-stage results, pushing the truly most relevant chunks to the top.

Why rerank is needed

First-stage (vector / BM25) uses bi-encoders: query and doc are independently encoded into vectors, then similarity is computed. Pros: fast, batchable. Cons: query-doc interaction info is lost.

Rerank uses cross-encoders: query and doc go into the model together, each pair gets a relevance score separately. Much higher precision, but N× slower — so can't be used on the full corpus.

Classic flow

1. First-stage (top-100): vector + BM25 recall 100 candidates
2. Rerank (top-5): use cross-encoder to refine 100 candidates down to 5
3. LLM generation: feed 5 chunks to LLM, get answer

First-stage ensures recall, rerank ensures precision, LLM gets high-quality context.

Mainstream rerank models

Model Type Performance / price
Cohere Rerank 3 SaaS Industry SOTA, cheap
bge-reranker-v2-m3 Open-source Multilingual, good value
Jina Reranker SaaS Lightweight
gte-rerank Open-source Alibaba DAMO
ms-marco-MiniLM Open-source Classic English baseline

Practical experience

  • cohere / bge-reranker-v2-m3 is the default starting point.
  • First-stage should be wide: top-100 reranks much better than top-20.
  • Multilingual rerank is required: Chinese → bge-reranker-v2-m3, English → bge-reranker-large.
  • Local deployment: bge-reranker on 4090 at 1ms/pair, single card handles 10k QPS.