Vector Database
A vector database is a database system purpose-built for storing and retrieving high-dimensional vectors. Embedded documents / images / user features live here, supporting nearest-neighbor (ANN) retrieval.
Core capabilities
- Vector index: HNSW, IVF, PQ algorithms reduce O(n) brute force to sub-millisecond.
- Metadata filtering: combine vector similarity + tag/category/date filter queries.
- Hybrid retrieval: dual-vector + BM25 retrieval with fusion.
- Horizontal scale: support 100M to 100B vectors.
Mainstream products
Open-source self-hosted
- Qdrant (Rust): high-performance, easy to deploy, REST + gRPC.
- Milvus (Go/C++): CNCF project, 100M-scale production choice.
- Weaviate (Go): modular, built-in RAG pipeline.
- Chroma (Python): lightweight, first choice for prototypes.
- LanceDB (Rust embedded): for embedded scenarios.
- pgvector: PostgreSQL plugin, no new components.
Cloud-hosted
- Pinecone: most mature SaaS.
- Weaviate Cloud, Qdrant Cloud: managed versions of open-source.
- Elasticsearch dense_vector: zero-migration for ES users.
Selection guide
| Scale | Recommendation |
|---|---|
| Prototype / < 100k vectors | Chroma / LanceDB / pgvector |
| Production / 1M-100M | Qdrant / Milvus / Weaviate |
| Hyperscale / > 100M | Milvus / Pinecone / Vespa |
| Don't want new components | pgvector / ES dense_vector |
Key metrics
- Recall@10: retrieval quality (vs brute force).
- QPS: queries per second.
- P99 latency: tail latency (vector DBs usually < 10ms).
- Memory: raw vector size × index inflation factor (HNSW ~2x, IVF ~1.2x).