Embedding
Embedding maps arbitrary data (text, image, audio, user behavior) into a fixed-dimensional dense vector representation. In this space:
- Semantically similar inputs are close in distance.
- Vector arithmetic carries semantic meaning ("king - man + woman ≈ queen").
Applications in LLM
- Retrieval (RAG): embed docs and queries, use cosine to find most relevant.
- Classification: a simple linear classifier on embeddings works well.
- Clustering: semantically similar docs cluster together.
- Recommendation: user vector × item vector = interest score.
Mainstream models
- OpenAI:
text-embedding-3-small/text-embedding-3-large(3072 dim). - Open source: BGE, M3E, GTE, Qwen3-Embedding, E5-Mistral.
- Multimodal: CLIP, BLIP-2, Qwen2-VL can embed images and text together.
Caveats
- Higher dimension isn't always better: look at benchmarks, not the number.
- Same source principle: docs and queries must use the same embedding model.
- Normalize: for cosine similarity, L2-normalize before storage.