Multimodal Models
Multimodal models process multiple data types simultaneously: text, images, audio, video. Representative works: GPT-4V / GPT-5, Claude 3.5 Sonnet, Gemini 1.5, Qwen-VL, LLaVA, InternVL.
Architecture
Most multimodal LLMs consist of three parts:
- Modality encoder: image (ViT/CLIP), audio (Whisper), video.
- Projection layer: maps different modality features into the LLM's token space.
- LLM backbone: language modeling and generation on a unified token stream.
Capability dimensions
- Image understanding: OCR, chart interpretation, VQA, object detection (some).
- Image generation: DALL-E, Midjourney, Stable Diffusion (not multimodal LLMs).
- Audio understanding / generation: Whisper (ASR), Vall-E (TTS).
- Video understanding: Gemini 1.5, long-video summarization.
Evaluation
- VQA, MME, MMMU (multimodal reasoning), MathVista (visual math).
Challenges
- OCR errors: small text in images still struggles.
- Spatial reasoning: models are weak on "upper-left" style spatial relations.
- Hallucination: models may invent content not present in the image.