Skip to main content
ZICQ

Wiki Concepts

Multimodal Models

Concepts
Aliases: MLLM VLM vision-language model ·2026-09-14

Multimodal Models

Multimodal models process multiple data types simultaneously: text, images, audio, video. Representative works: GPT-4V / GPT-5, Claude 3.5 Sonnet, Gemini 1.5, Qwen-VL, LLaVA, InternVL.

Architecture

Most multimodal LLMs consist of three parts:

  1. Modality encoder: image (ViT/CLIP), audio (Whisper), video.
  2. Projection layer: maps different modality features into the LLM's token space.
  3. LLM backbone: language modeling and generation on a unified token stream.

Capability dimensions

  • Image understanding: OCR, chart interpretation, VQA, object detection (some).
  • Image generation: DALL-E, Midjourney, Stable Diffusion (not multimodal LLMs).
  • Audio understanding / generation: Whisper (ASR), Vall-E (TTS).
  • Video understanding: Gemini 1.5, long-video summarization.

Evaluation

  • VQA, MME, MMMU (multimodal reasoning), MathVista (visual math).

Challenges

  • OCR errors: small text in images still struggles.
  • Spatial reasoning: models are weak on "upper-left" style spatial relations.
  • Hallucination: models may invent content not present in the image.