Skip to main content
ZICQ

Wiki Models & Products

NVIDIA

Models & Products
Aliases: NVIDIA Jensen Huang ·2026-09-19

NVIDIA

NVIDIA is the GPU and accelerated-computing company founded in 1993 by Jensen Huang (黄仁勋). After betting on deep learning with CUDA in the 2010s, NVIDIA pivoted from a gaming-GPU vendor into the picks-and-shovels provider for AI infrastructure. By 2025 its market cap briefly topped $4T, among the world's most valuable companies.

Hardware lineup

Datacenter GPUs

  • H100 (Hopper): 2022 release, the workhorse for AI training / inference through 2023-2024
  • H200: H100 refresh, 141GB HBM3e, larger memory bandwidth
  • B100 / B200 (Blackwell): 2024 release, FP4/FP6 support, ~2× training performance
  • GB200 NVL72: rack-scale with 72 B200 + 36 Grace CPUs
  • B300 / Rubin (2026): next generation, focused on inference cost

Workstation / edge

  • RTX 4090 / 5090: consumer-grade, dominant for AI inference / fine-tuning
  • Jetson Orin / Thor: edge AI for robotics / autonomous driving

Software stack

  • CUDA: low-level moat; nearly every deep-learning framework defaults to it
  • cuDNN / cuBLAS / NCCL: high-performance math and communication libs
  • TensorRT: inference compiler with quantization + kernel fusion
  • Triton Inference Server: inference serving
  • NIM (NVIDIA Inference Microservice): packages models as microservices — essentially an "AI app store"

In-house models

  • Nemotron family: open LLM / multimodal, focused on reasoning and enterprise
  • Cosmos: world foundation model, positioned against Sora
  • Earth-2: digital twin for climate / weather simulation

Platforms

  • DGX (single-node AI supercomputer), HGX (datacenter module), MGX (OEM)
  • DGX Cloud: GPU cloud with hyperscaler partners
  • AI Enterprise: enterprise AI software bundle (NeMo, BioNeMo, Earth-2)

Influence

  • Virtually every frontier LLM (GPT-4 / Claude / Gemini / Llama / DeepSeek / Qwen) trains on NVIDIA GPUs
  • Held >90% of the AI accelerator market in 2024-2025
  • The "AI picks-and-shovels" narrative: no matter which model wins, NVIDIA collects
  • Acquisitions: Mellanox (networking), the failed Arm bid, Run:ai (AI orchestration) — all to complete the stack

Limitations

  • Customer concentration in hyperscalers and top AI labs forces diversification pressure
  • Software stack is CUDA-locked; new entrants (AMD ROCm, Google TPU, Amazon Trainium) keep chipping away
  • Export controls (China restricted on H100 / H200, A800 / H20 as downgrades) hit the largest market
  • Heavy personality dependence on Jensen Huang (leather jacket + one-chip keynotes); succession and institutionalization are long-term risks