NVIDIA
NVIDIA is the GPU and accelerated-computing company founded in 1993 by Jensen Huang (黄仁勋). After betting on deep learning with CUDA in the 2010s, NVIDIA pivoted from a gaming-GPU vendor into the picks-and-shovels provider for AI infrastructure. By 2025 its market cap briefly topped $4T, among the world's most valuable companies.
Hardware lineup
Datacenter GPUs
- H100 (Hopper): 2022 release, the workhorse for AI training / inference through 2023-2024
- H200: H100 refresh, 141GB HBM3e, larger memory bandwidth
- B100 / B200 (Blackwell): 2024 release, FP4/FP6 support, ~2× training performance
- GB200 NVL72: rack-scale with 72 B200 + 36 Grace CPUs
- B300 / Rubin (2026): next generation, focused on inference cost
Workstation / edge
- RTX 4090 / 5090: consumer-grade, dominant for AI inference / fine-tuning
- Jetson Orin / Thor: edge AI for robotics / autonomous driving
Software stack
- CUDA: low-level moat; nearly every deep-learning framework defaults to it
- cuDNN / cuBLAS / NCCL: high-performance math and communication libs
- TensorRT: inference compiler with quantization + kernel fusion
- Triton Inference Server: inference serving
- NIM (NVIDIA Inference Microservice): packages models as microservices — essentially an "AI app store"
In-house models
- Nemotron family: open LLM / multimodal, focused on reasoning and enterprise
- Cosmos: world foundation model, positioned against Sora
- Earth-2: digital twin for climate / weather simulation
Platforms
- DGX (single-node AI supercomputer), HGX (datacenter module), MGX (OEM)
- DGX Cloud: GPU cloud with hyperscaler partners
- AI Enterprise: enterprise AI software bundle (NeMo, BioNeMo, Earth-2)
Influence
- Virtually every frontier LLM (GPT-4 / Claude / Gemini / Llama / DeepSeek / Qwen) trains on NVIDIA GPUs
- Held >90% of the AI accelerator market in 2024-2025
- The "AI picks-and-shovels" narrative: no matter which model wins, NVIDIA collects
- Acquisitions: Mellanox (networking), the failed Arm bid, Run:ai (AI orchestration) — all to complete the stack
Limitations
- Customer concentration in hyperscalers and top AI labs forces diversification pressure
- Software stack is CUDA-locked; new entrants (AMD ROCm, Google TPU, Amazon Trainium) keep chipping away
- Export controls (China restricted on H100 / H200, A800 / H20 as downgrades) hit the largest market
- Heavy personality dependence on Jensen Huang (leather jacket + one-chip keynotes); succession and institutionalization are long-term risks