Skip to main content
ZICQ

Wiki Infrastructure

Hugging Face

Infrastructure
Aliases: HuggingFace HF HF Hub ·2026-09-14

Hugging Face

Hugging Face (huggingface.co) is the "GitHub + AWS" of AI: open-source model hosting, datasets, Spaces (demos), Inference services all in one place. The de facto standard for the open-source AI ecosystem.

Core modules

Hub

  • 1.5M+ public models
  • 250k+ datasets
  • 300k+ Spaces (demo apps)

Transformers library

  • Unified loading / fine-tuning for any open-source LLM
  • pipeline() one-line API (summarize, translate, embed)
  • Full PyTorch / TensorFlow / JAX compatibility

Datasets library

  • Streaming load of large datasets
  • Built-in cache, memory mapping

PEFT / TRL / Accelerate

  • PEFT: LoRA / QLoRA fine-tuning (lora).
  • TRL: RLHF / DPO / PPO training (rlhf / dpo).
  • Accelerate: multi-GPU / TPU / mixed precision training.

Inference

  • Serverless Inference API: free tier, per-token billing.
  • Inference Endpoints: self-deploy model API (on AWS etc).
  • Spaces: free GPU demo hosting.

Mainstream model hosting examples

  • meta-llama/Meta-Llama-3-8B-Instruct
  • Qwen/Qwen2.5-7B-Instruct
  • BAAI/bge-large-en-v1.5 (embedding)
  • Salesforce/blip-image-captioning-large (multimodal)
  • openai/clip-vit-large-patch14 (multimodal)

Typical workflow

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct", device_map="auto")

inputs = tokenizer("Hello", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0]))

Commercial / Compliance

  • Model licenses vary by model, check HuggingFace model card.
  • Hub private repos charge monthly.
  • Training data ethics review (Opt-out mechanism).