Chapter 28 β The Hugging Face Ecosystem
βHugging Face is the GitHub of machine learning β a platform where models, datasets, and tools converge into a unified ecosystem that accelerates research and deployment.β
Ecosystem Overview
The Hub
The central repository for sharing ML artifacts:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
from huggingface_hub import HfApi, snapshot_download
# Browse and download
api = HfApi()
models = api.list_models(filter="llama", sort="downloads", direction=-1, limit=5)
# Download a model
snapshot_download("meta-llama/Llama-3.1-8B-Instruct", local_dir="./llama3")
# Upload your own model
api.upload_folder(
folder_path="./my_model",
repo_id="username/my-model",
repo_type="model",
)
Key Hub features:
- Model cards: Standardized documentation (training data, eval results, limitations)
- Gated models: Access control for restricted models (LLaMA, Gemma)
- Safetensors: Safe, fast model weight format (replaces pickle-based .bin)
- GGUF support: Quantized models for llama.cpp directly on the Hub
transformers Library
The core library providing a unified API across model families:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load any causal LM with identical API
model_id = "meta-llama/Llama-3.1-8B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto", # Automatic multi-GPU placement
attn_implementation="flash_attention_2",
)
# Chat-style generation
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain attention in one paragraph."},
]
input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
output = model.generate(
input_ids,
max_new_tokens=256,
temperature=0.7,
top_p=0.9,
do_sample=True,
)
print(tokenizer.decode(output[0][input_ids.shape[1]:], skip_special_tokens=True))
Auto Classes
The Auto* pattern dispatches to the correct architecture based on config:
| Auto Class | Purpose | Example Models |
|---|---|---|
AutoModel |
Base model (no head) | Feature extraction |
AutoModelForCausalLM |
Decoder-only generation | GPT-2, LLaMA, Mistral |
AutoModelForSeq2SeqLM |
Encoder-decoder | T5, BART |
AutoModelForSequenceClassification |
Text classification | BERT + head |
AutoModelForTokenClassification |
NER, POS tagging | BERT + token head |
AutoModelForQuestionAnswering |
Extractive QA | BERT/RoBERTa + QA head |
AutoModelForVision2Seq |
Image-to-text | LLaVA, PaliGemma |
Generation Config
1
2
3
4
5
6
7
8
9
10
11
12
from transformers import GenerationConfig
config = GenerationConfig(
max_new_tokens=512,
do_sample=True,
temperature=0.6,
top_p=0.9,
top_k=50,
repetition_penalty=1.1,
stop_strings=["<|eot_id|>"],
)
output = model.generate(input_ids, generation_config=config)
Pipeline API
Quick inference without manual setup:
1
2
3
4
5
6
7
8
9
10
11
12
from transformers import pipeline
# Text generation
generator = pipeline("text-generation", model="microsoft/Phi-3-mini-4k-instruct",
device_map="auto", torch_dtype="auto")
result = generator("The key insight of attention is", max_new_tokens=100)
# Other pipelines
classifier = pipeline("sentiment-analysis") # Default: distilbert
ner = pipeline("ner", grouped_entities=True) # Named entity recognition
qa = pipeline("question-answering") # Extractive QA
summarizer = pipeline("summarization", model="facebook/bart-large-cnn")
Trainer API
Structured training with logging, evaluation, and checkpointing:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
from transformers import Trainer, TrainingArguments
training_args = TrainingArguments(
output_dir="./checkpoints",
num_train_epochs=3,
per_device_train_batch_size=4,
gradient_accumulation_steps=8,
learning_rate=2e-5,
lr_scheduler_type="cosine",
warmup_ratio=0.1,
bf16=True,
logging_steps=10,
eval_strategy="steps",
eval_steps=500,
save_strategy="steps",
save_steps=500,
report_to="wandb",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
data_collator=data_collator,
tokenizer=tokenizer,
)
trainer.train()
Key Companion Libraries
PEFT β Parameter-Efficient Fine-Tuning
1
2
3
4
5
6
7
8
9
10
11
from peft import LoraConfig, get_peft_model
lora_config = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable params: 83,886,080 || all params: 8,114,769,920 || trainable%: 1.03%
TRL β Transformer Reinforcement Learning
1
2
3
4
5
6
7
8
9
10
11
12
from trl import SFTTrainer, DPOTrainer
# SFT
sft_trainer = SFTTrainer(model=model, args=training_args,
train_dataset=dataset, peft_config=lora_config)
sft_trainer.train()
# DPO
dpo_trainer = DPOTrainer(model=model, ref_model=ref_model,
args=training_args, train_dataset=preference_data,
beta=0.1, peft_config=lora_config)
dpo_trainer.train()
Accelerate β Distributed Training
1
2
3
4
5
6
7
8
9
10
11
12
from accelerate import Accelerator
accelerator = Accelerator(mixed_precision="bf16")
model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)
for batch in dataloader:
outputs = model(**batch)
accelerator.backward(outputs.loss)
optimizer.step()
optimizer.zero_grad()
# Launch: accelerate launch --multi_gpu --num_processes 8 train.py
datasets β Efficient Data Loading
1
2
3
4
5
6
from datasets import load_dataset
# Stream large datasets without downloading
dataset = load_dataset("HuggingFaceFW/fineweb", split="train", streaming=True)
for example in dataset:
tokens = tokenizer(example["text"])["input_ids"]
Whatβs Next
Now that we understand the tools, the next chapter dives into the transformers library itself β its class hierarchy, model internals, and how to extend it.
β Previous: Chapter 27 β ASICs & Accelerators Β· Next: Chapter 29 β Transformers Library Deep Dive β
Last updated: April 2026