Appendix C β HuggingFace Model Config Field Reference
βA
config.jsonis not just metadata β it is the single source of truth for every tensor shape in the model. A misread field propagates silently until your kernel crashes at 3 a.m.β
This appendix is a practitionerβs reference for every field you will encounter in a HuggingFace config.json. Each entry covers:
- What it is and the formula/derivation that depends on it
- Legal values and what they imply
- Caveats & pitfalls β especially for kernel developers, fine-tuners, and inference engineers
Jump to a section:
- Quick Reference Table
- Core Architecture Parameters
- Positional Encoding Parameters
- Normalization Parameters
- Attention Configuration
- Sliding-Window Attention
- Mixture of Experts
- Special Token IDs
- Generation & Runtime
- VLM Config β Qwen3-VL and LLaVA Style
- VLA Config β Alpamayo and RT-2 Style
- Cross-Model Comparison Table
- Kernel Development Pitfall Checklist
Quick Reference Table
| Field | Typical type | Example value | Section |
|---|---|---|---|
model_type |
string | "llama" |
Core Architecture |
architectures |
list[str] | ["LlamaForCausalLM"] |
Core Architecture |
vocab_size |
int | 128256 |
vocab_size |
hidden_size |
int | 4096 |
hidden_size |
num_hidden_layers |
int | 32 |
num_hidden_layers |
num_attention_heads |
int | 32 |
num_attention_heads |
num_key_value_heads |
int | 8 |
num_key_value_heads |
head_dim |
int (optional) | 128 |
head_dim β |
intermediate_size |
int | 14336 |
intermediate_size |
hidden_act |
string | "silu" |
hidden_act |
max_position_embeddings |
int | 131072 |
max_position_embeddings |
rope_theta |
float | 500000.0 |
rope_theta |
rope_scaling |
dict (optional) | {"type": "llama3", β¦} |
rope_scaling |
rms_norm_eps |
float | 1e-05 |
rms_norm_eps |
layer_norm_eps |
float | 1e-05 |
layer_norm_eps |
attention_bias |
bool | false |
attention_bias |
mlp_bias |
bool | false |
mlp_bias |
attention_dropout |
float | 0.0 |
attention_dropout |
use_sliding_window |
bool | true |
Sliding Window |
sliding_window |
int | 4096 |
Sliding Window |
max_window_layers |
int | 28 |
Sliding Window |
tie_word_embeddings |
bool | false |
tie_word_embeddings |
bos_token_id |
int | 1 |
Special Tokens |
eos_token_id |
int or list | 2 |
Special Tokens |
pad_token_id |
int (optional) | 0 |
Special Tokens |
use_cache |
bool | true |
Generation & Runtime |
torch_dtype |
string | "bfloat16" |
Generation & Runtime |
num_experts / num_local_experts |
int | 64 |
MoE |
num_experts_per_tok |
int | 2 |
MoE |
router_aux_loss_coef |
float | 0.001 |
MoE |
initializer_range |
float | 0.02 |
initializer_range |
transformers_version |
string | "4.46.1" |
Generation & Runtime |
Core Architecture Parameters
model_type / architectures
1
2
3
4
{
"model_type": "llama",
"architectures": ["LlamaForCausalLM"]
}
model_type is a short string key that maps to a registered model class inside the transformers source tree. architectures lists the full class name(s) that can load this config.
Legal values for model_type: "llama", "mistral", "qwen2", "qwen3", "gemma", "gemma2", "phi3", "falcon", "mpt", "bloom", "opt", "gpt_neox", "gpt2", "t5", "bert", β¦
Pitfalls:
- The
model_typestring must match exactly what the library registered viaAutoConfig. A typo silently falls through to a generic fallback class that may load with wrong shapes. - For custom / community models the field may not exist in the official registry. You must register it manually with
AutoConfig.register/AutoModelForCausalLM.registerbeforefrom_pretrainedwill work withouttrust_remote_code=True. architectures[0]is used by the Hub when someone clicks βUse in transformersβ. If you publish a model with a wrong value the generated snippet breaks.
vocab_size
1
{ "vocab_size": 128256 }
Number of tokens in the vocabulary. Sets the size of:
- The embedding matrix: $W_e \in \mathbb{R}^{V \times d}$
- The language-model head: $W_{\text{lm}} \in \mathbb{R}^{d \times V}$ (or $W_e^\top$ when
tie_word_embeddings=true)
Typical ranges: 32K (older GPT-2 era) β 256K (Qwen2/3 family)
Pitfalls:
- The tokenizerβs vocabulary size and
vocab_sizein the config must agree. When you add special tokens to a tokenizer and callmodel.resize_token_embeddings(), updateconfig.vocab_sizetoo β otherwisefrom_pretrainedwill load correctly but saving/reloading may corrupt the lm head. - Many kernels pre-allocate logit buffers as
[B, T, vocab_size]. With vocab sizes of 150K+ (Qwen3: 151,936), this dominates memory at long context. E.g. bfloat16, B=1, T=8192, V=151936 β 2.3 GB just for logits. vocab_sizeis sometimes padded to the nearest multiple of 64 or 128 for hardware efficiency (e.g. LLaMA pads to 32,064). The embedding matrix has that padded size, but tokens beyond the real vocabulary should never be produced. This padding breaks naive loop bounds in custom kernels.
hidden_size
1
{ "hidden_size": 4096 }
The residual stream dimension $d_{\text{model}}$. Every transformer layer reads and writes to a tensor of shape $[B, T, d_{\text{model}}]$.
Derived quantities: \(d_{\text{model}} = \texttt{hidden\_size}\) \(W_Q, W_K, W_V \in \mathbb{R}^{d_{\text{model}} \times d_{\text{model}}}\) \(\text{KV cache per layer} = 2 \cdot B \cdot T \cdot d_{\text{model}} \cdot \text{bytes\_per\_element}\)
Pitfalls:
hidden_sizedoes not by itself determine head shapes β seehead_dimbelow.- Tensor parallelism splits
hidden_sizeacross GPUs. Withtp=8andhidden_size=4096, each shard sees only 512 channels. Kernels that hard-code4096will break. - When
hidden_sizeis not a power of 2 (common in older encoder-only models like BERT-large: 1024), certain CUDA primitives that require power-of-2 tile sizes need guard branches.
num_hidden_layers
1
{ "num_hidden_layers": 32 }
Number of transformer decoder blocks $L$. Also called n_layer (GPT-2), num_layers (Falcon).
Total parameter contribution:
\(P_{\text{attn}} = L \cdot (4 \cdot d^2) \quad \text{(QKV + output projections, no bias)}\)
\(P_{\text{ffn}} = L \cdot (3 \cdot d \cdot d_{\text{ffn}}) \quad \text{(gate/up/down for SwiGLU)}\)
Pitfalls:
- MoE models have the same
num_hidden_layersbut some layers are dense and some are MoE β controlled bynum_expertsandmlp_type. Do not assume every layer has the same FLOPs. - Some architectures (Gemma 2, Mistral) interleave sliding-window and global attention layers. The index of a layer determines which variant is used. Kernels that treat all layers identically will produce wrong outputs.
num_attention_heads
1
{ "num_attention_heads": 32 }
Number of query heads $H_Q$ in multi-head attention (MHA) / grouped-query attention (GQA).
Memory layout (standard): \(Q \in \mathbb{R}^{B \times H_Q \times T \times d_{\text{head}}}\)
Pitfalls:
- This is the query head count, not the KV head count (see
num_key_value_heads). The KV cache is sized by the latter, notnum_attention_heads. - In tensor-parallel setups,
num_attention_headsmust be divisible bytp_degree. An odd number of heads (e.g.num_attention_heads=28in Qwen3-7B) meanstp=4is allowed buttp=3ortp=7is not. Always check divisibility before planning a TP configuration. - Do not use
num_attention_headsalone to computehead_dim. See the dedicated section.
num_key_value_heads
1
{ "num_key_value_heads": 8 }
Number of distinct KV heads $H_{KV}$. When $H_{KV} < H_Q$, each KV head is shared by $G = H_Q / H_{KV}$ query heads (GQA). When $H_{KV} = 1$ it is MQA.
KV cache size: \(\text{KV cache} = 2 \cdot L \cdot B \cdot T \cdot H_{KV} \cdot d_{\text{head}} \cdot \text{bytes}\)
Example β LLaMA 3.1 8B at T=8192, BF16: \(2 \times 32 \times 1 \times 8192 \times 8 \times 128 \times 2 = \approx 1.07 \text{ GB}\)
Pitfalls:
- Must divide evenly into
num_attention_heads: $G = H_Q / H_{KV}$ must be an integer. A non-integer ratio causes runtime errors. - In tensor-parallel setups,
num_key_value_headsmust be β₯tp_degreeand divisible by it, OR the framework must replicate the KV heads. vLLM replicates automatically; a custom kernel may not. num_key_value_headsis absent from older configs (pre-GQA) β treat missing as equal tonum_attention_heads(MHA).
head_dim
1
{ "head_dim": 128 }
β οΈ This is the most common source of silent shape bugs across model families.
$d_{\text{head}}$ is the per-head dimension. It determines the shape of every Q/K/V tensor and the KV cache.
Standard convention (no head_dim field in config)
\[d_{\text{head}} = \frac{\texttt{hidden\_size}}{\texttt{num\_attention\_heads}}\]
This is implicit β the field is absent from config.json and frameworks compute it at model instantiation time.
| Model | hidden_size |
num_attn_heads |
Derived head_dim |
|---|---|---|---|
| LLaMA 3.1 8B | 4096 | 32 | 128 |
| Mistral 7B | 4096 | 32 | 128 |
| Gemma 2 9B | 3584 | 16 | 224 |
| Phi-3 Mini | 3072 | 32 | 96 |
Qwen3: explicit override
All Qwen3 models set head_dim: 128 explicitly in config.json, regardless of hidden_size / num_attention_heads:
| Model | hidden_size |
num_attn_heads |
NaΓ―ve formula | Actual head_dim |
|---|---|---|---|---|
| Qwen3-0.6B | 1024 | 16 | 64 | 128 |
| Qwen3-1.7B | 2048 | 16 | 128 | 128 |
| Qwen3-4B | 2560 | 32 | 80 | 128 |
| Qwen3-8B | 4096 | 64 | 64 | 128 |
| Qwen3-14B | 5120 | 40 | 128 | 128 |
| Qwen3-32B | 5120 | 64 | 80 | 128 |
| Qwen3-235B-A22B | 4096 | 64 | 64 | 128 |
The implication: hidden_size β num_attention_heads Γ head_dim. The output projection is still $W_O \in \mathbb{R}^{(H_Q \cdot d_\text{head}) \times d_\text{model}}$, which for Qwen3-8B is $(64 \times 128) \times 4096 = 8192 \times 4096$.
Pitfalls:
- Kernel reshape bugs. Flash Attention, custom CUDA, and Triton kernels frequently reshape $[B, T, H \cdot d_h]$ β $[B, H, T, d_h]$. If your kernel assumes $d_h = d_\text{model} / H$, it produces wrong strides silently on Qwen3.
- RoPE buffer mismatch. Sinusoidal positional buffers are pre-computed for
head_dimfrequencies. A buffer computed withhead_dim=64applied to tensors ofhead_dim=128reads garbage for the upper half. - KV cache pre-allocation. If you pre-allocate a cache of shape $[L, B, T, H_{KV}, d_h]$ with $d_h$ derived from the formula, Qwen3 cache will be 2Γ too small on affected model sizes.
- Always read
head_dimfrom config first; fall back to the formula only if absent.
1
2
# Correct pattern:
head_dim = getattr(config, "head_dim", config.hidden_size // config.num_attention_heads)
intermediate_size
1
{ "intermediate_size": 14336 }
The inner dimension of the feed-forward (FFN) / MLP block $d_{\text{ffn}}$.
For SwiGLU (LLaMA/Mistral/Qwen style): \(\text{FFN}(x) = \bigl(\sigma(x W_{\text{gate}}) \odot x W_{\text{up}}\bigr) W_{\text{down}}\) \(W_{\text{gate}}, W_{\text{up}} \in \mathbb{R}^{d \times d_{\text{ffn}}}, \quad W_{\text{down}} \in \mathbb{R}^{d_{\text{ffn}} \times d}\)
Parameter count per layer: $3 \cdot d \cdot d_{\text{ffn}}$ (gate, up, down)
Common ratios $d_{\text{ffn}} / d$:
| Model family | Ratio | Notes |
|---|---|---|
| LLaMA 3 | ~3.5Γ | 4096 β 14336 |
| Qwen3 | ~5.3Γ | 4096 β 21888 (8B) |
| Gemma 2 | ~3.8Γ | 3584 β 14336 |
| GPT-4 style | 4Γ | Classic MLP, not SwiGLU |
Pitfalls:
- In MoE models each expert has its own
intermediate_size. Routing logic selects $k$ experts; total activated FFN params = $k \times d \times d_{\text{ffn,expert}}$. The effective ratio is much lower than dense. intermediate_sizeis sometimes not a multiple of 64/128. Padding to hardware-friendly sizes in custom kernels must preserve output equivalence.- Some models list a separate
moe_intermediate_size(DeepSeek, Qwen2-MoE) for the expert FFNs vs. a sharedintermediate_sizefor dense layers.
hidden_act
1
{ "hidden_act": "silu" }
The activation function used in the FFN block.
| Value | Formula | Used in |
|---|---|---|
"silu" / "swish" |
$x \cdot \sigma(x)$ | LLaMA, Mistral, Qwen, Phi |
"gelu" |
$x \cdot \Phi(x)$ | BERT, GPT-2, older models |
"gelu_new" |
OpenAI tanh approx. | GPT-2 family |
"gelu_fast" |
Sigmoid approx. | DistilBERT |
"quick_gelu" |
$x \cdot \sigma(1.702 x)$ | CLIP, vision encoders |
"relu" |
$\max(0, x)$ | Legacy, rare |
"geglu" |
$\text{GELU}(W_g x) \odot W_u x$ | T5 v1.1 |
Pitfalls:
- SwiGLU models (using
"silu") always have three FFN weight matrices (gate, up, down), not two. A kernel written for ReLU-style FFN (two matrices) will silently skip the gate projection. "gelu"vs"gelu_new"produces numerically different outputs. This matters when comparing perplexity across implementations.- Vision encoders embedded inside VLMs (e.g., Qwen3-VLβs ViT) often use
"quick_gelu"while the language model uses"silu". Make sure you apply the correct activation per sub-model.
initializer_range
1
{ "initializer_range": 0.02 }
Standard deviation of the normal distribution used for weight initialization (Xavier/truncated normal).
Pitfalls:
- This field only matters at training start. If you load pretrained weights, it is ignored. Never tune inference behaviour based on this field.
- When you add new layers to a pretrained model (e.g. adapter modules), you should use this value to initialise new weights consistently.
Positional Encoding Parameters
max_position_embeddings
1
{ "max_position_embeddings": 131072 }
The maximum sequence length the model was trained to handle. Sets the size of any learned position embedding table (for models with absolute PE). For RoPE models it is used to pre-compute frequency buffers.
Pitfalls:
- This is a training-time parameter. At inference you can often go beyond it using RoPE scaling tricks (see
rope_scaling), but without adjustments, attention scores degrade beyond this length. - For RoPE models, exceeding
max_position_embeddingswithoutrope_scalingcauses positions to wrap (the sinusoidal frequencies were only ever computed for positions 0 β max-1). In practice, perplexity spikes sharply. - Many models advertise a context window larger than their training context via
rope_scaling. Always check both fields.
rope_theta
1
{ "rope_theta": 500000.0 }
The base frequency $\theta$ for Rotary Position Embedding (RoPE). Controls the wavelengths of the sinusoidal position signals:
\[\theta_i = \theta^{-2i/d_{\text{head}}}, \quad i = 0, 1, \ldots, \tfrac{d_{\text{head}}}{2}-1\]Higher $\theta$ β longer wavelengths β better long-context generalisation.
| Model | rope_theta |
Context | Notes |
|---|---|---|---|
| LLaMA 2 | 10,000 | 4K | Original |
| LLaMA 3 | 500,000 | 128K | Extended |
| Mistral v0.1 | 10,000 | 8K | Original |
| Qwen3 | 1,000,000 | 32Kβ131K | Very long |
| DeepSeek-V3 | 10,000 | 128K | Uses YaRN |
Pitfalls:
- RoPE buffers (
cos_cache,sin_cache) are pre-computed fromrope_theta. If you cache these at model load time and then fine-tune with a differentrope_theta, you must invalidate and recompute the buffers. - A higher
rope_thetadoes NOT automatically enable longer context β you must also increasemax_position_embeddingsand possibly addrope_scaling. - Custom kernels that inline the $\theta_i$ constants lose the ability to adjust for different model families.
rope_scaling
1
2
3
4
5
6
7
8
9
{
"rope_scaling": {
"factor": 8.0,
"low_freq_factor": 1.0,
"high_freq_factor": 4.0,
"original_max_position_embeddings": 8192,
"rope_type": "llama3"
}
}
Optional dictionary controlling how RoPE frequencies are scaled to extend context beyond original_max_position_embeddings.
Common types:
rope_type / type |
Algorithm | Used in |
|---|---|---|
"linear" |
Divide all frequencies by factor |
Simple extension |
"dynamic" |
Scale factor grows with input length | LongRoPE |
"llama3" |
Interpolate low-freq, extrapolate high-freq | LLaMA 3 |
"yarn" |
NTK + attention temperature correction | Mistral, DeepSeek |
"longrope" |
Per-layer scaling factors | Phi-3 128K |
"mrope" |
Multi-modal RoPE β splits head_dim into temporal/height/width sections | Qwen3-VL |
Pitfalls:
- Every
rope_typerequires different code paths. A kernel that implements only"linear"scaling will silently apply wrong positional encodings to LLaMA 3 ("llama3") weights. "mrope"(Qwen3-VL) splits each headβs frequencies into three sections:mrope_section = [s_t, s_h, s_w]where $s_t + s_h + s_w = d_{\text{head}}/2$. Applying scalar RoPE to these will corrupt the spatial position encoding of image patches.- When
rope_scalingis absent, no scaling is applied; positions beyondmax_position_embeddingsare out-of-distribution. - LLaMA 3 uses
"llama3"type with two thresholds (low_freq_factor,high_freq_factor) that interpolate vs. extrapolate different frequency bands. The exact transition formula is not a simple linear interpolation.
Normalization Parameters
rms_norm_eps
1
{ "rms_norm_eps": 1e-05 }
The epsilon added to the RMS in RMSNorm to prevent division by zero:
\[\text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{d}\sum_i x_i^2 + \epsilon}} \cdot \gamma\]Used by LLaMA, Mistral, Qwen, Falcon, and most modern models.
Pitfalls:
- Default value varies:
1e-5(LLaMA),1e-6(Qwen),1e-8(Gemma 2). Using the wrong epsilon in a custom kernel causes numerically different outputs β almost always within floating-point noise, but can accumulate across layers. - In FP16 mixed-precision training, a too-small epsilon can cause numerical instability.
1e-6is safer than1e-8in FP16. - Different from
layer_norm_epsused in LayerNorm (mean-centred).
layer_norm_eps
1
{ "layer_norm_eps": 1e-05 }
Epsilon for LayerNorm (as opposed to RMSNorm):
\[\text{LayerNorm}(x) = \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}} \cdot \gamma + \beta\]Used by BERT, GPT-2, T5, encoder-only models, and vision encoders inside VLMs.
Pitfalls:
- Some VLM models (e.g. LLaVA) have a vision encoder with
layer_norm_epsand a language decoder withrms_norm_epsat different magnitudes. Mixing them produces subtle accuracy loss. - Certain models use both fields (the config has
rms_norm_epsfor transformer blocks andlayer_norm_epsfor adapters / projection layers).
Attention Configuration
attention_bias
1
{ "attention_bias": false }
Whether the Q, K, V, and output projection matrices include a bias term.
| Value | Memory impact | Used in |
|---|---|---|
false |
No extra params | LLaMA, Mistral, Qwen (most modern) |
true |
+4 Γ d per layer | GPT-2, older BERT-style, Falcon |
Pitfalls:
- A kernel that skips the bias term when
attention_bias=trueproduces incorrect outputs. This is especially insidious with models that have bias only on certain projections (Falcon has bias on QKV but not output). - Some models have a separate
qkv_biasfield (Qwen2-VL, CLIP ViT). Check the model class source if unsure.
mlp_bias
1
{ "mlp_bias": false }
Whether the FFN projections (gate, up, down) include bias. Almost always false in modern models.
Pitfalls:
- Same as
attention_biasβ missing a bias addition shifts outputs. - Phi-3 includes biases in QKV but not FFN (
attention_bias=falsein some versions). Always check the config, never assume.
attention_dropout
1
{ "attention_dropout": 0.0 }
Dropout probability applied to attention weights during training. At inference this is always 0 regardless of the field.
Pitfalls:
- This is only active when
model.train()mode is set. If your inference code accidentally leaves the model in train mode, attention dropout fires and produces stochastic (non-deterministic) outputs. - Flash Attention passes
dropout_pas a float; always pass0.0at inference even if the config has a non-zero value β check whether your inference wrapper does this correctly.
Sliding-Window Attention
1
2
3
4
5
{
"use_sliding_window": true,
"sliding_window": 4096,
"max_window_layers": 28
}
Some models (Mistral, Qwen2/3, Gemma 2) use a mix of local sliding-window attention (SWA) and global attention layers.
| Field | Meaning |
|---|---|
use_sliding_window |
Whether SWA is enabled at all |
sliding_window |
Number of tokens in the local window ($w$) |
max_window_layers |
Layers above this index use full global attention |
How it works:
1
2
Layers 0 β¦ max_window_layers-1 : SWA, window = sliding_window
Layers max_window_layers β¦ L-1 : Full attention
Memory formula (SWA layers):
\[\text{KV cache (SWA)} = 2 \cdot L_{\text{SWA}} \cdot B \cdot w \cdot H_{KV} \cdot d_h \cdot \text{bytes}\]Pitfalls:
- A kernel implementing only full attention and applied to an SWA model will produce correct but suboptimal outputs (the model still works because SWA is a subset of full attention, but memory usage explodes).
- Applying SWA to a model without it (or vice versa) silently produces wrong outputs. Always gate on
use_sliding_window. - Gemma 2 interleaves: every other layer alternates SWA/global.
max_window_layersdoesnβt apply; check the model source for the interleaving logic. - During KV cache eviction/paging (PagedAttention), the eviction policy differs for SWA layers β you can evict tokens outside the window freely. A paging system unaware of
sliding_windowmay under-evict.
Mixture of Experts Parameters
1
2
3
4
5
6
{
"num_local_experts": 64,
"num_experts_per_tok": 6,
"router_aux_loss_coef": 0.001,
"num_shared_experts": 2
}
Field naming is inconsistent across model families:
| Field name | Alternatives | Meaning |
|---|---|---|
num_local_experts |
num_experts, moe_num_experts |
Total experts per MoE layer |
num_experts_per_tok |
top_k, num_selected_experts |
Experts activated per token |
router_aux_loss_coef |
aux_loss_alpha |
Load balancing loss weight |
num_shared_experts |
β | Always-active experts (DeepSeek) |
first_k_dense_replace |
num_dense_layers |
How many leading layers are dense (not MoE) |
moe_layer_freq |
β | Every $n$-th layer is MoE (some configs) |
Formulae:
\(\text{Active params per token} = P_{\text{dense}} + k \cdot P_{\text{expert}}\) \(\text{Total params} = P_{\text{dense}} + N_e \cdot P_{\text{expert}}\)
Pitfalls:
num_experts_per_tok*num_local_expertsrouting logits are computed for every token. Kernels that pre-allocate expert buffers assuming all experts are active will allocate $N_e / k$ times too much memory.- Load balancing: if
router_aux_loss_coefis absent or 0 at training, experts will not be balanced, causing inference throughput collapse on certain tokens. num_shared_experts(DeepSeek-V3, Qwen2-MoE) refers to experts that are always active alongside the top-k selected experts. These must be summed into the active parameter count. Missing them underestimates FLOPs by ~20%.- Expert weights are usually stored contiguously for efficient dispatch. Custom kernels must match the layout expected by the routing code.
Special Token IDs
1
2
3
4
5
{
"bos_token_id": 1,
"eos_token_id": [2, 128009],
"pad_token_id": 0
}
| Field | Role |
|---|---|
bos_token_id |
Prepended to every sequence (beginning of sentence) |
eos_token_id |
Signals end of generation; can be a list of valid stop tokens |
pad_token_id |
Used to pad shorter sequences in a batch |
Pitfalls:
eos_token_idcan be a list (LLaMA 3:[128001, 128008, 128009]). A generation loop that only checks equality to a single integer will not stop on alternate EOS tokens, causing runaway generation.- When
pad_token_idis unset (null/absent) and you run batched inference, HuggingFace defaults toeos_token_id. This means padded positions will receive the EOS token β which is often fine for inference but can corrupt training batches if not masked. - For Qwen3/ChatML models, the BOS token is the same as the
<|im_start|>token. Omitting it causes the model to lose the chat template structure. - KV cache implementations that detect EOS and stop early must check against the full list, not just index 0.
Generation & Runtime Parameters
use_cache
1
{ "use_cache": true }
Whether to return and use past key-value states (the KV cache) during generation. Should be true for inference, false during training (saves memory, avoids caching overhead during forward-only passes).
Pitfalls:
- Leaving
use_cache=trueduring training wastes memory; disabling it at inference causes $O(T^2)$ computation per step instead of $O(T)$. - Gradient checkpointing (
use_reentrant=True) is incompatible withuse_cache=Truein some model classes. If you see recomputation errors, disable caching.
torch_dtype
1
{ "torch_dtype": "bfloat16" }
Preferred weight dtype. Informational β from_pretrained uses this as default but it can be overridden.
| Value | Range | Notes |
|---|---|---|
"float32" |
Β±3.4Γ10Β³βΈ | Default; safe but 2Γ memory |
"bfloat16" |
Β±3.4Γ10Β³βΈ | Modern standard; same range as FP32 |
"float16" |
Β±65504 | Overflow risk with large logits/activations |
"float8_e4m3fn" |
β | Emerging; requires explicit handling |
Pitfalls:
- FP16 overflows when logits or hidden states grow large. BF16 is strictly preferred for transformer training/inference. If you load a BF16 model in FP16 and the model has large embeddings (vocab_size > 100K), you may see NaNs in the logit layer.
torch_dtypein config does not enforce the dtype atfrom_pretrainedtime unless you passtorch_dtype="auto".
transformers_version
1
{ "transformers_version": "4.46.1" }
The library version used when the model was saved.
Pitfalls:
- A model saved with transformers 4.46 may not load correctly under 4.40 if it uses a feature added in-between (e.g. Qwen3βs
head_dimfield was added in a later version). Always use β₯ the saved version. - The config schema is not versioned separately from the library. Breaking changes land silently.
tie_word_embeddings
1
{ "tie_word_embeddings": false }
When true, the input embedding matrix $W_e$ and the LM head $W_{\text{lm}}$ share the same tensor.
Parameter savings:
\(\text{saved} = V \times d_{\text{model}} \times \text{bytes}\)
For LLaMA-3 8B: $128256 \times 4096 \times 2 \approx 1$ GB saved.
Pitfalls:
- With
tie_word_embeddings=true, optimizers maintain a single set of gradients for the shared matrix. If you apply separate learning rates to embeddings vs. head, the effective LR is the sum β usually unintentional. - When quantising, tying embeddings means quantising the embedding table also quantises the LM head. INT8 embeddings are fine; INT4 LM heads cause significant quality loss. Most quantisation libraries detect and skip tying for INT4.
- If you
resize_token_embeddings()but the config still hastie_word_embeddings=true, you must call it on the model (not just the weight), otherwise the two tensors fall out of sync.
VLM Config β Qwen3-VL and LLaVA Style
Vision-Language Models nest multiple sub-configs inside a single config.json. Understanding this structure is essential for both loading and kernel development.
General VLM Config Structure
1
2
3
4
5
6
7
8
9
10
{
"model_type": "qwen3_vl",
"architectures": ["Qwen3VLForConditionalGeneration"],
"text_config": { ... },
"vision_config": { ... },
"image_token_id": 151655,
"video_token_id": 151656,
"min_pixels": 200704,
"max_pixels": 1003520
}
A VLM config has:
- Root-level fields β model type, special tokens shared across modalities
text_configβ the language model config (same fields as a standalone LLM)vision_configβ the visual encoder config
The text_config.head_dim caveat applies identically inside VLMs. Qwen3-VL-7B has text_config.head_dim = 128 with text_config.num_attention_heads = 28 and text_config.hidden_size = 3584, giving naΓ―ve head_dim = 128 β correct only by coincidence for this size.
Vision Config Fields (Qwen3-VL / ViT-style)
| Field | Typical value | Meaning | Pitfalls |
|---|---|---|---|
depth |
32 | Number of ViT transformer layers | Different from LLMβs num_hidden_layers β do not confuse |
embed_dim |
1280 | Vision hidden size $d_\text{vis}$ | Must be projected to text_config.hidden_size via a connector |
num_heads |
16 | Vision attention heads | head_dim = embed_dim / num_heads (standard convention here, no override) |
image_size |
448 | Max input image side length | Dynamic resolution: actual size varies; this is just an upper bound |
patch_size |
14 | Spatial patch size in pixels | Must divide evenly into image dimensions after padding |
temporal_patch_size |
2 | Video frame patch grouping | Pairs of frames β single temporal token; odd frame counts need padding |
spatial_merge_size |
2 | Spatial pooling factor after encoding | Reduces token count by merge_sizeΒ²; affects the number of image tokens in the LLM context |
in_channels |
3 | Input image channels (RGB) | Some models support depth (4 channels) or infrared; changing this requires re-training |
mlp_ratio |
4 | FFN expansion in ViT | d_ffn = embed_dim Γ mlp_ratio = 5120 |
hidden_act |
"quick_gelu" |
ViT activation | Different from LLM activation (silu)! |
out_hidden_size |
3584 | Projection output size | Must equal text_config.hidden_size |
fullatt_block_indexes |
[7,15,23,31] |
ViT layers using full attention | Other layers use window attention β affects memory |
Image Tokenisation Formula
\[N_\text{patches} = \frac{H}{p} \cdot \frac{W}{p}\] \[N_\text{tokens} = \frac{N_\text{patches}}{m^2} = \frac{H \cdot W}{p^2 \cdot m^2}\]where $p$ = patch_size, $m$ = spatial_merge_size.
For a 448Γ448 image with $p=14$, $m=2$: \(N_\text{tokens} = \frac{448 \times 448}{14^2 \times 4} = \frac{200704}{784} = 256\)
For a 1024Γ1024 image: \(N_\text{tokens} = \frac{1024 \times 1024}{784} \approx 1337\)
Pitfalls (Vision Config):
- Dynamic resolution. Unlike text (fixed vocabulary), image token count varies per image. Pre-allocating a fixed KV cache for image tokens will break when a large image produces more tokens than expected.
- Connector shape mismatch. The projection MLP connecting the ViT to the LLM takes
embed_dimas input and outputstext_config.hidden_size. A shape error here is common when mixing different ViT and LLM sizes. fullatt_block_indexesmust be honoured. ViT layers not in this list use window attention. Applying full attention everywhere is functionally correct but wastes memory; applying window attention to full-attention layers corrupts quality.- Temporal patch size and video. For video input, frames are grouped into pairs (
temporal_patch_size=2). An odd number of frames must be handled (usually by duplicating the last frame). Missing this causes a shape error in the patching step. image_token_idandvideo_token_idoccupy specific positions in the vocabulary. These IDs must not collide with BOS/EOS or pad tokens. When extending the vocabulary for a new language, verify there are no collisions.min_pixels/max_pixelsconstrain the dynamic resolution selection. The processor will resize/tile the image to satisfy these bounds. Settingmax_pixelstoo low degrades OCR quality; too high inflates context length. These live at the root level, not insidevision_config.
LLaVA / LLaVA-Next Style VLM Config
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
{
"model_type": "llava",
"vision_config": {
"model_type": "clip_vision_model",
"hidden_size": 1024,
"intermediate_size": 4096,
"num_hidden_layers": 24,
"num_attention_heads": 16,
"image_size": 336,
"patch_size": 14
},
"text_config": { "model_type": "llama", ... },
"projector_hidden_act": "gelu",
"vision_feature_layer": -2,
"vision_feature_select_strategy": "default",
"image_grid_pinpoints": [[336, 672], [672, 336], [672, 672], [1008, 336], [336, 1008]]
}
| Field | Meaning | Pitfalls |
|---|---|---|
vision_feature_layer |
Which ViT layerβs output to use | -2 means penultimate; -1 is the final CLS-normalised output; mixing these changes the feature distribution |
vision_feature_select_strategy |
"default" (patch tokens only) or "full" (patch + CLS) |
"full" adds 1 extra token per image |
projector_hidden_act |
Activation in the MLP projector | "gelu" is common; using the wrong one mismatches pretrained projector weights |
image_grid_pinpoints |
Allowed tiled resolutions for high-res images | LLaVA-Next divides the image into multiple tiles; the number of image tokens = tiles Γ patches_per_tile |
VLA Config β Alpamayo and RT-2 Style
Vision-Language-Action (VLA) models extend VLMs with action prediction. Their configs add fields controlling the action space, observation modalities, and temporal dynamics.
What Makes a VLA Different
1
2
Input: [image tokens | text/instruction tokens | state tokens]
Output: [text tokens | action tokens]
Action tokens may be discrete (tokenised from the continuous action space via binning) or continuous (predicted via a regression head).
Alpamayo-Style Config Fields
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
{
"model_type": "alpamayo",
"architectures": ["AlpamayoForConditionalGeneration"],
"text_config": { ... },
"vision_config": { ... },
"action_config": {
"action_dim": 7,
"action_chunk_size": 16,
"action_vocab_size": 256,
"action_token_offset": 152064,
"num_action_layers": 4,
"action_hidden_size": 512,
"control_frequency_hz": 50,
"state_dim": 14,
"state_token_id": 151657
}
}
| Field | Meaning | Pitfalls |
|---|---|---|
action_dim |
Degrees of freedom of the action (e.g. 7 = 6-DoF arm + gripper) | Must match the robotβs URDF/API. A mismatch crashes the control loop silently (motor ignores extra dims) |
action_chunk_size |
Number of future timesteps predicted at once | Higher = smoother but higher latency per inference call; must align with the trajectory rollout buffer |
action_vocab_size |
Number of discrete bins per action dimension | Finer bins β higher resolution but more tokens; vocab must fit within text_config.vocab_size allocated range |
action_token_offset |
First ID in the vocabulary reserved for action tokens | Must not overlap with text/image tokens; validate against vocab_size |
num_action_layers |
Dedicated action-decoding transformer layers appended after the base LLM | Additional memory per inference; must be allocated separately from LLM KV cache |
action_hidden_size |
Hidden size of the action decoder | May differ from text_config.hidden_size; requires a projection bridge |
control_frequency_hz |
Robot control loop frequency | Inference latency must be < 1/frequency. E.g. 50 Hz β < 20 ms per inference. Batch size must be 1 |
state_dim |
Proprioceptive state vector dimension (joint angles, velocities, etc.) | Usually normalised per-dataset; normalisation stats are NOT stored in config.json β they live in a separate stats file |
state_token_id |
Special token marking the proprioceptive state in the sequence | Must not collide with image_token_id or video_token_id |
RT-2 / OpenVLA Style Notes
RT-2 (and OpenVLA which is based on LLaVA) discretises actions differently:
1
2
3
4
5
6
7
8
9
10
{
"norm_stats": {
"action": {
"mean": [...],
"std": [...],
"q01": [...],
"q99": [...]
}
}
}
Key distinction from Alpamayo:
- Actions are tokenised as text tokens (e.g. bin indices formatted as strings)
- The action vocabulary is carved out from the top of the existing text vocabulary:
action_token_id[i] = vocab_size - (action_vocab_size * action_dim) + i - No separate
action_configblock β actions are treated as structured text
Pitfalls (VLA):
- Action normalisation stats are separate.
config.jsondoes not store the per-dataset action mean/std. If you load a VLA for a different robot without updating normalisation stats, the predicted actions are physically wrong even if the model runs without error. - Control frequency vs. inference budget. At 50 Hz you have 20 ms. Typical 7B VLMs take 30β100 ms on a single A100. VLAs run smaller LLM backbones (3Bβ7B) and quantise aggressively for real-time use.
- Action chunking increases latency but reduces jitter.
action_chunk_size=16means you run inference every 16 timesteps at1/(16 Γ control_frequency_hz)call frequency. The robot executes pre-planned chunks; latency spikes between chunks cause jerky motion. - State dim normalisation must match the training distribution. Reusing a VLA on a robot with different joint limits without renormalising will produce poor actions even if the policy is otherwise applicable.
- Token collision. If
action_token_offset + action_vocab_size * action_dim > vocab_size, action tokens collide with the EOS/BOS region. This is not validated at load time. - Batch size must be 1 at inference for real-robot deployment (each inference is tied to the current observation state). VLAs are not meant for batched generation.
Cross-Model Comparison Table
| Param | LLaMA 3.1 8B | Mistral 7B v0.3 | Qwen3-8B | Gemma 2 9B | DeepSeek-V3 (MoE) | Qwen3-VL-7B (LLM part) |
|---|---|---|---|---|---|---|
hidden_size |
4096 | 4096 | 4096 | 3584 | 7168 | 3584 |
num_hidden_layers |
32 | 32 | 36 | 42 | 61 | 28 |
num_attention_heads |
32 | 32 | 64 | 16 | 128 | 28 |
num_key_value_heads |
8 | 8 | 8 | 8 | 128 | 4 |
head_dim |
(implicit 128) | (implicit 128) | 128 β override | 256 | 128 | (implicit 128) |
intermediate_size |
14336 | 14336 | 22016 | 14336 | 18432 | 18944 |
rope_theta |
500,000 | 1,000,000 | 1,000,000 | 10,000 | 10,000 | 1,000,000 |
vocab_size |
128,256 | 32,768 | 151,936 | 256,000 | 129,280 | 151,936 |
rms_norm_eps |
1e-5 | 1e-5 | 1e-6 | 1e-6 | 1e-6 | 1e-6 |
sliding_window |
β | 4096 | β | 4096 (odd) | β | β |
num_local_experts |
β | β | β | β | 256 | β |
num_experts_per_tok |
β | β | β | β | 8 | β |
tie_word_embeddings |
false | false | false | true | false | false |
Kernel Development Pitfall Checklist
Use this checklist when writing or porting a CUDA/Triton kernel for a new model family:
- head_dim: Always read from
config.head_dimif present; fall back tohidden_size // num_attention_heads. Never hard-code. - GQA reshape: KV head expansion ($H_Q / H_{KV}$ repetitions) must use
num_key_value_heads, notnum_attention_heads. - RoPE type: Check
rope_scaling.type/rope_typefield. Implement all types your model family uses, or assert on unsupported types. - mrope sections: For VLMs with
rope_type = "mrope", frequencies are split per modality dimension β apply per-section frequencies, not a uniform vector. - Sliding window: Check
use_sliding_windowper layer index vsmax_window_layers. Gemma 2 uses interleaved pattern (not threshold-based). - Activation function: SwiGLU uses three weight matrices. Check
hidden_actbefore assuming two-matrix FFN. - MoE layer type: Not all layers are MoE. Use
first_k_dense_replace/moe_layer_freqto determine which layers have expert routing. - vocab_size padding: Embedding tables may be padded beyond actual vocabulary. Mask logits beyond true vocab_size before sampling.
- Special token collisions: For VLMs/VLAs, verify
image_token_id,video_token_id,state_token_idare all withinvocab_sizeand non-overlapping. - Dynamic image token count: Never pre-allocate a fixed number of image tokens per image. Compute dynamically from the formula $N = H \cdot W / (p^2 \cdot m^2)$.
- EOS as list: Generation stop condition must check all IDs in
eos_token_id(can be a list). - dtype mismatch: Ensure KV cache dtype matches
torch_dtype. BF16 cache with FP32 accumulation is intentional (flash attention); FP16 cache with BF16 activations is not. - Tensor-parallel divisibility: Verify
num_attention_heads % tp_degree == 0andnum_key_value_heads >= tp_degreeor that KV heads will be replicated.
β Previous: Appendix B β Tokenization Deep Dive
Last updated: April 2026