pub struct MtpWeights {Show 17 fields
pub pre_fc_norm_embedding: DenseWeight,
pub pre_fc_norm_hidden: DenseWeight,
pub fc: DenseWeight,
pub input_layernorm: DenseWeight,
pub q_proj: DenseWeight,
pub k_proj: DenseWeight,
pub v_proj: DenseWeight,
pub o_proj: DenseWeight,
pub q_norm: DenseWeight,
pub k_norm: DenseWeight,
pub post_attn_layernorm: DenseWeight,
pub moe_gate: DenseWeight,
pub shared_expert: DenseExpertWeight,
pub shared_expert_gate: DenseWeight,
pub experts: Vec<DenseExpertWeight>,
pub dense_ffn: Option<DenseExpertWeight>,
pub norm: DenseWeight,
}Expand description
MTP (Multi-Token Prediction) head weights (all BF16 from safetensors).
Single decoder layer + concat projection. All projection weights are BF16 and get quantized to NVFP4 at load time by the weight loader.
Fields§
§pre_fc_norm_embedding: DenseWeightRMSNorm on token embedding before concat: [hidden_size] BF16.
RMSNorm on target hidden state before concat: [hidden_size] BF16.
fc: DenseWeightConcat projection: [hidden_size, 2*hidden_size] BF16.
input_layernorm: DenseWeightInput layernorm for the attention layer: [hidden_size] BF16.
q_proj: DenseWeightAttention projections (all BF16).
k_proj: DenseWeight§v_proj: DenseWeight§o_proj: DenseWeight§q_norm: DenseWeight§k_norm: DenseWeight§post_attn_layernorm: DenseWeightPost-attention layernorm: [hidden_size] BF16.
moe_gate: DenseWeightMoE router gate: [num_experts, hidden_size] BF16.
NULL when dense_ffn is Some (dense FFN MTP head).
Shared expert (BF16). NULL fields when dense_ffn is Some.
Shared expert gate: [1, hidden_size] BF16.
NULL when dense_ffn is Some.
experts: Vec<DenseExpertWeight>Per-expert weights (512 experts, BF16).
Empty when dense_ffn is Some.
dense_ffn: Option<DenseExpertWeight>Dense FFN triple (gate_proj, up_proj, down_proj) — used by MTP
heads bundled with dense (non-MoE) FP8 checkpoints, e.g.
Qwen/Qwen3.6-27B-FP8. When Some, the MoE fields above are unused
and the forward path takes the dense MLP shortcut.
norm: DenseWeightFinal output RMSNorm: [hidden_size] BF16.