pack_weight_sfb

Function pack_weight_sfb 

Source
pub fn pack_weight_sfb(
    scale_in: u64,
    scale_out: u64,
    n: u32,
    k: u32,
    src_n_major: bool,
    stream: u64,
) -> Result<()>
Expand description

Repack an Atlas E4M3 weight scale into the CUTLASS SM120 blockscaled SFB swizzle atom (tile_atom_to_shape_SFB, ue4m3) that the grouped collective reads. M-independent (the SFB atom depends only on N,K) so this runs once per expert at load. scale_out must hold the swizzled SFB region the grouped kernel consumes.

src_n_major selects the SOURCE layout: false = Atlas-transposed [K/16,N], true = checkpoint-native [N,K/16]. The N-major mode lets a checkpoint that already ships [N,K/16] scales (Laguna) build SFB without first materialising an Atlas-transposed copy. Output layout is identical either way.