requires_single_chunk_prefill

Function requires_single_chunk_prefill 

Source
pub fn requires_single_chunk_prefill(
    model_type: &str,
    kv_lora_rank: usize,
) -> bool
Expand description

Must chunked prefill run as a SINGLE chunk for this model?

True only for models that reach the chunk-LOCAL MLA prefill in qwen3_attention/prefill.rs, which attends over the current chunk’s K/V alone — multi-chunk there silently corrupts attention output (Mistral-Small-4, 2026-05-01: 8 K collapses to “The\nThe…”).

🔴 kv_lora_rank > 0 is a PROXY for that kernel and glm5_next breaks it: GLM-5.3 is MLA (rank 512) but prefills through Glm5NextLayer::prefill, a per-token walk that attends the whole paged prefix at each absolute position — chunk boundaries are invisible to it. Answering true capped every GLM prompt at 2 × --max-prefill-tokens, because prefill_a_step splits the FIRST chunk at the cap regardless and this gate then made the remainder one unsplit chunk the buffer arena refused. ANOMALIES A61.