pub struct PrefillSlice<'a> {
pub prompt_tokens: &'a [u32],
pub seq: &'a mut SequenceState,
pub chunk_start: usize,
pub chunk_len: usize,
pub is_last_chunk: bool,
}Expand description
Per-stream input slice for batched prefill.
One of these per concurrent prefilling stream — prefill_batch_chunk and
mixed_forward_batch accept a &mut [PrefillSlice<'_>] and process all
streams’ chunks in a single forward pass. See Q12 in
/workspace/atlas-internal/qwen-refactor/notes.md for the bug this
fixes (concurrent prefills serialized through prefilling.first_mut()
in the scheduler, causing 5× asymmetric TTFT).
Fields§
§prompt_tokens: &'a [u32]Full prompt tokens for this stream.
seq: &'a mut SequenceStatePer-stream sequence state (KV blocks, SSM slot, etc.).
chunk_start: usizeToken offset into prompt_tokens where this chunk starts.
chunk_len: usizeNumber of tokens in this chunk.
is_last_chunk: boolWhether this is the final chunk for this stream (controls whether the model emits last-token logits for sampling).
Auto Trait Implementations§
impl<'a> Freeze for PrefillSlice<'a>
impl<'a> !RefUnwindSafe for PrefillSlice<'a>
impl<'a> Send for PrefillSlice<'a>
impl<'a> Sync for PrefillSlice<'a>
impl<'a> Unpin for PrefillSlice<'a>
impl<'a> !UnwindSafe for PrefillSlice<'a>
Blanket Implementations§
Source§impl<T> BorrowMut<T> for Twhere
T: ?Sized,
impl<T> BorrowMut<T> for Twhere
T: ?Sized,
Source§fn borrow_mut(&mut self) -> &mut T
fn borrow_mut(&mut self) -> &mut T
Mutably borrows from an owned value. Read more