IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

kpool_tail_update_kernel

def kpool_tail_update_kernel[dtype: DType, TailLayoutType: TensorLayout, tail_origin: MutOrigin, OutLayoutType: TensorLayout, out_origin: MutOrigin, ClosedLayoutType: TensorLayout, closed_origin: MutOrigin, KLayoutType: TensorLayout, k_origin: ImmOrigin, GateLayoutType: TensorLayout, gate_origin: ImmOrigin, ApeLayoutType: TensorLayout, ape_origin: ImmOrigin, PosLayoutType: TensorLayout, pos_origin: ImmOrigin, SlotLayoutType: TensorLayout, slot_origin: ImmOrigin, TailEngine: TensorEngine, OutEngine: TensorEngine, ClosedEngine: TensorEngine, KEngine: TensorEngine, GateEngine: TensorEngine, ApeEngine: TensorEngine, PosEngine: TensorEngine, SlotEngine: TensorEngine, head_dim: Int, kpool: Int, next_n: Int = Int(1)](tail: TileTensor[dtype, TailLayoutType, tail_origin, Engine=TailEngine], pooled: TileTensor[dtype, OutLayoutType, out_origin, Engine=OutEngine], closed_pool: TileTensor[.int32, ClosedLayoutType, closed_origin, Engine=ClosedEngine], k: TileTensor[dtype, KLayoutType, k_origin, Engine=KEngine], gate: TileTensor[dtype, GateLayoutType, gate_origin, Engine=GateEngine], ape: TileTensor[.float32, ApeLayoutType, ape_origin, Engine=ApeEngine], positions: TileTensor[.int32, PosLayoutType, pos_origin, Engine=PosEngine], slot_idx: TileTensor[.uint32, SlotLayoutType, slot_origin, Engine=SlotEngine], num_requests: Int32)

Stashes a request's new tokens, and closes each pool as it fills.

A decoded token cannot be pooled on arrival, because its pool-mates arrived on earlier steps and have left the batch. Each request keeps its in-progress pool in tail, a ring of kpool slots addressed by position % kpool.

A speculative step appends next_n tokens at once, so several pools can close in one call.

The ring is indexed by slot_idx[r], not by r. A batch reorders between steps, so row r is not always the same request.

Every real token stashes, whether or not it closes a pool.

Rejected speculative tokens are the caller's problem. The ring holds no pointer to rewind, so a rejected token that has already stashed stays.

Parameters:

Args:

Was this page helpful?