For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
kv_cache_gather_rows_ragged
def kv_cache_gather_rows_ragged[cache_t: KVCacheT, dtype: DType, //, target: StringSpan[ImmStaticOrigin]](output: TileTensor[dtype, Engine=output.Engine, linear_idx_type=output.linear_idx_type], slots: TileTensor[.int32, Engine=slots.Engine, linear_idx_type=slots.linear_idx_type], row_offsets: TileTensor[.uint32, Engine=row_offsets.Engine, linear_idx_type=row_offsets.linear_idx_type], cache: cache_t, ctx: DeviceContext)
Copies cache rows at slots into output.
Parameters:
- cache_t (
KVCacheT): The key or value cache type (inferred); one head per slot. - dtype (
DType): The output element type (inferred); must be the cache's. - target (
StringSpan[ImmStaticOrigin]): Compilation target string, selects the CPU or GPU path.
Args:
- output (
TileTensor[dtype, Engine=output.Engine, linear_idx_type=output.linear_idx_type]):[num_rows, num_slots, head_dim], the last axis contiguous. - slots (
TileTensor[.int32, Engine=slots.Engine, linear_idx_type=slots.linear_idx_type]):[num_rows, num_slots]slot indices. Every slot must be non-negative and inside its request's allocated pages. - row_offsets (
TileTensor[.uint32, Engine=row_offsets.Engine, linear_idx_type=row_offsets.linear_idx_type]):[batch + 1]ragged offsets mapping each row ofslotsto a request. - cache (
cache_t): The key or value cache of one layer. - ctx (
DeviceContext): Device context used to enqueue the GPU kernel.