For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
indexer_score_ragged_paged
def indexer_score_ragged_paged[cache_t: KVCacheT, q_type: DType, //, target: StringSpan[ImmStaticOrigin], num_heads: Int](output: LayoutTensor[.float32, element_layout=output.element_layout, layout_int_type=output.layout_int_type, linear_idx_type=output.linear_idx_type, masked=output.masked, alignment=output.alignment], q: LayoutTensor[q_type, element_layout=q.element_layout, layout_int_type=q.layout_int_type, linear_idx_type=q.linear_idx_type, masked=q.masked, alignment=q.alignment], weights: LayoutTensor[.float32, element_layout=weights.element_layout, layout_int_type=weights.layout_int_type, linear_idx_type=weights.linear_idx_type, masked=weights.masked, alignment=weights.alignment], input_row_offsets: LayoutTensor[.uint32, element_layout=input_row_offsets.element_layout, layout_int_type=input_row_offsets.layout_int_type, linear_idx_type=input_row_offsets.linear_idx_type, masked=input_row_offsets.masked, alignment=input_row_offsets.alignment], base: LayoutTensor[.int32, element_layout=base.element_layout, layout_int_type=base.layout_int_type, linear_idx_type=base.linear_idx_type, masked=base.masked, alignment=base.alignment], cutoff: LayoutTensor[.int32, element_layout=cutoff.element_layout, layout_int_type=cutoff.layout_int_type, linear_idx_type=cutoff.linear_idx_type, masked=cutoff.masked, alignment=cutoff.alignment], cache: cache_t, ctx: DeviceContext)
Scores every query row against its live compressed entries.
Parameters:
- cache_t (
KVCacheT): The compressed leaf's key cache type (inferred); one head, paged by entry throughslots_per_page. - q_type (
DType): Query element type (inferred). - target (
StringSpan[ImmStaticOrigin]): Compilation target string, selects the CPU or GPU path. - num_heads (
Int): Query heads summed into each score.
Args:
- output (
LayoutTensor[.float32, element_layout=output.element_layout, layout_int_type=output.layout_int_type, linear_idx_type=output.linear_idx_type, masked=output.masked, alignment=output.alignment]):[num_rows, num_cand]scores,num_candeven; columnsnum_cand // 2on are the fresh windows. - q (
LayoutTensor[q_type, element_layout=q.element_layout, layout_int_type=q.layout_int_type, linear_idx_type=q.linear_idx_type, masked=q.masked, alignment=q.alignment]):[num_rows, num_heads, head_dim], the last axis contiguous. - weights (
LayoutTensor[.float32, element_layout=weights.element_layout, layout_int_type=weights.layout_int_type, linear_idx_type=weights.linear_idx_type, masked=weights.masked, alignment=weights.alignment]):[num_rows, num_heads]per-head weights, contiguous. - input_row_offsets (
LayoutTensor[.uint32, element_layout=input_row_offsets.element_layout, layout_int_type=input_row_offsets.layout_int_type, linear_idx_type=input_row_offsets.linear_idx_type, masked=input_row_offsets.masked, alignment=input_row_offsets.alignment]):[batch + 1]ragged row offsets. - base (
LayoutTensor[.int32, element_layout=base.element_layout, layout_int_type=base.layout_int_type, linear_idx_type=base.linear_idx_type, masked=base.masked, alignment=base.alignment]):[batch]entries each request had closed before this chunk. - cutoff (
LayoutTensor[.int32, element_layout=cutoff.element_layout, layout_int_type=cutoff.layout_int_type, linear_idx_type=cutoff.linear_idx_type, masked=cutoff.masked, alignment=cutoff.alignment]):[num_rows]entries each query may see, counting from 0. - cache (
cache_t): This layer's compressed leaf. - ctx (
DeviceContext): Device context used to enqueue the GPU kernel.