For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo module
kv_cache_gather
Gathers rows of a paged KV cache by an explicit per-row slot list.
The read side of kv_cache_store_ragged: row r of slots belongs to
the request b with row_offsets[b] <= r < row_offsets[b + 1], and
output[r, j, :] is the cache row at slot slots[r, j] of that request,
resolved through the request's lookup table exactly as load resolves a
token index. A leaf whose buffer holds slots_per_page rows per page is
paged by that count, because load reads the page size off the block shape,
so its slot is an entry index rather than a token index.
This is a pure copy in the cache's own dtype, so the result is bit-identical to indexing the block buffer by hand, on every target.
Functions
-
kv_cache_gather_rows_ragged: Copiescacherows atslotsintooutput.