IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo module

kv_cache_gather

Gathers rows of a paged KV cache by an explicit per-row slot list.

The read side of kv_cache_store_ragged: row r of slots belongs to the request b with row_offsets[b] <= r < row_offsets[b + 1], and output[r, j, :] is the cache row at slot slots[r, j] of that request, resolved through the request's lookup table exactly as load resolves a token index. A leaf whose buffer holds slots_per_page rows per page is paged by that count, because load reads the page size off the block shape, so its slot is an entry index rather than a token index.

This is a pure copy in the cache's own dtype, so the result is bit-identical to indexing the block buffer by hand, on every target.

Functions​

Was this page helpful?