For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
RecurrentKVLeafRegion
RecurrentKVLeafRegion
class max.nn.kv_cache.RecurrentKVLeafRegion(leaf_id, group_id, bytes_per_page, region, *, row_bytes=1)
Bases: KVLeafRegion
A leaf addressed by row: one block holds one request’s whole state.
-
Parameters:
-
- leaf_id (str)
- group_id (KVCacheGroupId)
- bytes_per_page (int)
- region (RecurrentStateRegion)
- row_bytes (int)
blocks_to_reserve()
blocks_to_reserve(num_blocks, *, enable_prefix_caching)
Returns the live block, and a checkpoint behind it with caching on.
Without prefix caching a state never checkpoints.
bound_row_copies()
bound_row_copies(src, dst)
Returns the state rows to copy, keyed where the pool is bound.
bound_row_span()
bound_row_span(block)
Returns the rows one block’s layers occupy.
region
region: RecurrentStateRegion
The state region that names this leaf and sizes its rows.
staged_input_shapes()
staged_input_shapes(batch_size, num_blocks)
Returns the row table each layer slices.
The table is layer-major, one layer per row, so a layer’s slice is a
contiguous row and the graph keeps it a view instead of materializing
a gather kernel. num_blocks is unused: a state is addressed by
row, not by block.
write_staged_inputs()
write_staged_inputs(plans, into)
Folds each request’s block into the rows its layers index.