For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
RecurrentKVGroupCoordinator
RecurrentKVGroupCoordinator
class max.pipelines.kv_cache.paged_kv_cache.recurrent_coordinator.RecurrentKVGroupCoordinator(pools, leaf_id, group_id, page_size=0, enable_prefix_caching=False, *, rows=<factory>, successors=<factory>, loading=<factory>)
Bases: KVGroupCoordinatorInterface
A group whose entry is one state, held in the last slot of a row.
row[-1] is the block the recurrence runs in, and it carries no hash.
Every block behind it is null, a published checkpoint, or one a
connector hit is still loading. The recurrence reads and writes one
block, so reaching a boundary publishes the block it ran in and copies
it into a successor rather than snapshotting it.
-
Parameters:
advance()
advance(req_id, num_committed_blocks, replica_idx)
Frees every published block behind the live one, nulling its slot.
blocks_held_of_connector_hit()
blocks_held_of_connector_hit(num_hit_blocks)
Returns one block, the state at the hit’s depth.
blocks_to_allocate()
blocks_to_allocate(req_id, num_required_blocks)
Returns the live block the row lacks, and its missing successor.
checkpoint()
checkpoint(ctx, replica_idx)
Publishes the block just run in and fills the one that succeeds it.
The block the forward ran in already holds the state the boundary hash names, so it is published where that hash lands rather than snapshotted. The request continues in a freshly drawn block, copied from it. Runs before the commit, so the copy reads a block nothing has freed yet.
Empty unless the forward ended exactly on a block boundary and the row holds no unpublished predecessor already.
claim_hit_blocks()
claim_hit_blocks(desired_hashes, replica_idx)
Takes the block at the granted block count, nulling the slots behind it.
Returns empty rows when no published state stands there.
claimable_hashes()
claimable_hashes(desired_hashes)
Returns the deepest hash: a state is one boundary, not a run.
commit()
commit(req_id, hashes, last_block, replica_idx)
Publishes the checkpoints below last_block.
Forgets a loading checkpoint that a twin holding its hash replaced.
enable_prefix_caching
enable_prefix_caching: bool = False
Whether the state checkpoints at page boundaries.
Without prefix caching nothing is published, so the state stays in the block it ran in.
extend()
extend(req_id, hit_blocks, loaded_blocks, replica_idx)
Appends the hit, recording a loaded checkpoint as loading.
forward_blocks()
forward_blocks(batch, num_blocks)
Returns the block each request’s recurrence runs in.
grow()
grow(req_id, num_required_blocks, replica_idx)
Draws the live block and successor the request lacks.
Also pads the row out.
live_blocks()
live_blocks(req_id)
Returns the block per leaf the recurrence runs in, if one is drawn.
loading
The checkpoint a connector hit is still loading, per request.
Hashless until its copy lands, so without this it reads as the live block.
page_size
page_size: int = 0
the granularity a state can be committed at.
-
Type:
-
Tokens per page
release()
release(req_id, replica_idx)
Frees every block the request holds, its successor too.
resume()
resume(ctx, replica_idx)
Returns the block this forward resumes from and the one it fills.
Filled from a checkpoint when the row holds one: what a prefix hit claimed or loaded, or the predecessor a checkpoint published. Filled with zeros when the request has processed nothing and matched no hit, since a drawn block holds whatever its last request wrote.
Empty once the request runs in a block it has written itself.
shrink_to_fit()
shrink_to_fit(req_id, num_committed_blocks, replica_idx)
Refits the row to the committed blocks, keeping the live block last.
Frees every other block in the row. A loading checkpoint’s copy holds its own pin.
successors
The block each request’s next checkpoint continues the state in.
Held from admission to release while prefix caching is on, which is
the second block blocks_to_reserve budgets. Drawing it in step
instead would be a draw no admission check priced.