For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python function
load_kv_manager
load_kv_manager()
max.pipelines.kv_cache.load_kv_manager(params, max_batch_size, max_seq_len, session, available_cache_memory, is_di_enabled, model_name)
Loads a KV cache manager from the given params.
Accepts both KVCacheParams (single cache) and MultiKVCacheParams
(multiple caches). The returned manager natively handles all caches
with a single BlockManager and KVConnector.
Only the Jenga manager can serve a state leaf, so a cache declaring one selects it whatever the cutover heuristic says.
TODO: remove is_di_enabled once Jenga supports DI.
-
Parameters:
-
- params (KVCacheParamInterface)
- max_batch_size (int | None)
- max_seq_len (int)
- session (InferenceSession)
- available_cache_memory (int | None)
- is_di_enabled (bool)
- model_name (str)
-
Return type:
-
PagedKVCacheManagerInterface