For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python class
ModuleV3PipelineModelWithKVCache
ModuleV3PipelineModelWithKVCache
class max.pipelines.lib.ModuleV3PipelineModelWithKVCache(pipeline_config, session, devices, kv_cache_config, weights, *, memory_plan, adapter=None, return_logits=ReturnLogits.LAST_TOKEN, return_hidden_states=ReturnHiddenStates.NONE, max_batch_size=1)
Bases: PipelineModelWithKVCache[BaseContextType]
The base class for a ModuleV3 model architecture that uses a KV cache.
A subclass implements _create_model_config(), which builds the
architecture’s config, and _instantiate_module(), which constructs the
root module and places it on a device or device mesh. load_model()
loads and adapts the checkpoint weights, builds the module under
F.lazy(), and compiles it with those weights. A subclass can also
override _prepare_state_dict(), _init_distributed_runtime(),
_module_default_dtype(), or _get_compile_input_types().
An encoder model without a KV cache subclasses
ModuleV3PipelineModel, and a vision-language model that compiles
several modules subclasses
ModuleV3MultiGraphPipelineModelWithKVCache.
Component models and unified speculative-decoding pipelines override
load_model().
The constructor compiles the model into model, and the default
execute() passes it model_inputs.buffers.
-
Parameters:
-
- pipeline_config (PipelineConfig) – The pipeline configuration, including the model path and its Hugging Face config.
- session (InferenceSession) – The inference session the pipeline runs in.
- devices (list[Device]) – The devices to run the model on.
- kv_cache_config (KVCacheConfig) – The KV cache configuration.
- weights (Weights) – The checkpoint weights to load.
- memory_plan (MemoryPlan) – The memory plan that sizes the KV cache.
- adapter (WeightsAdapter | None) – The weight adapter that renames checkpoint keys to the root
module’s parameter names. Defaults to
None, which loads the checkpoint keys unchanged. - return_logits (ReturnLogits) – Which logits the model returns. Defaults to
ReturnLogits.LAST_TOKEN. - return_hidden_states (ReturnHiddenStates) – Which hidden states the model returns. Defaults
to
ReturnHiddenStates.NONE. - max_batch_size (int) – The maximum number of requests in one batch. Defaults
to
1.
-
Raises:
-
ValueError – If the pipeline config enables LoRA while the KV cache config enables prefix caching, or enables LoRA for a subclass that doesn’t set
lora_modulev3.
execute()
execute(model_inputs)
Runs model on model_inputs.buffers.
-
Parameters:
-
model_inputs (ModelInputs) – The prepared inputs, whose
buffersmatch the compiled input order. -
Returns:
-
The outputs mapped by
_to_model_outputs(). -
Return type:
load_model()
load_model()
Build and compile the ModuleV3 callable.
model
model: Callable[..., Any]