IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Python class

ModuleV3PipelineModelWithKVCache

ModuleV3PipelineModelWithKVCache​

class max.pipelines.lib.ModuleV3PipelineModelWithKVCache(pipeline_config, session, devices, kv_cache_config, weights, *, memory_plan, adapter=None, return_logits=ReturnLogits.LAST_TOKEN, return_hidden_states=ReturnHiddenStates.NONE, max_batch_size=1)

source

Bases: PipelineModelWithKVCache[BaseContextType]

The base class for a ModuleV3 model architecture that uses a KV cache.

A subclass implements _create_model_config(), which builds the architecture’s config, and _instantiate_module(), which constructs the root module and places it on a device or device mesh. load_model() loads and adapts the checkpoint weights, builds the module under F.lazy(), and compiles it with those weights. A subclass can also override _prepare_state_dict(), _init_distributed_runtime(), _module_default_dtype(), or _get_compile_input_types().

An encoder model without a KV cache subclasses ModuleV3PipelineModel, and a vision-language model that compiles several modules subclasses ModuleV3MultiGraphPipelineModelWithKVCache. Component models and unified speculative-decoding pipelines override load_model().

The constructor compiles the model into model, and the default execute() passes it model_inputs.buffers.

Parameters:

  • pipeline_config (PipelineConfig) – The pipeline configuration, including the model path and its Hugging Face config.
  • session (InferenceSession) – The inference session the pipeline runs in.
  • devices (list[Device]) – The devices to run the model on.
  • kv_cache_config (KVCacheConfig) – The KV cache configuration.
  • weights (Weights) – The checkpoint weights to load.
  • memory_plan (MemoryPlan) – The memory plan that sizes the KV cache.
  • adapter (WeightsAdapter | None) – The weight adapter that renames checkpoint keys to the root module’s parameter names. Defaults to None, which loads the checkpoint keys unchanged.
  • return_logits (ReturnLogits) – Which logits the model returns. Defaults to ReturnLogits.LAST_TOKEN.
  • return_hidden_states (ReturnHiddenStates) – Which hidden states the model returns. Defaults to ReturnHiddenStates.NONE.
  • max_batch_size (int) – The maximum number of requests in one batch. Defaults to 1.

Raises:

ValueError – If the pipeline config enables LoRA while the KV cache config enables prefix caching, or enables LoRA for a subclass that doesn’t set lora_modulev3.

execute()​

execute(model_inputs)

source

Runs model on model_inputs.buffers.

Parameters:

model_inputs (ModelInputs) – The prepared inputs, whose buffers match the compiled input order.

Returns:

The outputs mapped by _to_model_outputs().

Return type:

ModelOutputs

load_model()​

load_model()

source

Build and compile the ModuleV3 callable.

Return type:

Callable[[…], Any]

model​

model: Callable[..., Any]

source