For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python module
max.pipelines.architectures.qwen2_5vl
Qwen2.5-VL vision-language architecture for multimodal text generation.
Qwen2_5VLConfig
class max.pipelines.architectures.qwen2_5vl.Qwen2_5VLConfig(*, devices, image_token_id, video_token_id, vision_start_token_id, spatial_merge_size, tokens_per_second, mrope_section, vision_config, llm_config, quantization_encoding=None)
Bases: ArchVLConfigWithTextSubconfig, ArchConfigWithKVCache
Configuration for Qwen2.5VL models.
-
Parameters:
-
- devices (list[DeviceRef])
- image_token_id (int)
- video_token_id (int)
- vision_start_token_id (int)
- spatial_merge_size (int)
- tokens_per_second (int)
- mrope_section (list[int])
- vision_config (VisionConfig)
- llm_config (Llama3Config)
- quantization_encoding (SupportedEncoding | None)
DEFAULT_ENCODING
DEFAULT_ENCODING: ClassVar[SupportedEncoding] = 'bfloat16'
SUPPORTED_ENCODINGS
SUPPORTED_ENCODINGS: ClassVar[set[SupportedEncoding]] = {'bfloat16', 'float32', 'float8_e4m3fn'}
devices
Devices that the Qwen2.5VL model is parallelized over.
finalize()
finalize(huggingface_config, pipeline_config, llm_state_dict, vision_state_dict, return_logits, norm_method='rms_norm')
Finalize the Qwen2_5VLConfig instance with state_dict dependent fields.
-
Parameters:
-
- huggingface_config (AutoConfig) – HuggingFace model configuration.
- pipeline_config (PipelineConfig) – The MAX Engine pipeline configuration.
- llm_state_dict (dict[str, WeightData]) – Language model weights dictionary.
- vision_state_dict (dict[str, WeightData]) – Vision encoder weights dictionary.
- return_logits (ReturnLogits) – Return logits configuration.
- norm_method (Literal['rms_norm', 'layer_norm']) – Normalization method.
-
Return type:
-
None
get_kv_params()
get_kv_params()
Returns the KV cache parameters from the embedded LLM config.
-
Return type:
get_num_layers()
static get_num_layers(huggingface_config)
-
Parameters:
-
huggingface_config (AutoConfig)
-
Return type:
image_token_id
image_token_id: int
Token ID used for image placeholders in the input sequence.
initialize()
classmethod initialize(pipeline_config, model_config=None)
Initializes a Qwen2_5VLConfig instance from pipeline configuration.
-
Parameters:
-
- pipeline_config (PipelineConfig) – The MAX Engine pipeline configuration.
- model_config (MAXModelConfig | None)
-
Returns:
-
A Qwen2_5VLConfig instance with fields initialized from config.
-
Return type:
initialize_from_config()
classmethod initialize_from_config(pipeline_config, huggingface_config)
Initializes a Qwen2_5VLConfig from pipeline and HuggingFace configs.
This method creates a config instance with all fields that can be determined from the pipeline and HuggingFace configurations, without needing the state_dict. Fields that depend on the state_dict should be set via the finalize() method.
-
Parameters:
-
- pipeline_config (PipelineConfig) – The MAX Engine pipeline configuration.
- huggingface_config (AutoConfig) – HuggingFace model configuration.
-
Returns:
-
A Qwen2_5VLConfig instance ready for finalization.
-
Return type:
llm_config
llm_config: Llama3Config
Language model configuration using Llama3 architecture.
mrope_section
List of indices for the mrope section.
quantization_encoding
quantization_encoding: SupportedEncoding | None = None
spatial_merge_size
spatial_merge_size: int
Size parameter for spatial merging of vision features.
tokens_per_second
tokens_per_second: int
Number of tokens per second.
video_token_id
video_token_id: int
Token ID used for video placeholders in the input sequence.
vision_config
vision_config: VisionConfig
Vision encoder configuration.
vision_start_token_id
vision_start_token_id: int
Token ID that marks the start of vision content.
Qwen2_5VLInputs
class max.pipelines.architectures.qwen2_5vl.Qwen2_5VLInputs(tokens, input_row_offsets, signal_buffers, position_ids, return_n_logits, image_token_indices=None, pixel_values=None, window_index=None, vision_position_ids=None, max_grid_size=None, cu_seqlens=None, cu_window_seqlens=None, max_seqlen=None, max_window_seqlen=None, *, kv_cache_inputs, lora=None, lora_buffers=(), vision_embeddings=<factory>, vision_scatter_indices=<factory>, hidden_states=None)
Bases: ModelInputs
A class representing inputs for the Qwen2.5VL model.
This class encapsulates the input tensors required for the Qwen2.5VL model execution, including both text and vision inputs. Vision inputs are optional and can be None for text-only processing.
-
Parameters:
-
- tokens (Buffer)
- input_row_offsets (list[Buffer])
- signal_buffers (list[Buffer])
- position_ids (Buffer)
- return_n_logits (Buffer)
- image_token_indices (list[Buffer] | None)
- pixel_values (list[Buffer] | None)
- window_index (list[Buffer] | None)
- vision_position_ids (list[Buffer] | None)
- max_grid_size (list[Buffer] | None)
- cu_seqlens (list[Buffer] | None)
- cu_window_seqlens (list[Buffer] | None)
- max_seqlen (list[Buffer] | None)
- max_window_seqlen (list[Buffer] | None)
- kv_cache_inputs (KVCacheInputsInterface[Buffer, Buffer])
- lora (LoRAInputs | None)
- lora_buffers (tuple[Buffer, ...])
- vision_embeddings (list[Buffer])
- vision_scatter_indices (list[Buffer])
- hidden_states (Buffer | list[Buffer] | None)
cu_seqlens
Cumulative sequence lengths for full attention.
cu_window_seqlens
Cumulative window sequence lengths for window attention.
has_vision_inputs
property has_vision_inputs: bool
Check if this input contains vision data.
image_token_indices
Per-device pre-computed multimodal merge indices for the image embeddings.
These are the locations of the image_token_id in the inputs fed to the model.
Some indices may be negative, which means that they are ignored by the multimodal merge.
input_row_offsets
Per-device tensors containing the offsets for each row in the ragged input sequence.
max_grid_size
Maximum grid size for vision inputs.
max_seqlen
Maximum sequence length for full attention for vision inputs.
max_window_seqlen
Maximum sequence length for window attention for vision inputs.
pixel_values
Pixel values for vision inputs.
position_ids
position_ids: Buffer
3D RoPE position IDs for the decoder.
return_n_logits
return_n_logits: Buffer
Number of logits to return, used by speculative decoding for example.
signal_buffers
Device buffers used for synchronization in communication collectives.
tokens
tokens: Buffer
Tensor containing the input token IDs.
vision_position_ids
1D RoPE position IDs for the visual inputs.
window_index
Window indices for vision attention mechanism.
Qwen2_5VLModel
class max.pipelines.architectures.qwen2_5vl.Qwen2_5VLModel(pipeline_config, session, devices, kv_cache_config, weights, adapter=None, return_logits=ReturnLogits.LAST_TOKEN, max_batch_size=1)
Bases: AlwaysSignalBuffersMixin, MultiGraphPipelineModelWithKVCache[TextAndVisionContext]
A Qwen2.5VL pipeline model for multimodal text generation.
-
Parameters:
-
- pipeline_config (PipelineConfig)
- session (InferenceSession)
- devices (list[Device])
- kv_cache_config (KVCacheConfig)
- weights (Weights)
- adapter (WeightsAdapter | None)
- return_logits (ReturnLogits)
- max_batch_size (int)
batch_processor_cls
batch_processor_cls
alias of Qwen2_5VLBatchProcessor
execute()
execute(model_inputs)
Executes the Qwen2.5VL model with the prepared inputs.
-
Parameters:
-
model_inputs (ModelInputs)
-
Return type:
language_model
language_model: Model
The compiled language model for text generation.
load_model()
load_model(session)
Override: incompatible tower graph capture signature.
_build_* is (module) -> Graph (not the base
(config, state_dict, module) -> (Graph, registry)), because
graphs are captured from a pre-instantiated Qwen2_5VL module after
load_state_dict in _create_model_config. Registry keys come
from the tower splits in _load_state_dict, not from per-tower
nn.state_dict() returns.
-
Parameters:
-
session (InferenceSession)
-
Return type:
model_config
model_config: Qwen2_5VLConfig | None
The Qwen2.5VL model configuration.
model_config_cls
model_config_cls
alias of Qwen2_5VLConfig
vision_model
The compiled vision model for processing images.
VisionConfig
class max.pipelines.architectures.qwen2_5vl.VisionConfig(dtype, llm_dtype, devices, patch_size, temporal_patch_size, in_channels, hidden_size, num_attention_heads, depth, intermediate_size, out_hidden_size, fullatt_block_indexes, rms_norm_eps, window_size, spatial_merge_size, quant_config=None)
Bases: object
Base configuration for Qwen2.5VL models with required fields.
-
Parameters:
-
- dtype (DType)
- llm_dtype (DType)
- devices (list[DeviceRef])
- patch_size (int)
- temporal_patch_size (int)
- in_channels (int)
- hidden_size (int)
- num_attention_heads (int)
- depth (int)
- intermediate_size (int)
- out_hidden_size (int)
- fullatt_block_indexes (list[int])
- rms_norm_eps (float)
- window_size (int)
- spatial_merge_size (int)
- quant_config (QuantConfig | None)
depth
depth: int
Number of vision transformer layers.
devices
Devices that the Qwen2.5VL vision encoder model is parallelized over.
dtype
dtype: DType
DType of the Qwen2.5VL vision model weights.
finalize()
finalize(huggingface_config, vision_state_dict, vision_dtype, llm_dtype)
Finalize VisionConfig with state_dict dependent fields.
-
Parameters:
-
- huggingface_config (AutoConfig)
- vision_state_dict (dict[str, WeightData])
- vision_dtype (DType)
- llm_dtype (DType)
-
Return type:
-
None
fullatt_block_indexes
Indexes of the full attention blocks in the vision encoder.
hidden_size
hidden_size: int
Hidden size of the vision encoder.
in_channels
in_channels: int
Vision transformer number of input channels.
initialize_from_config()
classmethod initialize_from_config(pipeline_config, hf_vision_config)
Initialize VisionConfig from HuggingFace vision config.
Note: dtype fields will be set to defaults and should be updated via finalize() once state_dict is available.
-
Parameters:
-
- pipeline_config (PipelineConfig)
- hf_vision_config (AutoConfig)
-
Return type:
intermediate_size
intermediate_size: int
Intermediate size in the vision encoder’s feed-forward layers.
llm_dtype
llm_dtype: DType
DType of the Qwen2.5VL language model weights.
num_attention_heads
num_attention_heads: int
Number of attention heads in the vision encoder.
out_hidden_size
out_hidden_size: int
Output hidden size of the vision encoder. Also the hidden size of the language model.
patch_size
patch_size: int
Vision transformer patch size.
quant_config
quant_config: QuantConfig | None = None
Scaled quantization configuration for the vision encoder.
rms_norm_eps
rms_norm_eps: float
Epsilon for layer normalization.
spatial_merge_size
spatial_merge_size: int
Spatial merge size for the vision encoder.
temporal_patch_size
temporal_patch_size: int
Vision transformer temporal patch size.
window_size
window_size: int
Window size for the vision encoder.