For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python module
max.pipelines.architectures.qwen_image
Qwen-Image diffusion architecture for image generation.
QwenImageArchConfig
class max.pipelines.architectures.qwen_image.QwenImageArchConfig(*, pipeline_config, quantization_encoding=None)
Bases: ArchConfig
Pipeline-level config for QwenImage (implements ArchConfig; no KV cache).
-
Parameters:
-
- pipeline_config (PipelineConfig)
- quantization_encoding (SupportedEncoding | None)
DEFAULT_ENCODING
DEFAULT_ENCODING: ClassVar[SupportedEncoding] = 'bfloat16'
SUPPORTED_ENCODINGS
SUPPORTED_ENCODINGS: ClassVar[set[SupportedEncoding]] = {'bfloat16'}
calculate_max_seq_len()
classmethod calculate_max_seq_len(huggingface_config, model_config)
Returns the resolved maximum sequence length.
Bounds or defaults the user’s model_config.max_length with the
model’s own limits. Construction runs this once and stores the
result on model_config.max_length; memory planning may lower
it further, but only on the memory plan.
-
Parameters:
-
- huggingface_config (AutoConfig) – The HuggingFace config to read model bounds from.
- model_config (MAXModelConfig) – The model config whose
max_lengthcarries the user’s setting.
-
Return type:
get_max_seq_len()
get_max_seq_len()
Returns the effective maximum sequence length for the model.
For configs that store a deployment length, this is the value
initialize received; for metadata-only configs it derives from
the checkpoint.
-
Return type:
initialize()
classmethod initialize(pipeline_config, model_config=None, *, max_seq_len)
Initialize the config from a PipelineConfig.
-
Parameters:
-
- pipeline_config (PipelineConfig) – The pipeline configuration.
- model_config (MAXModelConfig | None) – The model configuration to read from. When
None(the default),pipeline_config.modelis used. Pass an explicit config (e.g.pipeline_config.draft_model) to initialize the arch config for a different model. - max_seq_len (int) – The effective maximum sequence length to store on
the config. The value is received, never derived here: the
pipeline model passes the memory plan’s VRAM-clamped length,
while memory planning (which runs before a plan exists)
passes the construction-resolved
model_config.max_length. Configs whose sequence length is pure model metadata (e.g. diffusion components) ignore it.
-
Return type:
pipeline_config
pipeline_config: PipelineConfig
quantization_encoding
quantization_encoding: SupportedEncoding | None = None
QwenImageConfig
class max.pipelines.architectures.qwen_image.QwenImageConfig(*, config_file=None, section_name=None, patch_size=2, in_channels=64, out_channels=None, num_layers=60, attention_head_dim=128, num_attention_heads=24, joint_attention_dim=3584, guidance_embeds=False, axes_dims_rope=(16, 56, 56), rope_theta=10000, zero_cond_t=False, eps=1e-06, dtype=bfloat16, device=<factory>)
Bases: QwenImageConfigBase
-
Parameters:
-
- config_file (str | None)
- section_name (str | None)
- patch_size (int)
- in_channels (int)
- out_channels (int | None)
- num_layers (int)
- attention_head_dim (int)
- num_attention_heads (int)
- joint_attention_dim (int)
- guidance_embeds (bool)
- axes_dims_rope (tuple[int, ...])
- rope_theta (int)
- zero_cond_t (bool)
- eps (float)
- dtype (DType)
- device (DeviceRef)
generate()
static generate(config_dict, encoding, devices)
model_config
model_config: ClassVar[ConfigDict] = {'arbitrary_types_allowed': True, 'extra': 'forbid', 'strict': False}
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
QwenImageConfigBase
class max.pipelines.architectures.qwen_image.QwenImageConfigBase(*, config_file=None, section_name=None, patch_size=2, in_channels=64, out_channels=None, num_layers=60, attention_head_dim=128, num_attention_heads=24, joint_attention_dim=3584, guidance_embeds=False, axes_dims_rope=(16, 56, 56), rope_theta=10000, zero_cond_t=False, eps=1e-06, dtype=bfloat16, device=<factory>)
Bases: MAXModelConfigBase
-
Parameters:
-
- config_file (str | None)
- section_name (str | None)
- patch_size (int)
- in_channels (int)
- out_channels (int | None)
- num_layers (int)
- attention_head_dim (int)
- num_attention_heads (int)
- joint_attention_dim (int)
- guidance_embeds (bool)
- axes_dims_rope (tuple[int, ...])
- rope_theta (int)
- zero_cond_t (bool)
- eps (float)
- dtype (DType)
- device (DeviceRef)
attention_head_dim
attention_head_dim: int
axes_dims_rope
device
device: DeviceRef
dtype
dtype: DType
eps
eps: float
guidance_embeds
guidance_embeds: bool
in_channels
in_channels: int
joint_attention_dim
joint_attention_dim: int
model_config
model_config: ClassVar[ConfigDict] = {'arbitrary_types_allowed': True, 'extra': 'forbid', 'strict': False}
Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
num_attention_heads
num_attention_heads: int
num_layers
num_layers: int
out_channels
patch_size
patch_size: int
rope_theta
rope_theta: int
zero_cond_t
zero_cond_t: bool
QwenImageTransformerModel
class max.pipelines.architectures.qwen_image.QwenImageTransformerModel(config, encoding, devices, weights, session)
Bases: ComponentModel
-
Parameters:
load_model()
load_model()
Load and return a runtime model instance.