For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python module
max.pipelines.architectures.mpnet
MPNet sentence transformer architecture for embeddings generation.
MPNetConfig
class max.pipelines.architectures.mpnet.MPNetConfig(*, dtype, device, pool_embeddings, huggingface_config, max_seq_len, quantization_encoding=None)
Bases: ArchConfigWithBoundedMaxSeqLen, ArchConfig
Configuration for MPNet models.
-
Parameters:
-
- dtype (DType)
- device (DeviceRef)
- pool_embeddings (bool)
- huggingface_config (AutoConfig)
- max_seq_len (int)
- quantization_encoding (SupportedEncoding | None)
DEFAULT_ENCODING
DEFAULT_ENCODING: ClassVar[SupportedEncoding] = 'bfloat16'
SUPPORTED_ENCODINGS
SUPPORTED_ENCODINGS: ClassVar[set[SupportedEncoding]] = {'bfloat16', 'float32'}
device
device: DeviceRef
dtype
dtype: DType
huggingface_config
huggingface_config: AutoConfig
initialize()
classmethod initialize(pipeline_config, model_config=None)
Initializes an MPNetConfig instance from pipeline configuration.
-
Parameters:
-
- pipeline_config (PipelineConfig) – The MAX Engine pipeline configuration.
- model_config (MAXModelConfig | None)
-
Returns:
-
An initialized MPNetConfig instance.
-
Return type:
max_seq_len
max_seq_len: int
pool_embeddings
pool_embeddings: bool
quantization_encoding
quantization_encoding: SupportedEncoding | None = None
MPNetInputs
class max.pipelines.architectures.mpnet.MPNetInputs(next_tokens_batch, attention_mask, *, kv_cache_inputs=None, lora=None, lora_buffers=(), vision_embeddings=<factory>, vision_scatter_indices=<factory>, hidden_states=None)
Bases: ModelInputs
A class representing inputs for the MPNet model.
This class encapsulates the input tensors required for the MPNet model execution:
- next_tokens_batch: A tensor containing the input token IDs
- attention_mask: A tensor containing the extended attention mask
-
Parameters:
attention_mask
attention_mask: Buffer
next_tokens_batch
next_tokens_batch: Buffer
MPNetPipelineModel
class max.pipelines.architectures.mpnet.MPNetPipelineModel(pipeline_config, session, devices, kv_cache_config, weights, adapter=None, return_logits=ReturnLogits.ALL, max_batch_size=1)
Bases: GraphPipelineModel[TextContext]
-
Parameters:
-
- pipeline_config (PipelineConfig)
- session (InferenceSession)
- devices (list[Device])
- kv_cache_config (KVCacheConfig)
- weights (Weights)
- adapter (WeightsAdapter | None)
- return_logits (ReturnLogits)
- max_batch_size (int)
batch_processor_cls
batch_processor_cls
alias of MPNetBatchProcessor
execute()
execute(model_inputs)
Executes the graph with the given inputs.
-
Parameters:
-
model_inputs (ModelInputs) – The model inputs to execute, containing tensors and any other required data for model execution.
-
Returns:
-
ModelOutputs containing the pipeline’s output tensors.
-
Return type:
This is an abstract method that must be implemented by concrete PipelineModels to define their specific execution logic.
model
model: Model
model_config_cls
model_config_cls
alias of MPNetConfig