For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Python function
build_modelopt_nvfp4_config
build_modelopt_nvfp4_config()
max.pipelines.weights.build_modelopt_nvfp4_config(huggingface_config, state_dict, resolved_quant_config, *, state_dict_name_prefix='', ignored_modules_prefix='model.')
Builds the QuantConfig for modelopt’s NVFP4 two-level scheme.
Skips the whole-checkpoint quant_algo == "NVFP4" check that a
MIXED_PRECISION export cannot pass, so the caller must have
established which modules are NVFP4: handing this config to one
quantized some other way reads its payload as packed FP4 and yields
garbage.
-
Parameters:
-
- huggingface_config (AutoConfig) – The config the layer count is read from.
- state_dict (Mapping[str, WeightData]) – The checkpoint weights, read for the bias and embedding dtypes.
- resolved_quant_config (Mapping[str, Any] | None) – The resolved quantization config, from
resolve_hf_quant_config(). Only itsignoreglobs are read;Nonemeans nothing is ignored. - state_dict_name_prefix (str) – The optional prefix on the
state_dictkeys. - ignored_modules_prefix (str) – The prefix the
ignoreglobs are written against.
-
Returns:
-
The NVFP4 config, with the fused-kernel flags left at their defaults for
apply_fused_kernel_flags()to set. -
Return type: