For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
grouped_quantize_dynamic_scaled_fp4_async
def grouped_quantize_dynamic_scaled_fp4_async[input_dtype: DType, output_dtype: DType, scales_dtype: DType, //, OutputEngine: TensorEngine, ScalesEngine: TensorEngine, InputEngine: TensorEngine, RowOffsetsEngine: TensorEngine, ScalesOffsetsEngine: TensorEngine, ExpertIdsEngine: TensorEngine, SfEngine: TensorEngine, RowIndicesLayoutType: TensorLayout = Layout[TypeList[Int64](), TypeList[ComptimeInt[Int(1)]]()], RowIndicesEngine: TensorEngine = DefaultEngine](output_tensor: TileTensor[output_dtype, Engine=OutputEngine, linear_idx_type=output_tensor.linear_idx_type], scales_tensor: TileTensor[scales_dtype, Engine=ScalesEngine, linear_idx_type=scales_tensor.linear_idx_type], input_tensor: TileTensor[input_dtype, Engine=InputEngine, linear_idx_type=input_tensor.linear_idx_type], row_offsets: TileTensor[.uint32, Engine=RowOffsetsEngine, linear_idx_type=row_offsets.linear_idx_type], scales_offsets: TileTensor[.uint32, Engine=ScalesOffsetsEngine, linear_idx_type=scales_offsets.linear_idx_type], expert_ids: TileTensor[.int32, Engine=ExpertIdsEngine, linear_idx_type=expert_ids.linear_idx_type], sf_tensor: TileTensor[.float32, Engine=SfEngine, linear_idx_type=sf_tensor.linear_idx_type], ctx: DeviceContext, row_indices: OptionalReg[TileTensor[.int32, RowIndicesLayoutType, ImmutAnyOrigin, Engine=RowIndicesEngine]] = None)
Launches the grouped per-expert quantization kernel for NVFP4/MXFP4/MXFP8 on SM100 hardware.
Sets up the TMA scale-factor descriptor and enqueues the
grouped_quantize_dynamic_scaled_fp4_async_kernel over a grid
spanning the scale-factor tile dimensions.
Parameters:
- input_dtype (
DType): Element type of the input activation tensor (inferred). - output_dtype (
DType): Element type of the quantized output tensor (inferred). - scales_dtype (
DType): Element type of the block scale-factor tensor (inferred). - OutputEngine (
TensorEngine): Engine policy of theoutput_tensor. - ScalesEngine (
TensorEngine): Engine policy of thescales_tensor. - InputEngine (
TensorEngine): Engine policy of theinput_tensor. - RowOffsetsEngine (
TensorEngine): Engine policy of therow_offsetstensor. - ScalesOffsetsEngine (
TensorEngine): Engine policy of thescales_offsetstensor. - ExpertIdsEngine (
TensorEngine): Engine policy of theexpert_idstensor. - SfEngine (
TensorEngine): Engine policy of thesf_tensor. - RowIndicesLayoutType (
TensorLayout): Layout of the optionalrow_indicestensor. - RowIndicesEngine (
TensorEngine): Engine policy of therow_indicestensor.