For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
grouped_quantize_dynamic_scaled_fp4_async
def grouped_quantize_dynamic_scaled_fp4_async[input_dtype: DType, output_dtype: DType, scales_dtype: DType, //](output_tensor: TileTensor[output_dtype, Engine=output_tensor.Engine, linear_idx_type=output_tensor.linear_idx_type], scales_tensor: TileTensor[scales_dtype, Engine=scales_tensor.Engine, linear_idx_type=scales_tensor.linear_idx_type], input_tensor: TileTensor[input_dtype, Engine=input_tensor.Engine, linear_idx_type=input_tensor.linear_idx_type], row_offsets: TileTensor[.uint32, Engine=row_offsets.Engine, linear_idx_type=row_offsets.linear_idx_type], scales_offsets: TileTensor[.uint32, Engine=scales_offsets.Engine, linear_idx_type=scales_offsets.linear_idx_type], expert_ids: TileTensor[.int32, Engine=expert_ids.Engine, linear_idx_type=expert_ids.linear_idx_type], sf_tensor: TileTensor[.float32, Engine=sf_tensor.Engine, linear_idx_type=sf_tensor.linear_idx_type], ctx: DeviceContext)
Launches the grouped per-expert quantization kernel for NVFP4/MXFP4/MXFP8 on SM100 hardware.
Sets up the TMA scale-factor descriptor and enqueues the
grouped_quantize_dynamic_scaled_fp4_async_kernel over a grid
spanning the scale-factor tile dimensions.
Parameters: