For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
quantize_dynamic_block_scaled_mxfp4
def quantize_dynamic_block_scaled_mxfp4[in_dtype: DType](output: TileTensor[.uint8, Engine=output.Engine, address_space=output.address_space, linear_idx_type=output.linear_idx_type], output_scales: TileTensor[.float8_e8m0fnu, Engine=output_scales.Engine, address_space=output_scales.address_space, linear_idx_type=output_scales.linear_idx_type], input: TileTensor[in_dtype, Engine=input.Engine, address_space=input.address_space, linear_idx_type=input.linear_idx_type], ctx: DeviceContext)
Launches the AMD CDNA4 MXFP4 quantization kernel over a flat input buffer.
Enqueues quantize_dynamic_block_scaled_mxfp4_kernel with a
512-thread block grid, emitting a trace event for profiling.
Parameters:
- in_dtype (
DType): Element type of the input activation tensor (inferred). Must bebfloat16.