For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo struct
Struct_matmul_dynamic_block_scaled_mxfp6
struct Struct_matmul_dynamic_block_scaled_mxfp6[FP6_FORMAT: Int = Int(0), preshuffled_b: Bool = False]
Registers the mo.matmul.dynamic.block.scaled.mxfp6 graph op.
Separate from the MXFP4/MXFP8 op rather than another lane_bytes value:
both FP6 encodings put 24 bytes in a lane, so the byte count cannot choose
between them.
Parameters
- FP6_FORMAT (
Int): 0 selects E2M3, 1 selects E3M2, matchingFP6Format. - preshuffled_b (
Bool): When True,bandb_scalesmust already be in the plane-split / packed-scale layouts fromShuffler.preshuffle_b_planes/preshuffle_scale_4d(a one-time, load-time cost for the static weight; seepreshuffle_block_scaled_b_dense). Ignored (falls back to the row-major dispatch) whenMis within the decode range -- seemxfp6_block_scaled_matmul_amd.
Implemented traits
Methods
execute
static def execute[c_type: DType, a_type: DType, b_type: DType, //, target: StringSpan[ImmStaticOrigin]](c: ManagedTensorSlice[IOSpec[_, _].Output, static_spec=c.static_spec], a: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a.static_spec], b: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b.static_spec], a_scales: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=a_scales.static_spec], b_scales: ManagedTensorSlice[IOSpec[_, _].Input, static_spec=b_scales.static_spec], context: DeviceContext)