IMPORTANT: To view this page as Markdown, append `.md` to the URL (e.g. /get-started.md). For the complete documentation index, see llms.txt.
Skip to main content
For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).

Mojo function

quantize_mxfp6_lane_group

def quantize_mxfp6_lane_group[in_dtype: DType, width: Int, //, scales_dtype: DType, fmt: FP6Format, *, SF_VECTOR_SIZE: Int = Int(32)](val: SIMD[in_dtype, width]) -> Tuple[SIMD[.uint8, width], Scalar[scales_dtype]]

Quantizes one thread's slice of an MX block to MXFP6, cooperatively.

The MXFP6 counterpart of quantize_mxfp8_lane_group, for callers that quantize inside another kernel's epilogue and therefore hold fewer than a whole block per thread. quantize_mxfp6_amd gives one thread all 32 elements and needs no cross-lane step; here the block max spans SF_VECTOR_SIZE // width lanes, so the caller's thread-to-column map has to land each MX block on one aligned lane group.

Returns FP6 codes, one per element, rather than packed bytes: four codes share three bytes, so the group -- not the byte -- is the smallest addressable unit of packed FP6 (see pack_fp6_x4). Only the caller knows whether its store offset is group-aligned, so packing stays at the store site. A caller holding a multiple of four elements can pack its own codes with no further cross-lane traffic.

The scale and dead-block handling mirror quantize_mxfp6_amd exactly, so a caller that packs the returned codes reproduces that kernel byte for byte.

Parameters:

  • ​in_dtype (DType): Element type of the incoming values (bfloat16).
  • ​width (Int): Number of elements this thread holds.
  • ​scales_dtype (DType): Block-scale type (float8_e8m0fnu).
  • ​fmt (FP6Format): The FP6 encoding to produce (E2M3 or E3M2).
  • ​SF_VECTOR_SIZE (Int): Elements covered by one block scale (32).

Args:

Returns:

Tuple[SIMD[.uint8, width], Scalar[scales_dtype]]: (codes, e8m0_scale). Every lane in the group returns the same scale; the caller stores it once, from the group's first lane.

Was this page helpful?