For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
quantize_mxfp6_lane_group
def quantize_mxfp6_lane_group[in_dtype: DType, width: Int, //, scales_dtype: DType, fmt: FP6Format, *, SF_VECTOR_SIZE: Int = Int(32)](val: SIMD[in_dtype, width]) -> Tuple[SIMD[.uint8, width], Scalar[scales_dtype]]
Quantizes one thread's slice of an MX block to MXFP6, cooperatively.
The MXFP6 counterpart of quantize_mxfp8_lane_group, for callers that
quantize inside another kernel's epilogue and therefore hold fewer than a
whole block per thread. quantize_mxfp6_amd gives one thread all 32
elements and needs no cross-lane step; here the block max spans
SF_VECTOR_SIZE // width lanes, so the caller's thread-to-column map has to
land each MX block on one aligned lane group.
Returns FP6 codes, one per element, rather than packed bytes: four codes
share three bytes, so the group -- not the byte -- is the smallest
addressable unit of packed FP6 (see pack_fp6_x4). Only the caller knows
whether its store offset is group-aligned, so packing stays at the store
site. A caller holding a multiple of four elements can pack its own codes
with no further cross-lane traffic.
The scale and dead-block handling mirror quantize_mxfp6_amd exactly, so a
caller that packs the returned codes reproduces that kernel byte for byte.
Parameters:
- in_dtype (
DType): Element type of the incoming values (bfloat16). - width (
Int): Number of elements this thread holds. - scales_dtype (
DType): Block-scale type (float8_e8m0fnu). - fmt (
FP6Format): The FP6 encoding to produce (E2M3 or E3M2). - SF_VECTOR_SIZE (
Int): Elements covered by one block scale (32).
Args:
- val (
SIMD[in_dtype, width]): This thread'swidthcontiguous elements.
Returns:
Tuple[SIMD[.uint8, width], Scalar[scales_dtype]]: (codes, e8m0_scale). Every lane in the group returns the same scale;
the caller stores it once, from the group's first lane.