For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
fp4_materialize_kernel
def fp4_materialize_kernel[out_type: DType, w_layout: TensorLayout, s_layout: TensorLayout, out_layout: TensorLayout, w_engine: TensorEngine, s_engine: TensorEngine, out_engine: TensorEngine](out_w: TileTensor[out_type, out_layout, MutAnyOrigin, Engine=out_engine], packed: TileTensor[.uint8, w_layout, ImmutAnyOrigin, Engine=w_engine], scales: TileTensor[.float8_e4m3fn, s_layout, ImmutAnyOrigin, Engine=s_engine])
Materializes the packed-FP4 weight into a dense [N, K] out_type buffer.
One thread per output element (n, k). packed is [N, K//2] (lo-nibble
first), scales is [N, K//16]. Used by the Stage-1 oracle: it dequants the
weight to bf16 so the EXISTING AppleM5MatMul can consume it, proving the
dequant math against a host reference before the fused loader is written.
Parameters:
- out_type (
DType): Output element dtype (bf16 for the Apple W4A16 path). - w_layout (
TensorLayout):TileTensorlayout ofpacked, a plain rank-2[N, K//2]uint8buffer. - s_layout (
TensorLayout):TileTensorlayout ofscales, a plain rank-2[N, K//16]float8_e4m3fnbuffer. - out_layout (
TensorLayout):TileTensorlayout ofout_w, a plain rank-2[N, K]buffer ofout_type. - w_engine (
TensorEngine):TensorEngineofpacked. - s_engine (
TensorEngine):TensorEngineofscales. - out_engine (
TensorEngine):TensorEngineofout_w.
Args:
- out_w (
TileTensor[out_type, out_layout, MutAnyOrigin, Engine=out_engine]): Output weight buffer of shape[N, K]and dtypeout_type, written one element per thread. - packed (
TileTensor[.uint8, w_layout, ImmutAnyOrigin, Engine=w_engine]): Input packed FP4 weights asuint8of shape[N, K//2], two E2M1 nibbles per byte with the low nibble first. - scales (
TileTensor[.float8_e4m3fn, s_layout, ImmutAnyOrigin, Engine=s_engine]): Per-block FP8-E4M3 scales of shape[N, K//16], applied asabs(scale)over each 16-element K block.