For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
fp4_gemv_kernel
def fp4_gemv_kernel[c_type: DType, c_layout: TensorLayout, a_layout: TensorLayout, p_layout: TensorLayout, s_layout: TensorLayout, c_engine: TensorEngine, a_engine: TensorEngine, p_engine: TensorEngine, s_engine: TensorEngine, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]](c: TileTensor[c_type, c_layout, MutAnyOrigin, Engine=c_engine], a: TileTensor[.bfloat16, a_layout, ImmutAnyOrigin, Engine=a_engine], packed: TileTensor[.uint8, p_layout, ImmutAnyOrigin, Engine=p_engine], scales: TileTensor[.float8_e4m3fn, s_layout, ImmutAnyOrigin, Engine=s_engine], n_arg: Int32, k_arg: Int32)
One warp per output column; 32 lanes stride down K decoding FP4 -> fp32.
c is [1, N], a the bf16 activation [1, K], packed the FP4 weight
[N, K//2] (lo-nibble first), scales the FP8-E4M3 block scales
[N, ceil(K/16)]. Accumulation is fp32.
Parameters:
- c_type (
DType): Output element type (fp16, bf16, fp32). Accumulation is fp32. - c_layout (
TensorLayout):TileTensorlayout of the outputc. - a_layout (
TensorLayout):TileTensorlayout of the activationa. - p_layout (
TensorLayout):TileTensorlayout of the packed FP4 weight. - s_layout (
TensorLayout):TileTensorlayout of the FP8 block scales. - c_engine (
TensorEngine):TensorEngineof the outputc. - a_engine (
TensorEngine):TensorEngineof the activationa. - p_engine (
TensorEngine):TensorEngineof the packed FP4 weight. - s_engine (
TensorEngine):TensorEngineof the FP8 block scales. - elementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional fused epilogue applied on the width-1 store.
Args:
- c (
TileTensor[c_type, c_layout, MutAnyOrigin, Engine=c_engine]): Output tile tensor[1, N]receiving the GEMV result. - a (
TileTensor[.bfloat16, a_layout, ImmutAnyOrigin, Engine=a_engine]): Bf16 activation tile tensor[1, K], the single activation row. - packed (
TileTensor[.uint8, p_layout, ImmutAnyOrigin, Engine=p_engine]): FP4-packed weight tile tensor[N, K//2](lo-nibble first). - scales (
TileTensor[.float8_e4m3fn, s_layout, ImmutAnyOrigin, Engine=s_engine]): FP8-E4M3 block scales tile tensor[N, ceil(K/16)]. - n_arg (
Int32): Number of output columns (rows of the transposed weight). - k_arg (
Int32): Inner dimension length; must be a multiple ofNVFP4_SF_VECTOR_SIZE(16).