For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
matmul_qint4
def matmul_qint4[group_size: Int, elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None](a: TileTensor[.float32, Engine=a.Engine, linear_idx_type=a.linear_idx_type], b: TileTensor[.uint8, Engine=b.Engine, linear_idx_type=b.linear_idx_type], c: TileTensor[.float32, Engine=c.Engine, linear_idx_type=c.linear_idx_type], ctx: Optional[DeviceContext] = None)
Computes a matrix multiply of a float32 A matrix against block-wise quantized int4 B weights, producing a float32 result.
Dispatches to an architecture-specific kernel (VNNI, AVX2, NEON i8mm, or NEON dotprod) at compile time.
Parameters:
- group_size (
Int): Number of elements per quantization group. - elementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue applied to each output element.
Args:
- a (
TileTensor[.float32, Engine=a.Engine, linear_idx_type=a.linear_idx_type]): Input A tensor in float32. - b (
TileTensor[.uint8, Engine=b.Engine, linear_idx_type=b.linear_idx_type]): Input B tensor holding packed uint8 int4 weights. - c (
TileTensor[.float32, Engine=c.Engine, linear_idx_type=c.linear_idx_type]): Output C tensor in float32. - ctx (
Optional[DeviceContext]): Optional device context for parallel execution.