For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
matmul_qint4_pack_b
def matmul_qint4_pack_b[Engine: TensorEngine, //, group_size: Int](b: TileTensor[.uint8, Engine=Engine, linear_idx_type=b.linear_idx_type], b_rot: TileTensor[.uint8, Engine=Engine, linear_idx_type=b_rot.linear_idx_type])
Repacks block-wise quantized int4 weights into the tiled layout expected by the matmul_qint4 kernels.
Parameters:
- βEngine (
TensorEngine): Engine shared by both tile operands (inferred). - βgroup_size (
Int): Number of elements per quantization group.
Args:
- βb (
TileTensor[.uint8, Engine=Engine, linear_idx_type=b.linear_idx_type]): Source tensor holding packed uint8 weights with float16 scales. - βb_rot (
TileTensor[.uint8, Engine=Engine, linear_idx_type=b_rot.linear_idx_type]): Destination tensor for the repacked weights.
Raises:
If N is not a multiple of 32.