For the complete documentation index, see llms.txt. Markdown versions of all pages are available by appending .md to any URL (e.g. /get-started.md).
Mojo function
enqueue_apple_gemv_config
def enqueue_apple_gemv_config[c_type: DType, in_type: DType, //, *, tile_m: Int, rows_per_warp: Int, tile_k: Int = Int(8), unroll: Int = Int(2), warps_per_block: Int = Int(8), elementwise_lambda_fn: Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None] = None](c: TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type], a: TileTensor[in_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type], weight: TileTensor[in_type, Engine=weight.Engine, address_space=weight.address_space, linear_idx_type=weight.linear_idx_type], ctx: DeviceContext)
Enqueues apple_gemv_kernel with an explicit launch configuration.
Parameters:
- c_type (
DType): Output element type. Accumulation is fp32. - in_type (
DType): Activation and weight element type (bf16 or fp16). - tile_m (
Int): Largest activation row count the launch supports. - rows_per_warp (
Int): Weight rows (output columns) per warp. - tile_k (
Int): Elements per lane per K chunk. - unroll (
Int): K chunks per loop iteration. - warps_per_block (
Int): Warps per threadgroup. - elementwise_lambda_fn (
Optional[def[dtype: DType, width: SIMDLength, *, alignment: Int = Int(1)](IndexList[Int(2)], SIMD[dtype, width]) capturing thin -> None]): Optional epilogue applied to each output.
Args:
- c (
TileTensor[c_type, Engine=c.Engine, address_space=c.address_space, linear_idx_type=c.linear_idx_type]): Output[M, N]. - a (
TileTensor[in_type, Engine=a.Engine, address_space=a.address_space, linear_idx_type=a.linear_idx_type]): Activation[M, K]withM <= tile_mandK % tile_k == 0. - weight (
TileTensor[in_type, Engine=weight.Engine, address_space=weight.address_space, linear_idx_type=weight.linear_idx_type]): Weight[N, K]. - ctx (
DeviceContext): Device context to enqueue on.